Automatically deployed information technology (IT) system and method with enhanced security

The IT system addresses setup, configuration, and maintenance challenges by using a controller for automated management, enhancing security and scalability, and reducing downtime.

JP2025111560APending Publication Date: 2025-07-30NET THUNDER LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025068357
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-11
Filing Date
2025-04-17
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Current IT systems face challenges in setup, configuration, security, scalability, and maintenance, including vulnerabilities in bare metal cloud nodes, human errors, and difficulties in troubleshooting and recovering from outages, which can lead to significant downtime and economic losses.

Method used

An IT system with a controller that manages interoperability and security by using self-assembly rules, templates, and system states to automate setup, configuration, and maintenance, ensuring flexibility, security, and efficient resource management.

Benefits of technology

The system reduces human errors, enhances security, and improves scalability by automating setup, configuration, and maintenance, enabling efficient resource allocation and rapid recovery from outages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111560000001_ABST
    Figure 2025111560000001_ABST
Patent Text Reader

Abstract

To provide a system and method for allowing the flexibility, reducing the variability and human error and increasing the system security in IT systems.SOLUTION: An IT system includes a first service and a second service, wherein the first and second services have dependency on each other, the first service includes a depended service on the second service, and the second service includes a dependent service on the first service. The controller of the IT system is configured to manage an interoperability of the first service with respect to the second service.SELECTED DRAWING: Figure 19B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This patent application claims priority to U.S. Provisional Patent Application Serial No. 62 / 860,148, filed on June 11, 2019, entitled "Automatically Deployed Information Technology (IT) System and Method with Enhanced Security", the entire disclosure of which is incorporated herein by reference.

Background Art

[0002] The demand, utilization, and needs for computing have increased exponentially in the past few decades. Along with this, due to the need for greater storage, speed, computing power, applications, and accessibility, the field of computing has changed rapidly, and tools are provided to entities of various types and sizes. As a result, public virtual computing systems and cloud computing systems have been developed so that more computing resources can be provided to a large number of users and user types. This exponential growth is expected to continue. At the same time, as the risks of failures and security increase, the setup, management, change management, and updates of the infrastructure become more complex and costly. Scalability, that is, growing the system over time, has also become a major challenge in the field of information technology.

[0003] Most problems with IT systems, many of which are related to performance and security, can be difficult to diagnose and address. Constraints on time and resources allowed for system setup, configuration, and deployment can lead to errors and future IT problems. Over time, many different administrators may be involved in changes, patch applications, or updates to IT systems, including users, applications, services, security, software, and hardware. Often, the documentation and history of configurations and changes are insufficient or lost, making it difficult to understand later how a particular system is configured and functions. This can make future changes or troubleshooting difficult. When a problem or outage occurs, IT configurations and settings can be difficult to recover and reproduce. Also, system administrators can easily make mistakes, such as incorrect commands or other errors, which can cause computers and web databases and services to go down. Additionally, while an increased risk of security breaches is not uncommon, changes, updates, and patches to avoid security breaches can cause unwanted downtime.

[0004] When critical infrastructure is in place, functioning, and operating, the costs or risks often may seem to outweigh the benefits of changing the system. The problems associated with making changes to an operating IT system or environment can cause significant, sometimes catastrophic, problems for users or entities that depend on these systems. At a minimum, troubleshooting and resolving outages or problems that occur during change management can require significant time, personnel, and financial resources. Technical problems that can occur when changes are made to an operating environment can have a cascading effect and may not be resolved by simply reversing the applied changes. Many of these issues contribute to the inability to quickly rebuild a system when an outage exists during change management.

[0005] Furthermore, bare metal cloud nodes or resources within an IT system may be vulnerable to security issues, or be exposed to or accessible by unauthorized users. Hackers, attackers, or unauthorized users may pivot from such nodes or resources to access or hack into the network connected to other parts or nodes of the IT system. The bare metal cloud nodes or controllers of an IT system may also become vulnerable via resources connected to an application network that exposes the system to security threats or otherwise puts the system at risk. According to various exemplary embodiments disclosed herein, an IT system may be configured to improve the security of bare metal cloud nodes or resources from an application network, whether or not the application network interfaces with the Internet or is connected to an external network.

[0006] According to an exemplary embodiment, an IT system includes bare metal cloud nodes or physical resources. When a bare metal cloud node or physical resource is powered on, set up, managed, or used and may be connected to a network having nodes that may be used by other people or customers, in-band management may be omitted from the controller, may be switchable, may be disconnectable, or may be filtered. Also, an application network or networks of multiple applications within the system may be disconnected from, disconnectable from, switchable from, or filtered from the controller via the resource(s) to which the application network is coupled to the controller.

[0007] Physical resources that include virtual machines or hypervisors may also be vulnerable to security issues and may be at risk of being exposed to or accessed by unauthorized users when using a hypervisor to pivot to other hypervisors that are shared resources. An attacker may be able to escape from a virtual machine and may be able to gain network access to a management system and / or a monitoring system via a controller. According to various exemplary embodiments disclosed herein, an IT system may be configured to improve security, and one or more physical resources, including virtual resources on a cloud platform, may be disconnected from, disconnectable from, filterable, or filterable by, or disconnected from a controller via an in-band management connection.

[0008] According to an exemplary embodiment, the physical resources of an IT system may include one or more virtual machines or hypervisors, and the in-band management connection between the controller and the physical resources may be omitted from, disconnected from, disconnectable from, or filterable / filterable by that resource.

[0009] According to an exemplary embodiment, the system may include a controller that provisions and manages services related to each other within the system using the techniques described herein. By way of example, cleanup rules can be created and maintained that manage how to resolve modifications when deleting a service that has interdependencies with other services.

[0010] ]> According to an exemplary embodiment, the system may include a controller that provisions storage to computing resources and / or provisions and connects resources to cloud instances using the techniques described herein.

[0011] Furthermore, according to an exemplary embodiment, the system can assist in an efficient backup operation that includes a backup with multiple interdependent services, using the architecture described herein.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 2F

Figure 2G

Figure 2H

Figure 2I

Figure 2J

Figure 2K

Figure 2L

Figure 2M

Figure 2N

Figure 2O

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 7C

Figure 7D

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 9E

Figure 9F

Figure 9G-1

Figure 9G-2

Figure 9G-3

Figure 9G-4

Figure 9H

Figure 9I

Figure 9J

Figure 9K

Figure 9L

Figure 10

Figure 11A

Figure 11B

Figure 12

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 16A

Figure 16B

Figure 16C

Figure 16D

Figure 16E

Figure 16F

Figure 17A

Figure 17B

Figure 18A

Figure 18B

Figure 19A

Figure 19B

Figure 19C

Figure 19D

Figure 19E

Figure 19F

Figure 19G

Figure 20

Figure 21A

Figure 21B

Figure 21C-1

Figure 21C-2

Figure 21C-3

Figure 21C-4

Figure 21D

Figure 21E

Figure 21F

Figure 21G

Figure 21H

Figure 21I

Figure 21J

Figure 21K

Figure 21L-1

Figure 21L-2

Figure 22A

Figure 22B

Figure 22C

DETAILED DESCRIPTION OF THE INVENTION

[0013] In an effort to provide a technical solution to the needs of the art as described above, the present inventors disclose various embodiments of systems and methods for information technology that provide for the setup, configuration, maintenance, testing, change management, and / or upgrade of automated IT systems. For example, the present inventors disclose a controller configured to automatically manage a computer system based on a plurality of system rules, the system state of a computer system, and a plurality of templates. As another example, the present inventors disclose a controller configured to automatically manage the physical infrastructure of a computer system based on a plurality of system rules, the system state of a computer system, and a plurality of templates. Examples of automatic management executable by the controller may include remote or local access to and modification of settings or other information on a computer capable of operating an application or service, construction of an IT system, change of an IT system, construction of an individual stack within an IT system, creation of a service or application, loading of a service or application, configuration of a service or application, migration of a service or application, change of a service or application, deletion of a service or application, cloning of a stack to another stack on a different network, creation, addition, deletion, setup, configuration, reconfiguration, and / or change of a resource or system component, automatic addition, deletion, and / or restoration of a resource, service, application, IT system, and / or IT stack, configuration of interactions between applications, services, stacks, and / or other IT systems, and / or monitoring of the health of IT system components. In an exemplary embodiment, the controller can be embodied as a physical or virtual computing resource that can be remote or local. Additional examples of controllers that can be employed include, but are not limited to, any one or any combination of a process, virtual machine, container, remote computing resource, application deployed by another controller, and / or service.The controller may be distributed across multiple nodes and / or resources and may be located elsewhere or on a network.

[0014] IT infrastructure is most often built from individual hardware and software components. The hardware components used generally include servers, racks, power equipment, interconnections, display monitors, and other communication equipment. The ways and techniques for selecting and interconnecting these individual components become very complex due to a vast number of optional configurations that function with varying degrees of efficiency, cost-effectiveness, performance, and security. Individual technicians / engineers skilled in connecting these infrastructure components are costly to hire and train. Also, due to the vast number of hardware and software repetition possibilities, the maintenance and updating of that hardware and software become complex. This creates additional challenges when the individuals and / or engineering companies that originally installed the IT infrastructure are unable to perform the updates. Software components such as operating systems are designed either generically to operate across a wide range of hardware or are highly specialized for specific components. Most often, a complex plan or blueprint is drawn up and executed. Changes, growth, scaling, and other issues require updating the complex plan.

[0015] Some IT users purchase cloud computing services from the growing supplier industry, which does not solve the problems and issues of infrastructure setup, but rather transfers them from the IT users to the cloud service providers. Further, large cloud service providers are addressing the issues and problems of infrastructure setup in ways that can reduce flexibility, customizability, scalability, and the rapid adoption of new hardware and software technologies. Also, cloud computing services do not provide immediately available bare metal setups, deployment, and updates, nor do they enable migration to, from, or between bare metal and virtual IT infrastructure components. These and other limitations of cloud computing services can result in many computing, storage, and networking inefficiencies. For example, speed or latency inefficiencies in computing and networking can occur either in the cloud service or in an application or service that uses the cloud service.

[0016] The systems and methods of the exemplary embodiments provide for the deployment, utilization, and management of a novel and unique IT infrastructure. According to the exemplary embodiments, the complexity of resource selection, installation, interconnection, management, and update is rooted in the core controller system and its parameter files, templates, rules, and IT system state. This system includes a set of self-assembly rules and operating rules configured such that components perform self-assembly rather than requiring technicians to assemble, connect, and manage. Further, the systems and methods of the exemplary embodiments enable higher customizability, scalability, and flexibility using self-assembly rules without the need for current typical external specifications. They also enable efficient resource use and reuse.

[0017] A system and method are provided that improve many of the problems and issues of current IT systems, whether wholly or partially physical or virtual. The system and method of the exemplary embodiments provide a structure that enables flexibility, reduces variation and human error, and improves system security.

[0018] While there may be several individual solutions for one or more of the problems of current IT systems, such solutions do not comprehensively address the numerous problems as solved by the exemplary embodiments described herein. Further, such existing solutions may address a particular problem but may exacerbate other problems.

[0019] Among the current challenges addressed are problems related to, but not limited to, setup, configuration, infrastructure deployment, asset tracking, security, application deployment, service deployment, maintenance and compliance documentation, maintenance, scaling, resource allocation, resource management, load balancing, software failures, software and security updates / patch application, testing, IT system recovery, change management, and hardware updates.

[0020] As used herein, an IT system can include, but is not limited to, servers, virtual and physical hosts, databases and database applications, such as, but not limited to, IT services, business computing services, computer applications, customer support applications, web applications, mobile applications, backends, case number management, customer tracking, ticketing, business tools, desktop management tools, accounting, email, documentation, compliance, data storage, backup, and / or network management.

[0021] One of the problems a user may face before setting up an IT system is predicting infrastructure needs. First, or over time as it grows or changes, the user may not know how much storage, computing power, or other requirements will be needed. According to an exemplary embodiment, the IT system and infrastructure enable flexibility in that, when the system requires a change, it can be automatically added, removed, or reallocated within the infrastructure using the self-provisioning infrastructure (both physical and / or virtual) of the exemplary embodiment. Thus, the challenge of predicting future needs presented at system setup is addressed by providing a function to add to the system using global rules, templates, and system states, and tracking changes to such rules, templates, and system states.

[0022] Other issues may relate to correct configuration, configuration consistency, interoperability, and / or interdependence, which may include, for example, future incompatibilities due to changes in configured system elements or their configurations over time. For example, when an IT system is first set up, there may be missing elements or a failure in the configuration of some elements. Also, for example, when repetitions of elements or infrastructure components are set up, there may be a lack of consistency between the repetitions. When changes are made to the system, a review of the configuration may be required. Difficult choices are presented between an optimal configuration and flexibility with respect to future infrastructure changes. According to an exemplary embodiment, when initially deploying a system, the configuration is self-provisioned from a template to infrastructure components using global system rules, so that the configuration is unified, reproducible or predictable, enabling an optimal configuration. Such an initial system deployment can be performed on physical components, but subsequent components may be added or changed, which may or may not be physical. Further, such an initial system deployment can be performed on physical components, but subsequent environments may be cloned from the physical structure, which may or may not be physical. This makes it possible to optimize the system configuration while minimizing changes that cause future problems.

[0023] In the provisioning phase, typically, there are issues of interoperability of bare metal and / or software-defined infrastructure. There may also be issues of interoperability between software and other applications, tools, or infrastructure. These can include, but are not limited to, issues caused by deployed products from different vendors. The inventors disclose an IT system that can provide infrastructure interoperability regardless of whether it is bare metal, virtual, or any combination thereof. Thus, interoperability, i.e., the ability of parts to work together, can be incorporated into the disclosed infrastructure provisioning where the infrastructure is automatically configured and provisioned. For example, different applications may be interdependent and they may exist on separate hosts. To enable such applications to communicate with each other, the controller logic, templates, system state, and system rules described herein include the information and configuration instructions used for the configuration of application interdependencies and the tracking of interdependencies. Thus, the infrastructure features described herein provide a way to manage how each application or service interacts with each other. As an example, to enable an email service to communicate properly with an authentication service and / or to enable a groupware service to communicate properly with the email service. Even further, such management can go down to the infrastructure level, for example, making it possible to track how computing resources are communicating with storage resources. Otherwise, the complexity of the IT system may increase at O(n n )

[0024] As disclosed, the automatic provisioning of resources does not require pre-configuration of the operating system software by virtue of the ability of the controller to provision based on global system rules, templates, and the IT system state / system self-awareness. According to an exemplary embodiment, a user or IT specialist may not need to know whether the addition, assignment, or re-assignment of resources is coordinated in order to ensure interoperability. Additional resources according to an exemplary embodiment may be automatically added to the network.

[0025] To use an application, typically many different resources are required, including computing, storage, and networking. Also required is the interoperability of the resources with system components, which includes knowledge of what is located and operating in place and interoperability with other applications. The application may need to connect to other services to obtain configuration files and ensure that all components are properly coordinated. Therefore, the configuration of an application can be time-consuming and resource-intensive. If there are interoperability issues with other applications, the configuration of the application can have a cascading effect on the rest of the infrastructure. This can lead to outages or violations. The inventors disclose automated application provisioning to address these issues. Thus, as the inventors disclose, an application can be self-provisioned by reading from the IT system state, global system rules, and templates and configuring intelligently using knowledge of the current state of the system. Further, according to an exemplary embodiment, a pre-deployment test of the configuration can be performed using the change management function described herein.

[0026] Another problem addressed by the exemplary embodiments relates to issues that can arise with intermediate configurations when it is desirable to switch to a different vendor or other tool. According to one aspect of the exemplary embodiments, template conversion is provided between the rules and templates of the controller and the application templates of a particular vendor. This enables the system to automatically change the vendor of the software or other tool.

[0027] Many security issues arise from configuration mistakes, failed patch applications, and the inability to test patch applications before deployment. Often, security issues can occur during the setup configuration phase. For example, due to configuration mistakes, a highly confidential application may be left exposed to the Internet, or email forgery may be allowed from an email server. The inventors disclose a system setup that, by being automatically configured, protects against attackers, avoids unnecessary exposure to attackers, and provides more knowledge about the system to security engineers and application security architects. Automation reduces security vulnerabilities due to human or configuration errors. Also, the disclosed infrastructure can provide introspection between services, enable rule-based access, and limit communication between services to only what is actually necessary. The inventors disclose a system and method having the ability to safely test patches before deployment, as will be described, for example, with respect to change management.

[0028] Documentation is often an area of troubled IT management. During setup and configuration, the main goal can typically be to get components to work together. Usually, this involves a process of troubleshooting and trial-and-error, and it can be difficult to know exactly what made the system actually function. The exact commands executed are usually documented, but the troubleshooting or trial-and-error processes that may have led to a functioning system are often not well-documented or not documented at all. Problems or deficiencies in documentation can lead to issues with audit trails and auditing. The resulting documentation problems can cause issues when demonstrating compliance. When building a system or its components, compliance issues are often not well-known. Applicable compliance determinations may not be known until after the IT system has been set up and configured. Thus, documentation is essential for auditing and compliance. The inventors disclose a system that includes a global system rules database, templates, and an IT system state database that provide automatically documented setup and configuration. Every configuration that occurs is recorded in the database. According to an exemplary embodiment, the automatically documented configuration provides an audit trail and can be used to demonstrate compliance. In inventory management, information that is automatically documented and tracked can be used.

[0029] Another problem arising from the setup, configuration, and operation of IT systems relates to the inventory management of hardware and software. For example, it is usually important to know how many servers there are, whether they are up and still functioning, what their functions are, which rack each server is in, which power supply is connected to which server, which network cards and which network ports each server is using, in which IT system a component is operating, and many other important matters. In addition to the inventory information, the passwords and other confidential information used for inventory management need to be effectively managed. Particularly in large-scale IT systems, data centers, or data centers where equipment is frequently changed, the collection and maintenance of this information is a time-consuming task, which is often managed manually or using various software tools. Compliance protection with secure passwords is a major risk factor that can be an important issue when ensuring a secure computing environment. The inventors disclose an IT system in which the collection and maintenance of the inventory and operating status of all servers and other components are automatically updated, stored, and protected as part of the controller logic of the IT system status, global system rules, templates, and controllers.

[0030] In addition to problems related to the setup and configuration of IT systems, the inventors disclose an IT system that can also address problems and issues that occur during the maintenance of IT systems. For example, among other things, as data centers with hardware failures, such as power failures, memory failures, network failures, network card failures, and / or CPU failures, continue to function, many problems arise. When migrating hosts during a hardware failure, further failures occur. Therefore, the inventors disclose the migration of dynamic resources, for example, migrating resources from one resource provider to another when a host goes down. In such a situation, according to an exemplary embodiment, the IT system can migrate to other servers, nodes, or resources, or other IT systems. The controller can report the state of the system. The replication of data is on other hosts with a known automatically set up configuration. When a hardware failure is detected, any resources that the hardware may have provided can be automatically migrated after the failure is automatically detected.

[0031] An important issue regarding many IT systems is scalability. Growing businesses or other organizations typically add or reconfigure their IT systems as they grow and their needs change. For example, problems occur when more resources are needed for an existing IT system, such as additional hard drive capacity, storage capacity, CPU processing, more network infrastructure, more endpoints, more clients, and / or more security. Problems also occur in configuration, setup, and deployment when different services and applications, or infrastructure changes, are needed. According to an exemplary embodiment, a data center can be automatically scaled. Nodes or resources can be added to or removed from a resource pool dynamically and automatically. Resources added to and removed from the resource pool can be automatically allocated or reallocated. Services can be provisioned and quickly moved to new hosts. A controller can detect more resources and dynamically add them to the resource pool and know where resources should be allocated / reallocated. A system according to an exemplary embodiment can scale from a single-node IT system to a scaled system that requires multiple physical and / or virtual nodes or resources across multiple data centers or IT systems.

[0032] The inventors disclose a system that enables flexible resource allocation and management. This system may be within a resource pool and includes computational resources, storage resources, and networking resources that can be dynamically allocated. The controller can recognize new nodes or hosts on the network and then configure them to become part of the resource pool. For example, when a new server is plugged in, the controller can configure it as part of the resource pool, add it to the resources, and dynamically start using it. Nodes or resources can be detected by the controller and added to different pools. Resource requests can be made, for example, via an API request to the controller. The controller can then deploy or allocate the required resources from the pool according to rules. This enables the controller and / or an application via the controller to load balance and dynamically distribute resources based on the needs of the requests.

[0033] Examples of load balancing include, but are not limited to, deploying new resources in the event of a hardware or software failure, deploying one or more instances of the same application in response to an increase in user load, and deploying one or more instances of the same application in response to an imbalance in storage, computational, or networking requirements.

[0034] Problems related to making changes to an operating IT system or environment can pose significant, and sometimes catastrophic, problems for users or entities that rely on these systems to consistently come up and operate. These outages not only represent potential losses in the use of the system, but also economic losses due to data loss, the time, personnel, and significant resources of money required to resolve the problem. The problem can be exacerbated by difficulties in reconstructing the system if there are errors in the configuration documentation or a lack of understanding of the system. Due to this problem, many IT system users are reluctant to apply patches to IT resources to eliminate known security risks. In this way, they remain more vulnerable to security breaches.

[0035] Many problems that occur in the maintenance of IT systems are related to software failures due to change management or control that may be required for configuration. Situations in which such failures can occur include, but are not limited to, upgrades to new software versions, migrations to different software, changes to password or authentication management, switching between services or different providers of a service.

[0036] Infrastructure that is manually configured and maintained is usually difficult to recreate. Recreating the infrastructure can be important for several reasons including, but not limited to, rolling back problematic changes, power outages, or other disaster recovery. It is difficult to diagnose problems with systems that are manually configured. It is difficult to recreate infrastructure that is manually configured and maintained. Also, system administrators can easily make mistakes, such as an incorrect command that is known to have brought down a computer system.

[0037] Making changes to an operating IT system or environment can pose significant, and sometimes catastrophic, problems for users or entities that depend on these systems to consistently start up and operate. These outages not only represent potential losses in the use of the system, but such outages can also cause economic losses due to data loss, as well as significant resources of time, personnel, and money required to resolve the problem. The problem can be exacerbated by errors in the configuration documentation or lack of understanding of the system, making it difficult to reconstruct the system. Also, in many cases, it is very difficult to restore the system to its previous state after a major or large change.

[0038] Furthermore, technical problems that can occur when changes are made to an operating environment can have a cascading effect. These cascading effects can make it difficult, and in some cases impossible, to return to the state prior to the change. Thus, even if it is necessary to reverse a change due to problems with the change that was implemented, the state of the system has already been changed. In recent years, it has been stated that errors in infrastructure and system management, as well as incomplete changes to production environments that cannot be undone, are problems. Additionally, it is known that testing changes to a system before deployment to an operating environment is problematic.

[0039] Accordingly, the inventors disclose some exemplary embodiments of systems and methods configured to reverse changes to an operating system to its state prior to the change. Furthermore, the inventors disclose systems and methods configured to enable a significant restoration of the state of a system or environment that has received an in - operation change, which can prevent or ameliorate one or more of the above - mentioned problems.

[0040] According to a variation of the exemplary embodiment, the IT system has complete knowledge of the system based on global system rules, templates, and the IT system state. The infrastructure can be cloned using the complete knowledge of the system. The system or system environment can be cloned as a software-defined infrastructure or environment. A system environment including a volatile database in use, called the production environment, can be written to a non-volatile read-only database and used as a development environment in the development and test processes. Desired changes can be applied to the development environment and tested there. A user or controller logic can create a new version by making changes to the global rules. The versions of the rules can be tracked. Next, according to another aspect of the exemplary embodiment, the newly developed environment can be automatically implemented. The previous production environment may also be retained or fully functional, so that corrections to the previous state of the production environment are possible without data loss. Next, the development environment can be booted using the new specifications, rules, and templates, the database or system can be synchronized with the production database, and can be switched to a writable database. And the original production database can be switched to a read-only database, and the system can be restored to it if recovery is needed.

[0041] Regarding software upgrades or patch applications, when a service requiring an upgrade or patch is detected, a new host can be deployed. When a failure occurs due to an upgrade or patch, a new service can be deployed if it is possible to roll back the changes as described above.

[0042] Hardware upgrades are important in many situations, especially where the latest hardware is essential. An example of this type of situation occurs in the high-frequency trading industry, where IT systems with millisecond-level speed advantages can enable users to achieve excellent trading results and profits. In particular, problems arise when ensuring interoperability with the current infrastructure, such as recognizing how new hardware communicates using protocols and interoperates with the existing infrastructure. In addition to ensuring component interoperability, components need to be integrated with the existing setup.

[0043] Referring to FIG. 1, an IT system 100 of an exemplary embodiment is shown. System 100 can be one or more types of IT systems including, but not limited to, those described herein.

[0044] A user interface (UI) 110 coupled to a controller 200 via an application program interface (API) application 120 may or may not be present on a stand - alone physical or virtual server. The controller 200 may be deployed on one or more processors and one or more memories to perform any of the control operations described herein. Instructions executed by the processor(s) to perform such control operations may reside on a non - transitory computer - readable storage medium such as processor memory. The API 120 may include one or more API applications, which may be redundant and / or may operate in parallel. The API application 120 receives requests to configure system resources, analyzes the requests, and passes them to the controller 200. The API application 120 receives one or more responses from the controller, analyzes the response(s), and passes them to the UI (or application) 110. Alternatively, or in addition, an application or service may communicate with the API application 120. The controller 200 is coupled to computing resource(s) 300, storage resource(s) 400, and networking resource(s) 500. The resources 300, 400, 500 may or may not be present on a single node. One or more of the resources 300, 400, 500 may be virtual. The resources 300, 400, 500 may or may not be present on multiple nodes in various combinations. A physical device may include one or more or each of resource types including but not limited to computing resources 300, storage resources 400, and networking resources 500. The resources 300, 400, 500 may also include a pool of resources regardless of whether they are in different physical locations and regardless of whether they are virtual. Also, bare - metal computing resources may be used to enable the use of virtual or container computing resources.

[0045] In addition to the known definitions of nodes, the nodes used in this specification can be any system, device, or resource that executes functions on a stand-alone device or a network-connected device and is connected to a network (s) or other functional unit. Nodes can include, but are not limited to, for example, servers, services / applications / multiple services on a physical or virtual host, virtual servers, and / or multiple or single services operating on a multi-tenant server or within a container.

[0046] The controller 200 may include one or more physical or virtual controller servers, which may also be redundant and / or operate in parallel. The controller may operate on a physical or virtual host that functions as a computing host. As an example, the controller may be composed of controllers operating on a host that is also useful for other purposes, for example, to access highly confidential resources. The controller may receive requests from the API application 120, analyze the requests, optimize the task allocation to other resources, instruct other resources, monitor and receive information from the resources, maintain the state and change history of the system, and communicate with other controllers within the IT system. Also, the controller may include the API application 120.

[0047] The computing resources defined herein may include a physical or virtual single computing node, or a resource pool including one or more computing nodes. The computing resources or computing nodes may include one or more physical or virtual machines or container hosts that can host one or more services or run one or more applications. The computing resources may be on hardware designed for multiple purposes including, but not limited to, computing, storage, caching, networking, and specialized computing, including GPUs, ASICs, coprocessors, CPUs, FPGAs, and other specialized computing methods. Such devices may be added using a PCI Express switch or similar device and may be added dynamically in such a manner. The computing resources or computing nodes may include one or more hypervisors or container hosts that can run or be virtual computing resources including multiple different virtual machines that can run a service or application. The computing resources may be focused on providing computing capabilities, but may also include data storage capabilities and / or networking capabilities.

[0048] The storage resources defined herein may include a storage node or a pool of storage resources. The storage resources may include any data storage medium, such as fast, slow, hybrid, cache, and / or RAM. The storage resources may include one or more types of networks, machines, devices, nodes, or any combination thereof, which may or may not be directly connected to other storage resources. According to aspects of an exemplary embodiment, the storage resources may be bare metal or virtual or a combination thereof. The storage resources may be focused on providing storage capabilities, but may also include computing capabilities and / or networking capabilities.

[0049] The networking resource(s) 500 may include a single networking resource, multiple networking resources, or a pool of networking resources. The networking resource(s) may include a physical device or virtual device(s), tool(s), switch, router, or other interconnection between system resources, or an application for managing networking. Such system resources may be physical or virtual and may include computing resources, storage resources, or other networking resources. The networking resource may provide a connection between an external network and an application network and may host core network services including, but not limited to, Domain Name System (DNS or dns), Dynamic Host Configuration Protocol (DHCP), subnet management, Layer 3 routing, Network Address Translation (NAT), and other services. Some of these services may be deployed on computing resources, storage resources, or networking resources on a physical or virtual machine. The networking resource may utilize one or more fabrics or protocols including, but not limited to, Infiniband, Ethernet, Remote Direct Memory Access (DMA) over Converged Ethernet (RoCE), Fibre Channel, and / or Omnipath and may include an interconnection between multiple fabrics. The networking resource may or may not be Software Defined Networking (SDN) compliant. The controller 200 may be able to configure the topology of the IT system by directly changing the networking resource 300 using SDN, Virtual Local Area Network (VLAN), etc. The networking resource may focus on providing networking functions but may also have computing and / or storage functions.

[0050] As used herein, an application network means a networking resource or any combination thereof for connecting or coupling applications, resources, services, and / or other networks, or for connecting users and / or clients to applications, resources, and / or services. An application network can include a network used by a server to communicate with other application servers (physical or virtual) and with clients. An application network can communicate with machines or networks external to system 100. For example, an application network can connect a web front end to a database. A user can connect to a web application via the Internet or via another network that may or may not be managed by a controller.

[0051] According to an exemplary embodiment, computing resources 300, storage resources 400, and networking resources 500 can each be automatically added, removed, set up, allocated, reallocated, configured, reconfigured, and / or deployed by controller 200. According to an exemplary embodiment, additional resources can be added to a resource pool.

[0052] Although a user interface 110 such as a Web UI or other user interface through which user 105 can access and interact with the system is illustrated, alternatively or additionally, the application can communicate or interact with controller 200 via API application(s) 120 or in another way. For example, user 105 or the application can send requests including, but not limited to, construction of an IT system, construction of individual stacks within the IT system, creation of a service or application, migration of a service or application, change of a service or application, deletion of a service or application, cloning of a stack to another stack on a different network, creation, addition, deletion, setup or configuration, reconfiguration of a resource or system component.

[0053] The system 100 of FIG. 1 can include a server having connections or other communication interfaces to various elements, components or resources that can be physical or virtual or any combination thereof. According to a variant, the system 100 shown in FIG. 1 can include a bare metal server having connections.

[0054] As will be described in more detail herein, controller 200 can be configured to add, allocate, manage and update available resources by powering on a resource or component and automatically setting up, configuring and / or controlling the boot-up of the resource. The power-on process can start with powering on the controller so that the order of devices to be booted is consistent and does not depend on the user powering on the devices. This process can also include detection of the powered-on resources.

[0055] Referring to FIGS. 2A - 10, controller 200, controller logic 205, global system rules database 210, IT system state 220, and template 230 are shown.

[0056] System 100 includes global system rules 210. The global system rules 210 can declare rules for setting up, configuring, booting, allocating, and managing resources that can include, among other things, computing, storage, and networking. The global system rules 210 include the minimum requirements for the system 100 to be in a correct or desired state. These requirements can include IT tasks whose completion is expected and an updatable list of the expected hardware necessary to build the desired system as expected. The updatable list of the expected hardware enables the controller to confirm that the necessary resources are available (e.g., before the start of a rule or before the use of a template). The global rules can include a list of the actions required for various tasks and the corresponding instructions related to the ordering of the actions and tasks. For example, the rules can specify the order in which to power on components, the order in which to boot resources, applications, and services, dependencies, when to start various tasks, e.g., when to start the loading, configuration, start, reload, or update of hardware. The rules 210 can also, for example, include a list of resource allocations required for applications and services, a list of templates that can be used, a list of applications and configuration methods to be loaded, a list of services and configuration methods to be loaded, a list of application networks and which applications are compatible with which networks, a list of configuration variables specific to various applications and user-specific application variables, the expected state that enables the controller to check the system state and confirm that the state is as expected and the result of each instruction is as expected, and / or a list of changes to the rules (e.g., snapshots) that can enable tracking of rule changes and the ability to test or revert to different rules in different situations, including one or more of a version list. The controller 200 can be configured to apply the global system rules 210 to the IT system 100 on physical resources.The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on the virtual resources. The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on a combination of physical resources and virtual resources.

[0057] FIG. 2M shows an exemplary set of system rules 210 that can take the form of global system rules. The exemplary set of system rules 210 shown in FIG. 2M can be loaded into the controller 200 or derived by querying the system state (see 210.1). In the example of FIG. 2M, the system rules 210 include a set of instructions that can take the form of a configuration routine 210.2, and also include data 210.3 for creating and / or recreating an IT system or environment. The configuration rules within the system rules 210 may recognize a way to find the template 230 via the requested template list 210.7 (the template 230 may be present in a file system, disk, storage resource, or placed within the system rules). The controller logic 205 may also search for the template 230 before processing the template 230 and enable the system rules 210 after confirming the existence of the template 230. The system rules 210 may include a subset 210.15 of the system rules, and these subsets 210.15 may be executed as part of the configuration routine 210.2.

[0058] Also, subsystem rule 210.15 can be used, for example, as a tool for building a system of integrated IT applications (in which case, it is processed by system rule execution routine 210.16 and the system state and current configuration rules are updated to reflect the addition of 210.15). Also, subsystem rule 210.15 can be placed elsewhere and loaded into system state 220 by user interaction. For example, subsystem rules 210.15 can also be had as playbooks and can be made available and operated (since global system rule 210 is then updated, the playbooks can be replayed if the system is desired to be cloned).

[0059] Configuration routine 210.2 may be a set of instructions used to build a system. Configuration routine 210.2 may also include subsystem rule 210.15 or system state pointer 210.8 if the implementer desires. When executing configuration routine 210.2, controller logic 205 processes a series of templates in a specific order (210.9), optionally enabling parallel deployment, but appropriate dependency processing (210.12) can be maintained. Configuration routine 210.2 may optionally call API call 210.10 which can set configuration parameter 210.5 of an application that may be configured by processing templates according to 210.9. Also, requested service 210.11 is a service that needs to be up and running when the system makes API call(s) 210.10.

[0060] Routine 210.2 may include, but is not limited to, copying data, transferring the database to computing resources, pairing computing resources with storage resources, and / or updating the system state 220 based on the location of volatile data 210.6, and may include procedures, programs, or methods for data loading (210.13) related to volatile data 210.6. By holding a pointer to the volatile data (see 210.4) together with the data 210.3, volatile data that can be stored elsewhere can be found. The data loading routine 210.13 can also be used to load the configuration parameters 210.5 when they are located in a non-standard data store (e.g., included in a database).

[0061] System rules 210 can also include a resource list 210.18 that can indicate which components are assigned to which resources and enable the controller logic 205 to determine whether appropriate resources and / or hardware are available. System rules 210 can also include an alternative hardware and / or resource list 210.19 for alternative deployments (e.g., for a development environment where a software engineer may want to perform a demonstration test but may not want to allocate the entire data center). The system rules can also include a data backup and / or standby routine 210.17 that provides instructions on how to back up the system and how to use standby for redundancy. Examples of data backup systems and / or backup routines that implement the backup rules include, but are not limited to, those described herein with reference to FIGS. 21A - 21J.

[0062] After any action is taken, the system state 220 may be updated and a query (which may include writes) may be saved as a system state query 210.14.

[0063] Figure 2N shows an exemplary process flow in which controller logic 205 processes the system rule 210 (or subsystem rule 210.15) of Figure 2M. At step 210.20, the controller logic 205 checks and confirms that the appropriate resources are available (see 210.18 in Figure 2M). If not, at step 210.21, alternative configurations may be checked. A third option may include that the user is prompted to select an alternative configuration that can be supported by the template 230 referred to in list 210.7 of Figure 2M.

[0064] At step 210.22, the controller logic may then confirm that the computing resource (or any of the appropriate resources) accesses the volatile data. This may involve connecting to the storage resource or adding the storage resource to the system state 220. At step 210.23, next, the configuration routines are processed, and each time a routine is processed, the system state 220 is updated (step 210.24). The system state 220 may also be queried to check whether a particular step has been completed before processing (step 210.25).

[0065] The configuration routine processing steps shown in Figure 210.23 may include any (or combinations thereof) of the steps of 210.26. It may also include other steps. For example, the processing in 210.26 may include template processing (210.27), loading configuration data (210.28), loading static data (210.29), loading dynamic volatile data (210.30), and / or coupling services, applications, subsystems, and / or environments (210.31). Such steps within 210.26 may be repeated in a loop or executed in parallel since some system components are independent and others are dependent. Controller logic, service dependencies, and / or system rules may indicate which services can depend on each other, and services may be coupled to further build an IT system from the system rules.

[0066] Global system rule 210 may include storage expansion rules. The storage expansion rules provide, for example, a set of rules that automatically add storage resources to existing storage resources in the system. Further, a trigger point may be provided to ascertain when an application running on a computing resource(s) requests storage expansion (or the controller 200 may be able to ascertain when the storage of a computing resource or application should be expanded). The controller 200 may allocate and manage new storage resources and may merge or integrate the storage resources with existing storage resources for a particular running resource. Such a particular running resource may be, but is not limited to, a computing resource in the system, an application running a computing resource in the system, a virtual machine, a container, or a physical or virtual computing host, or a combination thereof. A running resource may inform the controller 200, for example, via a storage space query, that it is about to exhaust storage space. An in-band management connection 270, a SAN connection 280, or any networking or attachment to the controller 200 may be used in such a query. An out-of-band management connection 260 may also be used. Storage expansion rules (or a subset of those storage expansion rules) may also be used for non-running resources.

[0067] The storage expansion rules dictate how to find, connect, and set up new storage resources within the system. The controller registers the new storage resources with the system state 220 and notifies the running resources of where the storage resources exist and how to connect to them. The running sources use such registration information to connect to the storage resources. The controller 200 may merge new storage resources with existing storage resources or may add new storage resources to a volume group.

[0068] Figure 2B shows an exemplary flow of the operation of an exemplary set of storage expansion rules. At step 210.41, the running resource determines that there is low storage based on a trigger point or otherwise. At step 210.42, the running resource connects to the controller 200 via an in-band management connection 270, a SAN connection 280, or another type of connection visible to the operating system. Through this connection, the running resource can notify the controller 200 that there is low storage. At step 210.43, the controller configures the storage resource to increase the storage capacity for the running resource. At step 210.44, the controller provides the running resource with information regarding the location of the newly configured storage resource. At step 210.45, the running resource connects to the newly configured storage resource. At step 210.46, the controller adds a mapping of the location of the new storage resource to the system state 220. Next, the controller can add the new storage resource to the volume group assigned to the running resource (step 210.47), or the controller can add an assignment of the new storage resource to the running resource to the system state 220 (step 210.48).

[0069] Figure 2C shows an alternative for performing steps 210.41 and 210.42 in Figure 2B. At step 210.50, the controller sends key commands through the out-of-band management connection 260 to browse a monitor or console for storage state updates for the running resource. For example, the monitor may be an ipmi console, and the screen can be viewed through the ipmi console via the out-of-band connection 260. As an example, the out-of-band connection 260 can be connected to USB as a keyboard / mouse and can be connected to a VGA monitor port. At step 210.51, the running resource displays information on the screen. At step 210.52, the controller then reads the information presented on the monitor or console via the out-of-band management connection 260 and screen scraping or similar operations, and this read information may indicate a low storage state based on a trigger point. The process flow may then continue with step 210.43 of Figure 2B.

[0070] Figure 2D shows another alternative for performing steps 210.41 and 210.42 in Figure 2B. At step 210.55, the running resource automatically displays information on the monitor or console for the controller to read. At step 210.56, the controller automatically, periodically, or continuously reads the monitor or console to check the running resource. In response to this reading, the controller confirms that the running resource has low storage (step 210.57). The process flow may then continue with step 210.43 of Figure 2B.

[0071] The controller 200 also includes a library of templates 230 that may include bare metal and / or service templates. Those templates may include, but are not limited to, third-party applications that may be configurable by email, file storage, voice over IP, software accounting, software XMPP, wiki, version control, account authentication management, and user interface. The template 230 can have an association with a resource, application, or service, and can function as a recipe that defines how such a resource, application, or service is integrated into the system.

[0072] Thus, a template can include a set of established information used to create, configure, and / or deploy a resource, or an application or service loaded on the resource. Such information may include, but is not limited to, the kernel, initrd file, file system, or file system image, files, configuration files, configuration file templates, information used to determine an appropriate setup for different hardware and / or computing backends, and / or other available options for configuring resources for running application and operating system images that enable and / or facilitate the creation, boot, or execution of the application.

[0073] The template can include information that can be used to deploy an application on multiple supported hardware types / and or computing backends, including, but not limited to, multiple physical server types or components, multiple hypervisors running on multiple hardware types, and container hosts that can be hosted on multiple hardware types.

[0074] A template can derive a boot image of an application or service to be executed on computing resources. Using the template and the image derived from the template, an application can be created, an application or service can be deployed, and / or resources can be prepared for various system functions, which enables and / or facilitates the creation of the application. A template may have variable parameters in a file, file system, and / or operating system image that can be overwritten by configuration options from either default settings or settings provided by a controller. A template may have configuration scripts used to configure an application or other resources, and the template may utilize configuration variables, configuration rules, and / or default rules or variables, and those scripts, variables, and / or rules may include parameters specific to a particular hardware or other resources, such as specific rules, scripts, or variables for a hypervisor (when virtual), available memory. A template may have a file in the form of a binary resource, a binary resource, or compilable source code that yields parameters specific to hardware or other resources, a specific set of binary resources, or source code with compilation instructions for specific hardware or other resource-specific parameters, such as a hypervisor (when virtual), available memory. A template may include a set of information independent of what is being executed on the resources.

[0075] A template may include a base image. The base image may include a base operating system file system. The base operating system may be read-only. The base image may also include basic tools of an operating system independent of the one in execution. The base image may include a base directory and operating system tools. A template may include a kernel. The kernel or kernels may include an initrd kernel, or multiple kernels configured for different hardware types and resource types. An image may be derived from a template and loaded or deployed to one or more resources. The loaded image may also include a boot file such as the kernel of the corresponding template or the initrd kernel.

[0076] An image may include template file system information that can be loaded to a resource based on a template. The template file system may constitute an application or a service. The template file system may include a shared file system common to all resources or similar resources, for example, to save storage space where the file system is stored or to facilitate the use of read-only files. The template file system or the image may include a set of files common to the services to be deployed. The template file system may be pre-loaded on a controller or may be downloaded. The template file system may be updated. Since the template file system may not need to be rebuilt, it may enable relatively fast deployment. Sharing the file system with other resources or applications may enable storage reduction as files are not unnecessarily duplicated. This may also enable easier recovery from failures as only files different from the template file system need to be restored.

[0077] The template boot file may include the kernel and / or a similar file system used to assist the initrd or the boot process. The boot file may boot the operating system and may set up the template file system. The initrd may include a small-scale temporary file system having instructions on how to set up the template so that it can be booted.

[0078] The template may further include BIOS settings. The template BIOS settings may be used to set optional settings for running an application on the physical host. When used, the out-of-band management 260 may be used to boot a resource or an application as described herein with respect to FIGS. 1-12. The physical host may boot a resource or an application using the out-of-band management network 260 or a CDROM. The controller 200 may set the application-specific BIOS settings defined in such a template. The controller 200 may use the out-of-band management system to make direct BIOS changes through an API specific to a particular resource. The settings may be verified through the console and image recognition. Thus, the controller 200 may use the console function and may make BIOS changes using a virtual keyboard and mouse. The controller may also use the UEFI shell, may type directly into the console, may verify the successful result, may type the commands accurately, and may use image recognition to ensure the success of the setting changes. If there is a bootable operating system available for BIOS changes or updates to a particular BIOS version, the controller 200 may remotely load a disk image or ISO boot on which the operating system runs an application that updates the BIOS and enables configuration changes in a reliable manner.

[0079] The template may further include a list of supported resources specific to the template or a list of resources required to execute a particular application or service.

[0080] The template image, or a portion of the image or template, may be stored in the controller 200, or the controller 200 may move it to, or copy it to, the storage resource 410.

[0081] Figure 2E shows an exemplary template 230. The template includes all the information necessary to create an application or service. Template 230 may also include information about different hardware types that provide similar or identical functionality, alternative data, files, binaries. For example, there may be a file system blob 232 for / usr / bin and / bin, with binaries 234 compiled for different architectures. Template 230 may also include a daemon 233 or a script 231. The daemon 233 is a binary or script that can be executed when the host is powered on and ready, i.e., at boot time. In some cases, the daemon 233 may be accessible by the controller and may run an API that allows the controller to change the host's settings (and the controller may subsequently update the active system rules). The daemon may be powered off and restarted through the out-of-band management 260 or in-band management 270 described above and below. Those daemons may also run a general API to provide services that depend on new services (e.g., a general web server API that communicates with an API controlling nginx or apache). The script 231 may be an installation script that can be executed while or after the image is booted, or after the daemon is started, or after the service is enabled.

[0082] Template 230 may also include a kernel 235 and a pre-boot file system 236. Template 230 may also include multiple kernels 235 and one or more pre-boot file systems (such as Linux's initrd or initramfs, or a BSD read-only RAM disk) for different hardware and different configurations. As will be described below, initrd can also be booted into initramfs 236, which can optionally connect to storage resources through a SAN connection 280, to mount a file system blob 232 presented as an overlay and used to mount a root file system on remote storage.

[0083] The file system blob 232 is a file system image that can be split into separate blobs. The blobs may be mutually changeable based on configuration options, hardware types, and other differences in the setup. The host booted from template 230 may be booted from a union file system (such as overlayfs) that includes multiple blobs or images created from one or more file system blobs.

[0084] Template 230 may also include or be linked to additional information 237 such as volatile data 238 and / or configuration parameters 239. For example, the volatile data 238 may be included in or external to the template 230. The volatile data 238 may be in the form of a file system blob 232, or other data store formats including a database, flat file, files stored in a directory, a tarball of files, git, or other version control repositories, but is not limited thereto. Further, the configuration parameters 239 may be included inside or outside the template 230 and are optionally included in system rules and applied to the template 230.

[0085] System 100 further includes an IT system state 220 that tracks, maintains, changes, and updates the state of System 100, including but not limited to resources. The system state 220 may track available resources, which notifies the controller logic of whether there are resources available for rule implementation and templates and which resources are available. The system state may track used resources, which enables the controller logic 205 to inspect and utilize efficiency regardless of whether there is a need to switch for upgrades or other reasons such as efficiency improvement or priority. The system state may track which applications are running. The controller logic 205 may compare the expected applications to be executed with the actual applications running, according to the system state and whether revisions are needed. The system state 220 may also track where the applications are running. The controller logic 205 may use this information for purposes of evaluating efficiency, change management, updates, troubleshooting, or audit trails. The system state may track networking information, e.g., which networks are on or currently running, or configuration values and history. The system state 220 may also track the history of changes. The system state 220 may also track which templates are being used in which deployments based on global system rules that define which templates are used. The history may be used for audits, warnings, change management, build reports, tracked versions correlated with hardware and applications and configurations, or configuration variables. The system state 220 may maintain a history of configurations for purposes of audits, compliance tests, or troubleshooting.

[0086] The controller has logic 205 for managing all the information included in the system state, templates, and global system rules. The controller logic 205, the global system rules database 210, the IT system state 220, and the templates 230 are managed by the controller 200, and may or may not exist within the controller 200. The controller logic or application 205, the global system rules database 210, the IT system state 220, and the templates 230 may be physical or virtual, may be distributed services, distributed databases, and / or files, or may not be. The API application 120 may be included together with the controller logic / controller application 205.

[0087] The controller 200 may run as a stand-alone machine and / or may include one or more controllers. The controller 200 may include a controller service or application and may run within another machine. The controller machine may first start the controller service so as to ensure that the boot of the entire stack or a group of stacks is ordered and / or made consistent.

[0088] The controller 200 may control one or more stacks having computing resources, storage resources, and networking resources. Each stack may or may not be controlled by a different subset of the rules within the global system rules 210. For example, there may be pre-created, creation, development, inspection stacks, parallel, backup, and / or other stacks having different functions within the system.

[0089] The controller logic 205 may be configured to read and interpret global system rules so as to achieve a desired IT system state. The controller logic 205 may be configured to use templates according to global rules to build system components such as applications or services and to allocate, add, or delete resources so as to achieve a desired IT system state. The controller logic 205 may read global system rules, develop a task list to reach the correct state, and issue instructions that satisfy the rules based on available operations. The controller logic 205 may include logic for executing operations, for example, logic for starting the system, adding, deleting, or reconfiguring resources, and identifying what is available to do so. The controller logic may check the system state at startup and at periodic intervals to confirm whether the hardware is available, and if available, may execute tasks. If the required hardware is not available, the controller logic 205 may present alternative options using the available hardware from the global system rules 210, templates 220, and system state 230, and may modify the global rules and / or system state 220 accordingly.

[0090] The controller logic 205 may recognize which variables are needed, what the user needs to input to continue, or what the user needs in the system to function. The controller logic may use a list of templates from global system rules and compare them with the templates required in the system state to confirm that the required templates are available. The controller logic 205 may identify whether the resources on the list of supported resources specific to the template are available from the system state database. The controller logic may allocate resources, update the state, and proceed to the next set of tasks to implement the global rules. The controller logic 205 may start / run an application on the allocated resources as specified in the global rules. The rules may specify how to construct an application from a template. The controller logic 205 may obtain the template(s) and construct an application from variables. The template may notify the controller logic 205 of which kernel, boot file, file system, and supported hardware resources are required. Next, the controller logic 205 may add information regarding application deployment to the system state database. After each instruction, the controller logic 205 may check the system state database against the predicted state of the global rules to verify whether the predicted actions have been completed accurately.

[0091] The controller logic 205 may use a version that complies with the version rules. The system state 220 may have a database that correlates which rule versions are being used in different deployments.

[0092] The controller logic 205 may include efficient logic and an efficient order for rule optimization. The controller logic 205 may be configured to optimize resources. Information on the system state, rules, and ten related to an application that is running or whose execution is predicted may be used by the controller logic to implement efficiency or priority for the resources. The controller logic 205 may use the information in the "used resources" in the system state 220 to determine the efficiency or necessity of switching resources for upgrade, reuse, or other purposes.

[0093] The controller may check the running applications according to the system state 220 and compare them with the applications for which the execution of global rules is predicted. If the application is not running, the application may be started. If the application should not be running, it may be stopped and the resources may be reallocated appropriately if applicable. The controller logic 205 may include a database of resource (computing, storage, networking) specifications. The controller logic may include logic for recognizing the resource types available to the system that can be used. This may be executed using the out-of-band management network 260. The controller logic 205 may be configured to recognize new hardware using the out-of-band management 260. The controller logic 205 may also retrieve information on the change history, rules used, and versions from the system state 220 for the purposes of auditing, report construction, and change management.

[0094] Figure 2F shows an exemplary process flow of controller logic 205 for processing template 230, booting, powering on, and / or enabling a resource that may be referred to as a host for this exemplary purpose, and deriving an image. This process may also include configuring storage resources and coupling storage hosts and compute hosts and / or resources. Controller logic 205 ascertains the hardware resources available in system 100, and system rules 210 may indicate which hardware resources are available. At step 205.1, controller logic 205 parses template 230, and template 230 may include an instruction file that is executed to cause controller logic to collect files external to template 230 shown in Figure 2E. The instruction file may be in json format. At step 205.2, controller logic collects a list of required file buckets. Also, at step 205.3, controller logic 205 collects into the buckets the necessary hardware-specific files that are referenced by the hardware and optionally by a hypervisor (or container host system, multi-tenancy type). The reference to the hypervisor (or container host system, or multi-tenancy type) may be required if the hardware is to execute on a virtual machine.

[0095] If there is a hardware-specific file, the controller logic collects the hardware-specific file in step 205.4. In some cases, the file system image may include the kernel and initramfs along with a directory containing a kernel module (or a kernel module that will ultimately be placed in a directory). The controller logic 205 then selects an appropriate base image with compatibility in step 205.5. The base image includes operating system files that may not be specific to the image derived from the application or template 230. Compatibility in this context means that the base image includes the files necessary to change the template for the operating application. The base image may be managed outside of the template as a space-saving mechanism (and, in many cases, the base image may be the same for several applications or services). Further, in step 205.6, the controller logic 205 selects a bucket (s) having the executable file, source code, and hardware-specific configuration file. The template 230 may refer to other files including, but not limited to, configuration files, configuration file templates (configuration files that may contain placeholders or variables satisfied by variables in the system rules 210 known in the template 230 so that the controller 200 can change the configuration template to a configuration file and optionally change the configuration file through an API endpoint), binaries, and source code (which may be compiled when the image is booted). In step 205.7, the hardware-specific instructions corresponding to the elements selected in steps 205.4, 205.5, and 205.6 may be loaded as part of the image to be booted. The controller logic 205 derives an image from the selected components. For example, there may be different pre-installation scripts for physical host versus virtual machine, or differences for Powerpc versus x86.

[0096] In step 205.8, the controller logic 205 mounts overlayfs and repackages the target files into a single file system blob. When multiple file system blobs are used, multiple blobs may be used to create an image, untar the tarball, and / or fetch the git. If step 205.8 is not executed, the file system blobs may remain separate, and the image is created as a set of file system blobs and mounted using a file system that can mount multiple smaller file systems (such as overlayfs) together. The controller logic 205 may then find a compatible kernel (or the kernel specified in system rule 210) in step 205.9 and find an applicable initrd in step 205.10. A compatible kernel may be a kernel that satisfies the dependencies of the template and the resources used to implement the template. A compatible initrd may be an initrd that loads the template onto the desired computing resources. Often, the initrd may be used on physical resources so that it can mount storage resources (since the root file system may be remote) before fully booting. The kernel and initrd may be packaged into the file system blob, and the file system blob may be used directly for kernel booting using kexec to change the kernel on the running system after booting a preliminary operating system, or may be used on the physical host.

[0097] The controller then configures the storage resource(s) to enable running the application(s) and / or image(s) using any of the techniques indicated by computing resource(s) 205.11, 205.12, and / or 205.13. According to 205.11, an overlayfs file can be provided as a storage resource. According to 205.12, a file system is presented. For example, the storage resource may present a combined file system or multiple file system blobs on which computing resources can be simultaneously mounted using a file system similar to overlayfs. According to 205.13, the blob is sent to the storage resource before presenting the file system.

[0098] Figures 2G and 2H show exemplary process flows for steps 205.11 and 205.12 of FIG. 2F. Further, the system can adopt processes and rules for connecting computer resources to storage resources, which may be referred to as a storage connection process. Examples of such storage connection processes, in addition to the processes shown by FIGS. 2G and 2H, are provided in Appendix A included in the specification. FIG. 2G shows an exemplary process flow for connecting storage resources. Some storage resources may be read-only, and others may be writable. The storage resources may manage their write locks so that there is no concurrent writing that causes a race condition, or the system state 220 may track which connections can write to the storage resources (see, for example, step 205.20), and / or prevent multiple read-write connections to the resources (step 205.21). The controller logic or the resource itself may query the system state 220 of the controller about the location and transmission type of the storage resource (e.g., Internet Small Computer System Interface (ISCSI, iSCSI, or iscsi), ISCSI Extension for Remote Direct Memory Access (RDMA or rdma) (ISER, iSER, or iser), Non-Volatile Memory Express over Fabrics (NVMEOF or nvmeof), Fibre Channel (FC or fc), Fibre Channel over Ethernet (FCOE, FCoE, or fcoe), Network File System (NFS or nfs), nfs over rdma, Andrew File System (AFS or afs), Common Internet File System (CIFS or cifs), windows share) (step 205.22). If the computing resources are virtual, a hypervisor (e.g., via a hypervisor daemon) may handle the connection to the storage resources (step 205.23). This may have desirable security benefits since the virtual machine may not recognize the SAN 280.

[0099] Referring to step 205.24, the process of connecting the computing resources and the storage resources may be directed in system rule 210. The controller logic then queries the system state 220 to confirm that the resources are available and writable if necessary (step 205.22). The system state 220 can be queried via any of a plurality of techniques such as SQL queries (or other types of database queries), JSON parsing, etc. The query returns the information necessary for the computing resources to connect to the storage resources. The controller 200, the system state 220, or the system rule 210 may provide authentication credit information for the computing resources to connect to the system state (step 205.25). The computing resources then update the system state 220 either directly or via the controller (step 205.26).

[0100] Figure 2H shows an exemplary boot process for a physical, virtual, or other type of computing resource, application, service, or host to power on and connect to a storage resource. The storage resource may optionally utilize a fused file system and / or expandable volumes. In situations where a controller or other system enables a physical host, the physical host may be preloaded by an operating system for configuring the system. Thus, at step 205.31, the controller may preload the boot disk with initramfs. Also, the controller 200 may use an out-of-band management connection 260 to network boot a preliminary operating system (step 205.30), and then, optionally, preload the host with the preliminary operating system (step 205.31). The initramfs is then loaded at step 205.32, and the storage resource is connected at step 205.33 using the method shown in Figure 2G. Next, if expandable volumes exist, the subvolumes or devices to be combined together are optionally assembled as a volume group at step 205.34 if logical volume management (LVM) is in use. Or they may be combined at step 205.34 using other methods of combining disks.

[0101] If a merged file system is being used, in step 205.36, files may be combined and then the boot process may continue (step 205.46). If overlayfs is used in Linux to fix some known issues, the following subprocess may be executed. In each mounted file system blob that may be volatile, a / data directory may be created (step 205.37). Next, in step 205.38, a new_root directory may be created, and in step 205.39, overlayfs is mounted on the directory. Next, initramfs executes exec_root on / new_root (step 205.40).

[0102] When the host is a virtual machine (VM), additional tools such as direct kernel boot may be available. In this situation, the hypervisor may connect to the storage resource before booting the VM (step 205.41), or it may do so while the VM is booting. The VM may then be a direct kernel that boots while loading the initramfs (step 205.42). The initramfs is then loaded in step 205.43, and the hypervisor may connect to the storage resource, which may be remote at this point (step 205.44). To achieve this, the hypervisor host may need to pass in to the interface (for example, if inifiniband is required to connect to the iSER target, it may pass in to SR-IOV based on virtual functions using pci-passhtru, or in some situations, a paravirtualized network interface may be used). Those connections are available for use by the initramfs. The virtual machine may then connect to the storage resource in step 205.45 if it has not already been connected. The virtual machine may also receive its storage resource through the hypervisor (optionally, through paravirtualized storage). The process may optionally be similar for a virtual machine that has a fused file system and LVM style disks mounted.

[0103] As shown at 205.13, FIG. 2O shows an exemplary process flow for constructing storage resources from a file system blob or other group of files. The blobs are collected at step 205.75 and may be copied directly to the storage resource host at 205.73 (if the storage resource host is different from the device holding the file system blob 232). Once the storage resources are in place, the system state is then updated at 205.74 based on the location of the storage resources and available transports (e.g., iSER, nvmeof, iSCSI, FcoE, Fibre Channel, nfs, nfs over rdma). Some of those blobs may be read-only, in which case the system state remains the same and new compute resources or hosts may connect to that read-only storage resource (e.g., when connecting to a base image). In some cases, as shown at 205.70, it may be desirable to place files into a single file system image to avoid the overhead of any fusion file system. This may be accomplished by mounting the blobs as a fusion file system (step 205.71), then copying them to a new file system or repackaging them as a single file system (step 205.72), and then optionally copying the new file system image to an appropriate location for presentation as a storage resource. Some fusion file systems may allow merging without first mounting at step 205.71, potentially allowing the fusion file system to be merged in a single step.

[0104] FIG. 2I shows another exemplary template 230 as shown in FIG. 2E. In this example, the controller may be configured to use the template 230 as shown in FIG. 2I using an intermediate configuration tool. According to an exemplary embodiment, the intermediate configuration tool may include a common API used to couple a new application or service to a dependent application or service. Thus, the template 230 may additionally include a dependency list 244 that may be required for setting up the services of the template. The template 230 may also include connection rules 245 that may include calls to the common API of the dependencies. The template 230 may also include one or more common APIs 243 and a list 242 of common APIs and versions. The common API 243 may have methods, functions, scripts, or instructions that may be callable (or non-callable) from an application or controller that enable the controller to configure the dependent application or service, whereby the dependent application or service may be coupled to the new application constructed by the template 230. The controller may communicate with the common API 243 and / or make API calls to configure the coupling of the new service or application to the dependent service or application. Alternatively, the instructions may enable an application or service to communicate directly with the common API 243 on the dependent application or service and / or send calls. The connection rules 245 of the template 230 are a set of rules and / or instructions that may include API calls regarding connecting a new service or application to a dependent service or application.

[0105] System state 220 may further include a list 246 of running services. The list 246 of running services may be queried by the controller logic 205 to determine that it meets the dependencies 244 from the template 230. The controller may also include a list 247 of various common APIs available for a particular service / application or type of service / application, and may also include a template that includes the common APIs. The list may exist in the controller logic 205, the system rules 210, the system state 220, or the template storage accessible by the controller. The controller also maintains an index 248 of common APIs compiled from all existing or loaded templates.

[0106] Figure 2J relates to the controller logic 205 processing the template 230 as shown by Figure 2F, and step 255 shows an exemplary process flow in which the controller manages service dependencies. Figure 2K shows the exemplary process flow of step 255 of Figure 2J. In step 255.1, the controller collects the dependency list 244 from the template. The controller also collects a list 243 of common APIs from the template. (A) In step 255.2, the controller filters the list of possible dependent applications or services by comparing the list 243 of common APIs from the template with the common API index 248 and based on the type of application or service that is considered to meet the dependencies. In step 255.3, the controller determines whether the system rules 210 specify a way to meet the dependencies.

[0107] If yes in step 255.3, the controller determines whether the dependent service or application is running by querying the list of templates being executed (step 255.4). If no in step 255.4, a service application that may include the controller logic processing the template of the dependent service / application is executed (and / or configured and then executed) (step 255.5). If it is determined in step 255.4 that the dependent service or application is running, the process flow proceeds to step 255.6. In step 255.6, the controller uses the template to couple the new service or application to be constructed to the dependent service or application. When coupling the new service or application to the dependent application / service, the controller considers the template it is processing and executes connection rule 245. The controller sends commands to the common API 243 regarding how to satisfy the dependency 244 and / or how to couple the application / service based on connection rule 245. The common API 243 converts the commands from the controller to connect the new service or application to the dependent application or service, which may include but is not limited to calling the API functions of the service, changing the configuration, executing the script, and calling other programs. Following step 255.6, the process flow proceeds to step 205.2 of FIG. 2J.

[0108] If step 255.3 determines that system rule 210 does not specify a way to satisfy the dependency, the controller queries system state 220 in step 255.7 to check whether the appropriate dependent application or service is running. In step 255.8, the controller makes that determination based on the query about whether the appropriate dependent application or service is running. If it is no in step 255.8, the controller notifies the action administrator or user (step 255.9). If it is yes in step 255.8, the process flow proceeds to step 255.6 where it can operate as described above. Optionally, the user may be queried about whether to connect to the dependent application with the new application running, in which case, in step 255.6, the controller may couple the new application or service to the dependent application or service as follows, and the controller considers the template 230 it is processing and executes the connection rule 245. The controller then sends a command to the common API 243 based on the connection rule 245 regarding how to satisfy the dependency 244. The common API 243 translates the command from the controller to connect the new service or application and the dependent application or service.

[0109] The user may communicate with the controller 200 through an external user interface or web UI, or through an API application 120 that may also be incorporated into the controller application or logic 205 by the application.

[0110] The controller 200 communicates with the stack or resources by one or more of a plurality of networks, interconnections, or other connections through which the controller can be instructed to operate on computing resources, storage resources, and networking resources. Such connections may include an out-of-band management connection 260, an in-band management connection 270, a SAN connection 280, and an optional network in-band management connection 290.

[0111] Out-of-band management may be used by the controller 200 to detect, configure, and manage components of the system 100 through the controller 200. The out-of-band management connection 260 may enable the controller 200 to detect resources that are plugged in and available but not powered on. Resources may be added to the IT system state 220 when plugged in. Out-of-band management may be configured to load a boot image and configure and monitor resources belonging to the system 100. Out-of-band management may also boot a temporary image for operating system diagnostics. Out-of-band management may be used to change BIOS settings and may also use console tools to execute commands on a running operating system. Settings may also be changed by the controller using image recognition of video signals from physical or virtual monitor ports on hardware resources such as a console, keyboard, and VGA, DVI, or HDMI ports, and / or using APIs provided by out-of-band management, such as Redfish.

[0112] Out-of-band management, as used herein, may include, but is not limited to, a management system that can connect to resources or nodes independent of the operating system and the main motherboard. The out-of-band management connection 260 may include a network, or multiple types of direct or indirect connections or interconnections. Examples of types of out-of-band management connections include, but are not limited to, IPMI, Redfish, SSH, telnet, other management tools, keyboard video and mouse (KVM), or KVM over IP, serial console, or USB. Out-of-band management can turn the power of a node or resource on or off, can monitor temperature and other system data, can make BIOS and other low-level changes that may not be under the control of the operating system, can connect to a console and send commands, and can control inputs including, but not limited to, a keyboard, mouse, and monitor, and is a tool that can be used over a network. Out-of-band management may be coupled to an out-of-band management circuit within a physical resource. Out-of-band management may connect a disk image as a disk that can be used to boot an installation media.

[0113] The management network or in-band management connection 270 may enable the controller to collect information regarding computing resources, storage resources, networking resources, or other resources and communicate directly with the operating system running on the resources. The storage resources, computing resources, or networking resources may include a management interface that interfaces with connection 260 and / or 270, whereby they can communicate with the controller 200, can notify the controller of what is being executed and available with respect to the resources, and can receive commands from the controller. An in-band management network as used herein includes a management network capable of communicating directly with the resources and the operating systems of the resources. Examples of in-band management connections may include, but are not limited to, SSH, telnet, other management tools, serial console, or USB.

[0114] Out-of-band management is described herein as a network physically or virtually separate from the in-band management network, but they may be combined with each other for efficiency or operate in cooperation with each other as described in more detail herein. Also, thus, out-of-band and in-band management or aspects thereof may communicate through the same port of the controller or may be coupled with a combined interconnect. Optionally, one or more of connections 260, 270, 280, 290 may be separated from or combined with the rest of such networks, may include or may not include the same fabric.

[0115] Furthermore, the computing resources, storage resources, and the controller may or may not be coupled to a storage network (SAN) 280 in a manner that allows the controller 200 to use the storage network to boot each resource. The controller 200 may send a boot image or other template to a separate storage or other resource, such that the other resource can boot from the storage or other resource. The controller may indicate where to boot from in such a situation. The controller may turn on the power of the resource and indicate to the resource where to boot from and how to configure itself. The controller 200 indicates to the resource how to boot, which image to use, and where the image is located if the image is on another resource. The BIOS resources may be pre-configured. The controller may additionally or alternatively configure the BIOS through out-of-band management such that the BIOS boots from the storage area network. The controller 200 may also be configured to boot an operating system from an ISO and enable the resource to copy data to a local disk. Next, the local disk may subsequently be used to boot. The controller may configure other resources, including other controllers, such that the resources can boot. Some resources may include applications that provide computing, storage, or networking functionality. Additionally, the controller can perform the role of booting up the storage resources and then supplying subsequent resources or service boot images to the storage resources. The storage may also be managed through a different network that is used for other purposes.

[0116] Optionally, one or more of the resources may be coupled to the network in-band management connection 290. The connection 290 may include one or more types of in-band management as described for the in-band management connection 270. The connection 290 may connect the controller to the application network to utilize or manage the network through the in-band management network.

[0117] FIG. 2L shows an image 250 that can be loaded from template 230 directly or indirectly (through another resource or database) to boot a resource, or an application or service loaded on the resource. Image 250 may include a boot file 240 for the resource type and hardware. Boot file 240 may include a kernel 241 corresponding to the resource, application, or service to be deployed. Boot file 240 may also include an initrd or similar file system used to assist the boot process. Boot system 240 may include multiple kernels or initrds configured for different hardware types and resource types. Further, image 250 may include a file system 251. File system 251 may include a base image 252 and corresponding file system, a service image 253 and corresponding file system, and a volatile image 254 and corresponding file system. The file system and the loaded data may vary according to the resource type and the running application or service. Base image 252 may include a base operating system file system. The base operating system may be read-only. Base image 252 may also include basic tools of an operating system independent of what is being executed. Base image 252 may include a base directory and operating system tools. Service file system 253 may include configuration files and specifications for the resource, application, or service. Volatile file system 254 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may be configured as variables including, but not limited to, passwords, session keys, and private keys, or may not be configured.The file system may be mounted as a single unified file system using technologies such as overlayFS, enabling some read-only and some read-write file systems to reduce the amount of replicated data used for applications.

[0118] As described above, the controller 200 can be used to add resources such as computing resources, storage resources, and / or networking resources to the system. FIG. 11A shows an exemplary method of adding a physical resource such as a bare metal node to the system 100. A resource, i.e., a computing resource, a storage resource, or a networking resource, is plugged into the controller via a network connection (1110). The network connection may include an out-of-band management connection. The controller recognizes that the resource has been plugged in via the out-of-band management connection (1111). The controller recognizes information related to the resource, which may include but is not limited to the type, capabilities, and / or attributes of the resource (1112). The controller adds the resource and / or the information related to the resource to its system state (1113). An image derived from a template is loaded onto a physical component of the system, which may include but is not limited to a resource, another resource such as a storage resource, or the controller (1114). The image includes one or more file systems that may include a configuration file. Such a configuration may include BIOS and boot parameters. The controller instructs the physical resource to boot using the file system of the image (1115). Additional resources, or multiple bare metal or physical resources of different types, may thus be added using the image of the template or at least a portion thereof.

[0119] FIG. 11B illustrates an exemplary method for automatically allocating resources using the global system rules and templates of an exemplary embodiment. A request is made to the system (1120) that requires resource allocation to meet the request. The controller identifies the resource pool based on its system state database (1121). The controller uses the template to determine the required resources (1122). The controller allocates the resources and stores the information in the system state (1123). The controller deploys the resources using the template (1124).

[0120] Referring to FIG. 12, an exemplary method for automatically deploying an application or service using the system 100 described herein is shown. A user or application requests a service (1210). The request is converted into an API application (1220). The API application routes the request to the controller (1230). The controller interprets the request (1240). The controller considers the state of the system and its resources (1250). The controller uses its rules and templates for service deployment (1260). The controller sends the request to the resources (1270), deploys the image derived from the template (1280), and updates the IT system state.

[0121] Additional and more detailed examples of operations such as adding resources, allocating resources, and deploying an application or service are described in more detail below.

[0122] Adding Computational Resources to the System Referring to FIG. 3A, the addition of computing resource 310 to system 100 is shown. When computing resource 310 is added, computing resource 310 may be coupled to controller 200 and powered off. If computing resource 310 is preloaded by an image, alternative steps may follow in which any of the network connections are used to communicate with the resource, boot the resource, and add information to the system state. If the computing resource and the controller are on the same node, the service running the computing resource is off.

[0123] As shown in FIG. 3A, the computing resource 310 is coupled to the controller by a network, namely, the out-of-band management connection 260, the in-band management connection 270, and optionally, the SAN 280. The computing resource 310 is also coupled to one or more application networks 390 through which services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or circuit of the computing resource 310 that is turned on when the computing resource 310 is plugged in. The device 415 may enable features including, but not limited to, power on / off of the device, attachment to the console and entry of commands, monitoring of temperature and other computer health-related elements, and setting of BIOS settings and other features outside the scope of the operating system. The controller 200 can identify the computing resource 310 through the out-of-band management network 260. The controller 200 may also identify the type of computing resource and its configuration using in-band management or out-of-band management. The controller logic 205 is configured to look for added hardware and examine the out-of-band management 260 or in-band management 270. If the computing resource 310 is detected, the controller logic 205 may use the global system rules 220 to determine whether the resource is configured automatically or by interacting with the user. If the resource is added automatically, the setup follows the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may query the user to confirm the addition of the resource and what the user wishes to do with the computing resource. The controller 200 may query the API application(s) to confirm that the new resource has been approved, or in other cases, request it from the user or any program that controls the stack.The approval process may also use encryption to automatically and securely complete the verification of the legitimacy of new resources. The controller logic 205 adds the computing resource 310 to the IT system state 220, which includes the switch or network to which the computing resource is plugged in.

[0124] If the computing resource is physical, the controller 200 may turn on the power of the computing resource through the out-of-band management network 260. The computing resource 310 may boot from the image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, by the SAN 280. The image may be loaded indirectly through other network connections or by other resources. Once booted, the information related to the computing resource 310 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 310 may then be added to the storage resource pool and become a resource managed by the controller 200 and tracked in the IT system state 220.

[0125] If the computing resource is virtual, the controller 200 may turn on the power of the computing resource through the in-band management network 270 or through the out-of-band management 260. The computing resource 310 may boot from the image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, by the SAN 280. The image may be loaded indirectly through other network connections or by other resources. Once booted, the information related to the computing resource 310 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 310 may then be added to the storage resource pool and become a resource managed by the controller 200 and tracked in the IT system state 220.

[0126] The controller 200 can automatically turn resources on and off according to global system rules for reasons determined by the IT system user, such as to save power by turning off resources, or to improve application performance by turning on resources, or for other reasons that the IT system user may have, and can update the IT system state.

[0127] Figure 3B shows an image 350 that is loaded directly or indirectly (through another resource or database) from template 230 to computing resource 310 to boot the computing resource and / or load an application. Image 350 may include a boot file 340 for the resource type and hardware. Boot file 340 may include a kernel 341 corresponding to the resource, application, or service to be deployed. Boot file 340 may also include an initrd, or a similar file system used to assist the boot process. Boot system 340 may include multiple kernels or initrds configured for different hardware types and resource types. Further, image 350 may include a file system 351. File system 351 may include a base image 352 and corresponding file system, a service image 353 and corresponding file system, and a volatile image 354 and corresponding file system, along with the corresponding file systems. The file system and loaded data may vary according to the resource type and the running application or service. Base image 352 may include the file system of the base operating system. The base operating system may be read-only. Base image 352 may also include the basic tools of an operating system independent of what is being executed. Base image 352 may include a base directory and operating system tools. Service file system 353 may include configuration files and specifications for the resource, application, or service. Volatile file system 354 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may be configured as variables including, but not limited to, passwords, session keys, and private keys, or may not be configured at all.The file system may be mounted as a single unified file system using technologies such as overlayFS, enabling some read-only and some read-write file systems to reduce the amount of replicated data used for applications.

[0128] Figure 3C shows an exemplary process flow for adding a resource, such as a computing resource 310, to system 100. In this example, the resource of interest is described as a computing resource 310, but it should be understood that the resource of interest for the process flow of Figure 3C may be a storage resource 410 and / or a networking resource 510. In the example of Figure 3C, the resource 310 being added is not on the same node as controller 200. In step 300.1, resource 310 is coupled to a powered-off controller 200. In the example of Figure 3C, an out-of-band management connection 260 is used to connect resource 310. However, it should be understood that other network connections may be used if required by an implementer. In steps 300.2 and 300.3, controller logic 205 examines the system's out-of-band management connection and uses the out-of-band management connection 260 to recognize and identify the type and its configuration of the added resource 310. For example, the controller logic can check the BIOS or other information about the resource (such as serial number information) as a reference for obtaining type and configuration information.

[0129] In step 300.4, the controller uses the global system rules to determine whether a particular resource 310 should be automatically added. If it should not be automatically added, the controller waits until its use is approved (step 300.5). For example, the user may respond to the query that they do not wish to use a particular resource 310, or that it may be automatically held until it is to be used in step 300.4. If step 300.4 determines that the resource 310 should be automatically added, the controller uses that rule for automatic setup (step 300.6) and proceeds to step 300.7.

[0130] In step 300.7, the controller selects and uses template 230 associated with the resource to add the resource to system state 220. In some cases, template 230 may be specific to a particular resource. However, some templates 230 may cover multiple resource types. For example, some templates 230 may be hardware-independent. In step 300.8, the controller turns on the power of resource 310 through its out-of-band management connection 260 in accordance with global system rule 210. In step 300.9, using global system rule 210, the controller finds and loads a boot image for the resource from the selected template(s). Resource 310 then boots from the image derived from the target template 230 (step 300.10). Additional information about resource 310 may then be received from resource 310 through the in-band management connection 270 after resource 310 has booted (step 300.11). Such information may include, for example, firmware version, network card, any other device to which the resource may be connected. In step 300.12, new information may be added to system state 220. Resource 310 may then be considered added to the resource pool and prepared for allocation (step 300.13).

[0131] Regarding FIG. 3C, it should be understood that when the resource and the controller are on the same node, the service that executes the resource may be outside that node. In such a case, the controller may use an inter - process communication technique with the resource, such as, for example, a unix socket, a loopback adapter, or other inter - process communication techniques between processes that communicate with the resource. From the system rules, the controller may install a virtual host, or a hypervisor or container host from the controller to execute an application using a known template. Next, the resource application information can be added to the system state 220, and the resource will be ready for allocation.

[0132] Addition of Storage Resources to the System FIG. 4A shows the addition of storage resource 410 to system 100. In an exemplary embodiment, the exemplary process flow of FIG. 3C can be followed to add storage resource 410 to system 100, where the storage resource 410 being added is not on the same node as controller 200. Also, it should be noted that if an image is pre - loaded on storage resource 410, alternative steps may be followed to communicate with storage resource 410, boot storage resource 410, and add information to system state 220 using any of the network connections.

[0133] When a storage resource 410 is added, the storage resource 410 may be coupled to the controller 200 and be powered off. The storage resource 410 is coupled to the controller via a network, namely, an out-of-band management network 260, an in-band management connection 270, a SAN 280, and optionally, a connection 290. The storage resource 410 may also be coupled, or not coupled, to one or more application networks 390 through which services, application users, and / or clients can communicate with each other. An application or client may have direct or indirect access to the storage of the resource through the application, whereby the application or client is not accessed through the SAN. The application network may have storage incorporated therein, or be accessed and be identified as a storage resource in the IT system state. The out-of-band management connection 260 may be coupled to an independent out-of-band management device 415 or circuit of the storage resource 410 that turns on when the storage resource 410 is plugged in. The device 415 may enable features including, but not limited to, power on / off of the device, attachment to the console and input of commands, monitoring of temperature and other computer health-related elements, and setting of BIOS settings and other features outside the scope of the operating system. The controller 200 may refer to the storage resource 410 through the out-of-band management network 260. The controller 200 may also identify the type of the storage resource and identify its configuration using in-band or out-of-band management. The controller logic 205 is configured to look for added hardware and check the out-of-band management 260 or in-band management 270. When the storage resource 410 is detected, the controller logic 205 may use the global system rules 220 to determine whether the resource 410 should be configured automatically or through interaction with the user.When resource 410 is added automatically, the setup follows the global system rules 210 within the controller 200. When resource 410 is added by the user, the global system rules 210 within the controller 200 may request the user to confirm the addition of the resource and what the user wants to do with the storage resource. The controller 200 may query the API application(s) to confirm that the new resource has been approved, or in other cases, request it from the user or any program that controls the stack. The approval process may also be automatically and securely completed using encryption to confirm the legitimacy of the new resource. The controller logic 205 adds the storage resource 410 to the IT system state 220, which includes the switch or network to which the storage resource 410 is plugged in.

[0134] The controller 200 may turn on the power of the storage resource 410 through the out-of-band management network 260, and the storage resource 410 boots from the image 450 loaded from the template 230, for example, via the SAN 280, using the global system rules 210 and the controller logic 205. The image may also be loaded indirectly through other network connections or via another resource. Once booted, information received through the in-band management connection 270 regarding the storage resource 410 is also collected and added to the IT system state 220. Here, the storage resource 410 is added to the storage resource pool, and the storage resource 410 becomes a resource that is managed by the controller 200 and tracked in the IT system state 220.

[0135] Storage resources may comprise a storage resource pool or multiple storage resource pools that an IT system can use or access independently or simultaneously. When storage resources are added, the storage resources may provide a storage pool, multiple storage pools, a portion of a storage pool, and / or multiple portions of a storage pool to the IT system state. The controller and / or storage resources may manage the various storage resources of the pool, or groups of such resources within the pool. A storage pool may include multiple storage pools executed on multiple storage resources. For example, a flash storage disk or array that caches a platter disk or array, or a storage pool on a dedicated computing node coupled to a pool on a dedicated storage node optimizes bandwidth and latency simultaneously.

[0136] Figure 4B shows an image 450 that is directly or indirectly loaded from template 230 (or from another resource or database) to storage resource 410 to boot the storage resource and / or load an application. The image 450 may include a boot file 440 for the resource type and hardware. The boot file 440 may include a kernel 441 corresponding to the resource, application, or service to be deployed. The boot file 440 may also include an initrd or similar file system used to assist in the boot process. The boot system 440 may include multiple kernels or initrds configured for different hardware types and resource types. Further, the image 450 may include a file system 451. The file system 451 may include a base image 452 and corresponding file system, a service image 453 and corresponding file system, and a volatile image 454 and corresponding file system, along with the corresponding file systems. The file systems and data to be loaded may vary depending on the resource type and application, or the service being executed. The base image 452 may include a base operating system file system. The base operating system may be read-only. The base image 452 may also include the basic tools of an operating system that is independent of what is being executed. The base image 452 may include a base directory and operating system tools. The service file system 453 may include configuration files and specifications for the resource, application, or service. The volatile file system 454 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may be configured as variables including, but not limited to, passwords, session keys, and private keys, or may not be configured at all.The file system may be mounted as a single unified file system using technologies such as overlayFS, enabling some read-only and some read-write file systems to reduce the amount of replicated data used for applications.

[0137] FIG. 5A shows an example in which another storage resource, namely, direct attached storage 510, which may take the form of a node by JBOD or other type of direct attached storage, is coupled to storage resource 410 as an additional storage resource for the system. JBOD is typically an external disk array connected to a node that provides storage resources. In FIG. 5A, JBOD is used as an exemplary form of direct attached storage 510, but it should be understood that other types of direct attached storage may be employed as 510.

[0138] The controller 200 may add, for example, the storage resource 410 and the JBOD 510 to its system as described with respect to FIG. 5A. The JBOD 510 is coupled to the controller 200 via the out-of-band management connection 260. The storage resource 410 is coupled to a network, namely, the out-of-band management connection 260, the in-band management connection 270, the SAN 280, and optionally, the connection 290. The storage node 410 communicates with the storage of the JBOD 510 through the SAS or other disk drive fabric 520. The JBOD 510 may also include an out-of-band management device 515 that communicates with the controller through the out-of-band management connection 260. Through the out-of-band management 260, the controller 200 may detect the JBOD 510 and the storage resource 410. The controller 200 may also detect other parameters not controlled by the operating system, such as those described herein with respect to various out-of-band management circuits. The controller 200 global system rules 210 provide configuration startup rules for booting or starting JBODs and storage nodes that have not yet been added. The order in which the storage resources are powered on may be controlled by the controller logic 205 using the global rules 220. According to one set of the global system rules 220, the controller may first power on the JBOD 510, and then the controller 200 may power on the storage resources using the loaded image 450 in a manner similar to that described with respect to FIG. 4. In another set of the global system rules, the controller 200 may first power on the storage resource 410 and then the JBOD 510. In other global system rules, the power-on timing or delay between various devices may be specified. Detection of the readiness or operating state of various resources may be determined by the controller logic 205, the global system rules 210, and / or the template 230, and / or may be used in the device allocation management by the controller 200.The IT system state 220 may be updated by communication with the storage resource 410. The storage node 410 recognizes the storage parameters and configuration of the JBOD 510 by accessing the JBOD through the disk fabric 520. The storage resource 410 then provides information to the controller 200 that updates the IT system state 220 with information regarding the amount of available storage and other attributes. When the storage resource 410 is booted and recognized as part of the pool of storage resources 400 of the system 100, the controller updates the IT system state 220. The storage node handles the logic for controlling the JBOD storage resources using the configuration set by the controller 200. For example, the controller may instruct the storage node to configure the JBOD to create a pool from RAID10 or other configurations.

[0139] FIG. 5B shows an exemplary process flow for adding the storage resource 410 and the direct attached storage 510 for the storage resource 410 to the system 100. In step 500.1, the direct attached storage 510 is coupled to the power-off controller 200 via the out-of-band management connection 260. In step 500.2, the storage resource 410 is coupled to the power-off controller 200 via the out-of-band management connection 260 and the in-band management connection 270, and at the same time, the storage resource 410 is coupled to the direct attached storage 510 via, for example, SAS 520 such as a disk drive fabric.

[0140] Next, the controller logic 205 may examine the out-of-band management connection 260 to detect the storage resource 410 and the direct attached storage 510 (step 500.3). Although any network connection can be used, in this example, out-of-band management may be used for the controller logic to recognize and identify the types of resources being added (in this case, the storage resource 410 and the direct attached storage 510) and their configurations (step 500.4).

[0141] In step 500.5, the controller 200 selects and uses a template 230 for a particular type of storage for each type of storage device to add resources 410 and 510 to the system state 220. In step 500.6, the controller powers on the direct storage and the storage nodes in that order, through the out-of-band management connection 260, according to the global system rule 210 that can specify the boot order to power on. Using the global system rule 210, the controller finds and loads the boot image for the storage resource 410 from the template 230 selected for that storage resource 410, and then the storage resource boots from the image (step 500.7). The storage resource 410 recognizes the storage parameters and configuration of the direct attached storage 510 by accessing the direct attached storage 510 through the disk fabric 520. Next, additional information about the storage resource 410 and / or the direct attached storage 510 may be provided to the controller through the in-band management connection 270 to the storage resource (step 500.8). In step 500.9, the controller updates the system state 220 with the information obtained in step 500.8. In step 500.10, the controller sets the configuration for the storage resource 410 that handles the direct attached storage 510 and the way to configure the direct attached storage. Next, in step 500.11, the new resource comprising the storage resource 410 together with the direct attached storage 510 may be added to the resource pool and be ready to be allocated within the system.

[0142] According to another aspect of the exemplary embodiment, the controller may use out-of-band management to recognize other devices in a stack that may not be involved in the computation or service. For example, such devices may include, but are not limited to, cooling towers / air conditioning units, lighting, temperature, sound, alarms, power systems, or any other device associated with the system.

[0143] Adding Networking Resources to the System FIG. 6A shows the addition of networking resource 610 to system 100. In one exemplary embodiment, the exemplary process flow of FIG. 3C can be followed to add networking resource 610 to system 100, where the networking resource 610 being added is not on the same node as controller 200. Also, if an image is pre-loaded on networking resource 610, it should be noted that alternative steps can be followed to communicate with networking resource 610 using any of the network connections, boot networking resource 610, and add information to system state 220.

[0144] When a networking resource 610 is added, the networking resource 610 may be coupled to the controller 200 and be powered off. The networking resource 610 may be coupled to the controller 200 through a connection, i.e., an out-of-band management connection 260 and / or an in-band management connection 270. Optionally, the networking resource 610 is connected to the SAN 280 and / or the connection 290. The networking resource 610 may also be coupled, or may not be coupled, to one or more application networks 390 through which services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 615 or circuit of the networking resource 610 that is turned on when the networking resource 610 is plugged in. The device 615 may enable features including, but not limited to, power on / off of the device, attachment to the console and command entry, monitoring of temperature and other computer health-related elements, and setting of BIOS settings and other features outside the scope of the operating system. The controller 200 may refer to the networking resource 610 through the out-of-band management connection 260. The controller 200 may identify the type of the networking resource and / or the network fabric, and may identify the configuration using in-band or out-of-band management. The controller logic 205 is configured to look for the added hardware and examine the out-of-band management 260 or the in-band management 270. When the networking resource 610 is detected, the controller logic 205 may determine whether the resource 610 should be configured automatically or through interaction with the user using the networking global system rules 220. If the resource 610 is added automatically, the setup follows the global system rules 210 within the controller 200. If added by the user, the global system rules 210 within the controller 200 may ask the user to confirm the addition of the resource and what the user wants to do with the resource.The controller 200 may query the API application(s) to confirm that a new resource has been approved, or, in other cases, request it from any program that controls the user or stack. The approval process may also be automatically and securely completed using encryption to verify the legitimacy of the new resource. Next, the controller logic 205 may add the networking resource 610 to the IT system state 220. For switches that cannot identify themselves to the controller, the user may manually add them to the system state.

[0145] If the networking resource is physical, the controller 200 may turn on the power of the networking resource 610 through the out-of-band management connection 260. The networking resource 610 may boot from the image 605 loaded from the template 230, for example, via the SAN 280, using the global system rules 210 and the controller logic 205. The image may also be loaded indirectly through other network connections or via other resources. Once booted, information received through the in-band management connection 270 for the networking resource 610 may also be collected and added to the IT system state 220. The networking resource 610 is added to the storage resource pool and becomes a resource that is managed by the controller 200 and tracked in the IT system state 220. Optionally, some networking resource switches may be controlled through a console port connected to the out-of-band management 260, configured upon power-on, or have a switch operating system installed through a bootloader, for example, through ONIE.

[0146] When the networking resources are virtual, the controller 200 may turn on the power of the networking resources through the in-band management network 270 or through the out-of-band management 260. The networking resource 610 may boot from an image 650 loaded from the template 230 via the SAN 280 using the global system rules 210 and the controller logic 205. Once booted, information received through the in-band management connection 270 for the networking resource 610 may be collected and added to the IT system state 220. Next, the networking resource 610 is added to the storage resource pool, and the networking resource 610 becomes a resource managed by the controller 200 and tracked in the IT system state 220.

[0147] The controller 200 may direct port assignments, reassignments, or moves to physical or virtual networking resources to connect to different physical or virtual resources, i.e., connections, storage, or computing as defined herein. This may be done using techniques including but not limited to SDN, Infiniband partitioning, VLAN, vXLAN. The controller 200 may direct a virtual switch to move or assign virtual interfaces for network or interconnect communication with the virtual switch or the resources hosting the virtual switch. Some physical or virtual switches may be controlled by an API coupled to the controller.

[0148] When such changes are possible, the controller 200 may also direct a change in the fabric type to a computing resource, a storage resource, or a networking resource. Ports may be configured to switch to different fabrics, e.g., to switch the fabric of a hybrid Infiniband / Ethernet interface.

[0149] The controller 200 may issue commands to a networking resource that may include a switch for switching among a plurality of application networks or other networking resources. The switch or network device may comprise different fabrics or may be plugged into, for example, preferably, an Infiniband switch, a ROCE switch, and / or other switches by means of SDN capabilities and multiple fabrics.

[0150] Figure 6B shows an image 650 that is directly or indirectly loaded from template 230 (e.g., via another resource or database) to networking resource 610 to boot networking resources and / or load an application. Image 650 may include a boot file 640 for the resource type and hardware. Boot file 640 may include a kernel 641 corresponding to the resource, application, or service to be deployed. Boot file 640 may also include an initrd or similar file system used to assist in the boot process. Boot system 640 may include multiple kernels or initrds configured for different hardware types and resource types. Further, image 650 may include a file system 651. File system 651 may include a base image 652 and corresponding file system, a service image 653 and corresponding file system, and a volatile image 654 and corresponding file system, along with the corresponding file systems. The file systems and data loaded may vary depending on the resource type and application, or the service being executed. Base image 652 may include a base operating system file system. The base operating system may be read-only. Base image 652 may also include basic tools of the operating system independent of the one in execution. Base image 652 may include a base directory and operating system tools. Service file system 653 may include configuration files and specifications for the resource, application, or service. Volatile file system 654 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may be configured as variables including, but not limited to, passwords, session keys, and private keys, or may not be configured.The file system may be mounted as a single unified file system using technologies such as overlayFS, enabling some read-only and some read-write file systems to reduce the amount of replicated data used for applications.

[0151] Deployment of an application or service onto a resource FIG. 7A shows a system 100 comprising a controller 200, physical and virtual computing resources including a first computing node 311, a second computing node 312, and a third computing node 313, storage resources 410, and network resources 610. The resources are shown as being set up and added to an IT system state 220 in a manner as described herein with respect to FIGS. 1 - 6B.

[0152] Although multiple computing nodes are shown in this figure, a single computing node may be used according to an exemplary embodiment. The computing node may host physical or virtual computing resources and may execute an application on a physical or virtual computing node. Similarly, although a single network provider node and storage node are shown, it is contemplated that multiple resource nodes of these types may or may not be used in the system of the exemplary embodiment.

[0153] A service or application may be deployed in any of the systems according to the exemplary embodiments. An example of deploying a service on a computing node may be described with respect to FIG. 7A, but may be used similarly in different arrangements of system 100. For example, the controller 200 in FIG. 7A may automatically configure the computing resources 310 in the form of computing nodes 311, 312, 313 according to the global system rules 210. Next, they may be added to the IT system state 220. Thus, the controller 200 may recognize the computing resources 311, 312, 313 (which may be powered off or not powered off), and optionally any physical or virtual applications running on the computing resources or nodes. The controller 200 may automatically configure the storage resources (s) 410 and the networking resources (s) 610 according to the global system rules 210 and the template 230, and add them to the IT system state 220. The controller 200 may recognize the storage resources 410 and the networking resources 610 that may or may not start in the powered-off state.

[0154] FIG. 7B shows an exemplary process for adding resources to the IT system 100. At step 700.1, a new physical resource is coupled to the system. At step 700.2, the controller recognizes the new resource. The resource may be connected to remote storage (step 700.4). At step 700.3, the controller configures how to boot the new resource. Any connections made to the resource can be logged to the system state 220 (step 700.5). FIG. 3C described above provides further details regarding an exemplary embodiment of a process flow such as that shown in FIG. 7B.

[0155] Figures 7C and 7D illustrate an exemplary process flow for the deployment of an application to a plurality of computing resources, a plurality of servers, a plurality of virtual machines, and / or a plurality of sites. The process for this example differs from standard template deployment in that the IT system 100 requires components that couple redundant and interrelated applications and / or services. The controller logic may process the meta-template at step 700.11, where the meta-template may include a plurality of templates 230, a file system blob 232, and other components (which may be in the form of other templates 230) required to configure the multi-homed service.

[0156] At step 700.12, the controller logic 205 checks the system state 220 of the available resources. However, if there are not sufficient resources, the controller logic may reduce the number of redundant services that can be deployed (see 700.16 where the number of redundant services is identified). At step 700.13, the controller logic 205 configures the networking resources and interconnections required to connect the services together. If a service or application is deployed across multiple sites, the meta-template may include services optionally composed of templates that enable data synchronization and interoperability across the sites (or the controller logic 205 may configure them) (see 700.15).

[0157] In step 700.16, the controller logic 205 may determine the number of redundant services from the system rules, the meta-template data, and the resource availability (if the redundant services are on multiple hosts). At 700.17, there are couplings with other redundant services and with the master. If there are multiple redundant hosts, the controller logic 205 or the logic within the template (which may include a file system blob containing the binary 234, the daemon 232, or a configuration file instructing settings in the operating system) may prevent network address and host name conflicts. Optionally, the controller logic provides a network address (see 700.18) and registers each redundant service with DNS (700.19) and the system state 220 (700.18). The system state 220 tracks the redundant services and, if it notices that a redundant service with conflicting parameters such as a host name (e.g., a software-defined access (SDA) name), a dns name, a network address, etc., is already in the system state 220, the controller logic 205 does not permit duplicate registration.

[0158] The configuration routine shown in FIG. 7D processes the template(s) of the meta-template. The configuration routine processes all redundant services, deploys multi-host or clustered services to multiple hosts, and deploys services that couple the hosts. Any process that can deploy an IT system from the system rules can execute the configuration routine. In the case of a multi-host service, an exemplary routine may process a service template as in 700.32, provision storage resources as in 700.33, turn on the power of the host as in 700.35, and couple (and register with the system state 220) the host / compute resources with the storage resources as in 700.36, (then repeat for the number of redundant services (700.38)). Each time, register with the system state 220 (see 700.20) and use the controller logic to track individual services and record information to prevent conflicts (see 700.31).

[0159] Some service templates may include services and tools that can combine multi-host services. Some of these services may be treated as dependencies (700.39), and then the binding routine of 700.40 may be used for the binding of services and the registration of the binding to system state 220. Further, one of the service templates may be a master template, in which case the dependent service templates of 700.39 are slave or secondary services, and the binding routine of 700.40 connects them. The routine can be defined in a meta-template. For example, for a redundant dns configuration, the binding routine of 700.40 may include the connection of a slave dns to a master dns and the configuration of zone transfer with dnssec. Some services may use physical storage (see 700.34) to improve performance, and the preliminary OS disclosed in FIG. 5B may be loaded therein. Tools for combined services may be included in the template itself, and the configuration between services may be performed by an API accessible to the controller and / or other hosts in a multi-node application / service.

[0160] The controller 200 may enable a user or the controller to determine an appropriate computing backend for use by an application. The controller 200 may enable a user or the controller to optimally place an application on appropriate physical or virtual computing resources by determining the usage status of resources. When a hypervisor or other computing backend is deployed to a computing node, they may report back the controller resource usage statistics through the in-band management connection 270. When the controller decides to create an application on a virtual computing resource from either its own logic and global system rules or user input, the controller may automatically select a hypervisor on the optimal host and turn on the power of the virtual computing resources on that host.

[0161] For example, the controller 200 deploys an application or service to one or more computing resources using a template (s). Such an application or service may be, for example, a virtual machine that runs the application or service. In one example, FIG. 7A shows the deployment of multiple virtual machines (VMs) onto multiple computing nodes, and the illustrated controller 200 can recognize that multiple computing resources 310 are in its computing resource pool in the form of computing nodes 311, 312, 313. The computing nodes may, for example, have a hypervisor deployed, or alternatively, may be on bare metal if the use of virtual machines is not preferred for speed. In this example, the computing resource 310 has a VM(1) 321 and a VM(2) 322 configured and deployed on the computing node 311 with a hypervisor application loaded. For example, if the computing node 311 does not have resources for additional VMs, or if other resources are preferred for a particular service, the controller 200 may recognize, based on the stack state 220, that there are no available resources on the computing node 311 or that it is preferable to set up a new VM with different resources. It can also be recognized that the hypervisor is loaded on the computing resource 312 and not on the resource 313, which may be a bare metal computing node used for other purposes. Thus, according to the requirements of the installed service or application template and the state of the system state 220, the controller in this example may then select the computing node 313 for the deployment of the next required resource VM(3) 323.

[0162] The computing resources of the system may be configured to share storage on the storage resources for the storage nodes.

[0163] The user may request that a service be set up for system 100 through user interface 110 or an application. The service may include, but is not limited to, an email service, a web service, a user management service, a network provider, LDAP, Dev tools, VOIP, authentication tools, accounting.

[0164] API application 120 converts the user's or application's request and sends a message to controller 200. The service template or image 230 of controller 200 is used to identify which resources are needed for that service. Next, the resources to be used are identified based on availability according to IT system state 220. Controller 200 makes requests for one or more of compute nodes 311, 312, or 313 for the required compute service, for storage resource 410 for the required storage resources, and for network resource 610 for the required networking resources. Next, IT system state 220 is updated to identify the allocated resources. Next, the service is installed on the allocated resources using global system rules 210 according to template 230 for the service or application.

[0165] According to an exemplary embodiment, multiple compute nodes may be used whether for the same service or different services, while, for example, a storage service and / or a network provider pool may be shared among the compute nodes.

[0166] Referring to FIG. 8A, a system 100 is shown in which a controller 200, a computing resource 300, a storage resource 400, and a networking resource 600 are on the same or shared physical hardware, such as a single node. The various features shown and described in FIGS. 1 - 10 may be incorporated into a single node. When the node is powered on, the controller image is loaded onto the node. The computing resource 300, the storage resource 400, and the networking resource 600 are configured by a template 230 using global system rules 210. The controller 200 may be configured to load compute backends 318, 319 as computing resources, and the compute backends may or may not be added on the node or on different node(s). Such backends 318, 319 may include, but are not limited to, virtualization, container, and multi - tenant processes to create virtual computing resources, networking resources, and storage resources.

[0167] An application or service 725, such as a web service, an email service, a core network service (DHCP, DNS, etc.), a collaboration tool, may be installed on virtual resources on a node / device shared with the controller 200. These applications or services may be moved to physical or virtual resources independent of the controller 200. The application may be executed on a virtual machine on a single node.

[0168] Figure 8B shows an exemplary process flow for extending from a single node system to a plurality of node systems (e.g., by nodes 318 and / or 319 as shown in Figure 8A). Thus, referring to Figures 8A and 8B, an IT system having a controller 200 running on a single server can be considered, and it is desirable to extend the IT system as a multi-node IT system. Thus, before the extension, the IT system is in a single-node state. As shown in Figure 8A, the controller 200 can include, but is not limited to, storage resources, computing resources, a hypervisor, and / or container hosts, and runs on a multi-tenant single-node system to run various IT system management applications and / or resources.

[0169] In step 800.2, the new physical resources are coupled to the single node system by connecting the new physical resources through the out-of-band management connection 260, the in-band management connection 270, the SAN 280, and / or the network 290. For this example, this new physical resource can also be referred to as hardware or a host. The controller 200 may detect the new resources on the management network and then query the device. Alternatively, the new device may broadcast a message to inform the controller 200 of itself. For example, the new device can be identified by the MAC address, out-of-band management, and / or the use of the pre-boot OS and in-band management and thereby the identification of the hardware type. In any case, in step 800.3, the new device provides information about its node type and its currently available hardware and software resources to the controller. Next, the controller 200 recognizes the new device and its capabilities.

[0170] In step 800.4, a task assigned to the system executing controller 200 may be assigned to a new host. For example, when an operating system (such as a storage host operating system or a hypervisor) is preloaded on the host, controller 200 assigns new hardware resources and / or capabilities. Next, the controller may provide an image to provision the new hardware, or the new hardware may request an image from the controller and configure itself using the methods disclosed above and below. If the new host can host storage resources or virtual computing resources, the new resources can be made available to controller 200. Next, controller 200 may move and / or assign existing applications to the new resources, or use the new resources for newly created or subsequently created applications.

[0171] In step 800.5, the IT system may retain its current applications running on the controller or migrate them to new hardware. When migrating virtual computing resources, VM migration techniques (such as the qemu + kvm migration tool) may be used to update the system state with the new system rules. These changes can be made reliably and safely using the change management techniques described below. Since more applications can be added to the system, the controller may use any of a variety of techniques to determine how to allocate the system's resources, including but not limited to round-robin techniques, weighted round-robin techniques, least utilized techniques, weighted least utilized techniques, prediction techniques based on utilization-based reinforcement learning, planning techniques, desired capacity techniques, and maximum size techniques.

[0172] FIG. 8C shows an exemplary process flow for migrating a storage resource to a new physical storage resource. Next, the storage resource may be mirrored, migrated, or a combination thereof (e.g., the storage may be mirrored and then the original storage resource is disconnected). At step 820, the storage resource is coupled to the system by having its new storage resource contact the controller or by having the controller discover the new storage resource. This can be done using the out-of-band management connection 260, the in-band management connection 270, the SAN network 280, or a flat network that the application network may use, or a combination thereof. By in-band management, the operating system may be pre-booted and the new resource may be connected to the controller.

[0173] At step 822, a new storage target is created on the new storage resource, which can be logged to the database at step 824. In one example, the storage target may be created by copying a file. In another example, the storage target may be created by creating a block device and copying its data (which may be in the form of a file system blob(s)). In another example, by mirroring two or more storage resources between block devices (e.g., creating a RAID) and optionally connecting through a remote storage transport(s) including but not limited to iscsi, iser, nvmeof, nfs, nfs over rdma, fc, fcoe, srp, etc., the storage target may be created. At step 824, if the storage resource is on the same device as another resource or host, the database entry may include information for the compute resource (or other type of resource and / or host) to connect to the new storage resource remotely or locally.

[0174] In step 826, the storage resources are synchronized. For example, the storage can be mirrored. As another example, the storage can be synchronized offline. Technologies such as raid1 (or other types of raid, usually raid1 or raid0, but optionally raid110 (mirrored raid10)) (mdadm, zfs, btrfs, hardware raid) may be employed in step 826.

[0175] Next, data from the old storage resources may optionally be connected after a database login in step 828 (when that is done later, the database may include information related to the state of copying such data if such data has to be recorded). When the storage target is migrating away from the previous host (for example, moving from a single-node system to a multi-node and / or distributed IT system as shown previously, as in FIGS. 8A and 8B), the new storage resources may be designated as primary storage resources in step 830 by a controller, system state, computing resources, or a combination thereof. This may be done as a step of deleting the old storage resources. In some cases, it may then be necessary to update the physical or virtual hosts connected to the resources, and in some cases, in step 832, the power may be turned off during the migration (and then turned back on) (the techniques disclosed herein for turning on the power of a physical or virtual host can be used).

[0176] FIG. 8D shows an exemplary process flow for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for computing and storage. At step 850, the controller 200 creates new storage resources that may be on the new node (see, e.g., nodes 318 and 319 of FIG. 8A). Next, at step 852, the old application host may be powered off. Next, at step 854, the data is copied or synchronized. Prior to the copy / synchronization at step 854, by powering off at step 852, the migration is safer if the migration involves the migration of VMs from a single node. Powering off would also be beneficial for the movement of VMs from a virtual machine to a physical machine. Step 854 may be implemented via a data pre-synchronization step 862 before powering off, thereby minimizing the associated downtime. Further, the host may not be powered off as in step 852, in which case the old host remains online until the new host is ready (or the new storage resources are ready). Techniques for avoiding the power-off step 852 are described in more detail below. At step 854, the data can be optionally synchronized unless the storage resources are mirrored or synchronized using hot standby.

[0177] Here, the new storage resources are operational and may be logged into the database at step 856, whereby the controller 200 can connect the new host to the new storage resources at step 858. When migrating from a single node by multiple virtual hosts, it may be necessary to repeat this process for multiple hosts (step 860). If they are tracked, the boot order may be determined by the controller logic using the application dependencies.

[0178] FIG. 8E shows another exemplary process flow for the expansion from a single node to multiple nodes in a system. At step 870, new resources are coupled to a single node system. The controller may have a set of system rules and / or expansion rules for the system (alternatively, the controller may derive expansion rules based on service execution, service templates, and service dependencies on each other). At step 872, the controller checks such rules used to facilitate the expansion.

[0179] If the new physical resources include storage resources, the storage resources may be moved at step 874 from a single node or other form of a simpler IT system (alternatively, the storage resources may be mirrored). If the storage resources are moved, after the storage resources are moved, the computing resources or the running resources may be reloaded or rebooted at step 876. In another example, the computing resources may be connected at step 876 to the mirrored storage resources and may remain running, while the old storage resources on the single node system or the hardware resources of the previous system may be disconnected or disabled. For example, a running service may be coupled to two mirrored block devices. That is, one may be on a single node server (e.g., using mdadm raid1), and the other may be coupled on the storage resources. When the data is synchronized, the drive on the single node server may be disconnected. The previous hardware may remain part of the IT system and may be run on the same node as the controller in a hybrid mode (step 878). The system may continue to repeat this migration process until the original node only runs the controller, where the system is then decentralized (step 880). Further, at each of the steps of the process flow of FIG. 8E, the controller can update the system state 220 and can log changes to the system in a database (step 882).

[0180] Referring to FIG. 9A, application 910 is installed on resource 900. Resource 900 may be a computing resource 310, a storage resource 410, or a networking resource 610 as described with respect to FIGS. 1-10 herein. Resource 900 may be a physical resource. The physical resource may include a physical machine or a physical IT system component. Resource 900 may be, for example, a physical computing resource, a physical storage resource, or a physical networking resource. Resource 900 may be coupled to controller 200 of system 100 by other resources among the computing resources, networking resources, or storage resources as described with respect to FIGS. 2A-10 of this specification.

[0181] Resource 900 may initially be powered off. Resource 900 may be coupled to the controller via a network, namely, an out-of-band management connection 260, an in-band management connection 270, a SAN 280, and / or a network 290. Resource 900 may also be coupled to one or more application networks 390 through which services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 may be coupled to an independent out-of-band management device 915 or circuit of resource 900 that turns on when resource 900 is plugged in. The device may enable features including, but not limited to, power on / off of the device, attachment to the console and input of commands, monitoring of temperature and other computer health-related elements, and setting of BIOS settings 195 and other features outside the scope of the operating system.

[0182] The controller 200 may detect the resource 900 through the out-of-band management network 260. The controller 200 may also identify the type of the resource and may identify its configuration using in-band management or out-of-band management. The controller logic 205 may be configured to search for additional hardware and examine the out-of-band management 260 or in-band management 270. When the resource 900 is detected, the controller logic 205 may use the global system rules 220 to determine whether the resource 900 should be configured automatically or through interaction with the user. If the resource 900 is added automatically, the setup follows the global system rules 210 within the controller 200. If the resource 900 is added by the user, the global system rules 210 within the controller 200 may request the user to confirm the addition of the resource and what the user wants to do with the computing resource. The controller 200 may query the API application(s) to confirm that the new resource has been approved, or in other cases, may request the user or any program controlling the stack. The approval process may also be automatically and securely completed using encryption to verify the legitimacy of the new resource. Next, the resource 900 is added to the IT system state 220 including the switch or network to which the resource 900 is connected.

[0183] The controller 200 may power on the resources through the out-of-band management network 260. The controller 200 may use the out-of-band management connection 260 to power on the physical resources and configure the BIOS 195. The controller 200 may automatically use the console 190 and may select the desired BIOS options, which may be achieved by the controller 200 reading the console image with image recognition and controlling the console 190 through out-of-band management. The boot-up state may be determined by image recognition through the console of the resource 900, out-of-band management with a virtual keyboard, querying the services operating on the resource, or querying the services of the application 910. Some applications may have a process that allows the controller 200 to monitor or, in some cases, change the settings of the application 910 using in-band management 270.

[0184] An application 910 on a physical resource 900 (or resources 300, 310, 311, 312, 313, 400, 410, 411, 412, 600, 610 as described with respect to FIGS. 1 - 10 of this specification) may boot via a SAN 280 or another network using a BIOS boot option or other means of configuring a remote boot such as enabling PXE boot or Flex boot. Further, or alternatively, the controller 200 may use out - of - band management 260 and / or in - band management connection 270 to instruct the physical resource 900 to boot an application image of image 950. The controller may configure the boot options on the resource or use an existing available remote boot method such as PXE boot or Flex boot. The controller 200 may optionally or alternatively use out - of - band management 260 to boot from an ISO image to configure the local disk and then instruct the resource to boot from the local disk(s) 920. The local disk(s) may be loaded with boot files. This may be achieved by using out - of - band management 260, image recognition, and a virtual keyboard. The resource may also have a boot file and / or a boot loader installed. The resource 900 and the application may use global system rules 210 and controller logic 205 to boot, for example, from an image 950 loaded from a template 230 via a SAN 280. The global system rules 220 may specify the boot order. For example, the global system rules 220 may require that the resource 900 be booted first and then the application 910 be booted. When the resource 900 is booted using the image 950, information received through the in - band management connection 270 regarding the resource 900 may be collected and added to the IT system state 220. The resource 900 is added to a storage resource pool and the resource 900 becomes a resource that is managed by the controller 200 and tracked in the IT system state 220.Application 910 may be booted in the order specified by global system rule 220 using application image 956 loaded on image 950 or resource 900.

[0185] Controller 200 may configure networking resource 610 to connect application 910 to application network 390 via out-of-band management connection 260 or another connection. Physical resource 900 may be connected to remote storage such as block storage resources including, but not limited to, ISER (ISCSI over RDMA), NVMEOF FCOE, FC, or ISCSI, or to another storage backend such as SWIFT, GLUSTER, or CEPHFS. When a service or application is launched and running, IT system state 220 may be updated using out-of-band management connection 260 and / or in-band management connection 270. Controller 200 may determine the power state of physical resource 900, i.e., on or off, using out-of-band management connection 260 or in-band management connection 270. Controller 200 may determine whether a service or application is running or in a boot-up state using out-of-band management connection 260 or in-band management connection 270. The controller may perform other actions based on the information it receives and global system rule 210.

[0186] FIG. 9B shows image 950 loaded directly or indirectly from template 230 to a compute node (e.g., via another resource or database) to boot application 910. Image 950 may include a custom kernel 941 for application 910.

[0187] Image 950 may include a boot file 940 for resource type and hardware. The boot file 940 may include a kernel 941 corresponding to the resources, applications, or services to be deployed. The boot file 940 may also include an initrd or similar file system used to assist the boot process. The boot system 940 may include multiple kernels or initrds configured for different hardware types and resource types. Further, image 450 may include a file system 951. The file system 951 may include a base image 952 and corresponding file system, a service image 953 and corresponding file system, and a volatile image 954 and corresponding file system, along with the corresponding file systems. The loaded file system and data may vary according to the resource type and application, or the service being executed. The base image 952 may include a base operating system file system. The base operating system may be read-only. The base image 952 may also include the basic tools of an operating system independent of what is being executed. The base image 952 may include a base directory and operating system tools. The service file system 953 may include configuration files and specifications for resources, applications, or services. The volatile file system 594 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may be configured as variables including, but not limited to, passwords, session keys, and private keys, or may not be configured. The file system may be mounted as a single file system using technologies such as overlayFS, enabling some read-only and some read-write file systems to reduce the amount of replicated data used for the application.

[0188] FIG. 9C shows an example of installing an application from an NT package that can be one type of template 230. At step 900.1, the controller determines that the package blob needs to be installed. At step 900.2, the controller creates storage resources on the default data store for the blob type (block, file, file system). At step 900.3, the controller connects to the storage resources via the storage transport available for the storage resource type. At step 900.4, the controller copies the package blob to the connected storage resources. Next, the controller disconnects from the storage resources (step 900.5) and sets the storage resources to read-only (step 900.6). Next, the package blob is successfully installed (step 900.7).

[0189] In another example, Appendix B attached hereto describes exemplary details regarding how the system connects computing resources to overlayfs. Using such techniques, it is possible to facilitate installing an application on a resource according to FIG. 9A or booting a computing resource from a storage resource according to step 205.11 of FIG. 2F.

[0190] FIG. 9D shows an application 910 deployed on a resource 900. The resource 900 may comprise virtual computing resources, for example, a hypervisor 920, one or more virtual machines 921, 922, and / or a computing node that may comprise containers. The resource 900 may be configured in a manner similar to that described herein with respect to FIGS. 1-10 using an image 950 loaded on the resource 900. In this example, the resource 920 is shown as a hypervisor that manages virtual machines 921, 922. The controller 200 may communicate with a resource 900 that hosts a hypervisor 920 to create a resource using in-band management 270, and may configure the resource to allocate appropriate hardware resources. Appropriate hardware resources may include, but are not limited to, a CPU, RAM, GPU, a remote GPU (which may use RDMA to remotely connect to another host), a network connection, a network fabric connection, and / or virtual and physical connections to a partitioned and / or segmented network. The controller 200 may control the resource 900 and the hypervisor 920 using a virtual console 190 (including, but not limited to, SPICE or VNC) and image recognition. Additionally or alternatively, the controller 200 may use out-of-band management 260 or an in-band management connection 270 to instruct the hypervisor 920 to boot an application image 950 from a template 230 using global system rules 210. The image 950 may be stored on the controller 200, or the controller 200 may move or copy them to a storage resource 410. Boot images for VMs 921, 922 may be stored locally, for example, as the image 950, or as a block device or a file on a remote host, and shared, for example, by file sharing such as NFS over RDMA / NFS using an image type such as qcow2 or raw, or a remote block device using ISCSI, ISER, NVMEOF, FC, FCOE may also be used.Part of the image 950 may be stored on the storage resource 410 or the compute node 310. The controller 200 may configure the networking resource 610 to appropriately support the application by using global rules and / or templates via the out-of-band management connection 260 or another connection. The application 910 on the resource 900 may use the image 950 loaded by the SAN 280 or another network using the BIOS boot option, or the hypervisor 920 on the resource 900 may boot via connecting to a block storage resource including but not limited to ISER (ISCSI over RDMA), NVMEOF FCOE, FC, or ISCSI, or another storage backend such as SWIFT, GLUSTER, or CEPHFS. The storage resource may be copied from a template target on the storage resource. The IT system state 220 may be updated by querying the hypervisor 920 for information. The in-band management connection 270 may communicate with the hypervisor 920 and may be used to determine the power state of the resource, i.e., on or off, or to determine the boot-up state. The hypervisor 920 may use the virtual in-band connection 923 to the virtualized application 910 or may use the hypervisor 920 for functions similar to out-of-band management. This information may indicate whether the service or application has started up and is running depending on whether the power is on or it has been booted.

[0191] The boot-up state may be determined by image recognition through the console 190 of the resource 900, out-of-band management 260 by a virtual keyboard, querying services operating on the resource, or querying services of the application 910 itself. Some applications may have a process that enables the controller 200 to monitor or, in some cases, change the settings of the application 910 using in-band management 270. Some applications may be on virtual resources, and the controller 200 may monitor by communicating with the hypervisor 920 using in-band management 270 (or out-of-band management 260). The application 910 may not have such a process for monitoring inputs and / or for additional processes (or such a process that may be switched to save resources). In such cases, the controller 200 may use the out-of-band management connection 260 and may use an image process and / or a virtual keyboard for logging on to the system to make changes and / or switch on the management process. Similar to virtual computing resources, the virtual machine console 190 may be used.

[0192] FIG. 9E shows an exemplary process flow for adding a virtual computing resource host to the IT system 100. At step 900.11, a host that can be used as a virtual computing resource is added to the system. The controller may configure the bare metal server according to the process flow of FIG. 15B (step 900.12), or the operating system may be preloaded, and / or the host may be pre-configured (step 900.13). Next, the resource is added to the system state 220 as a virtual computing resource pool (step 900.14), and the resource becomes accessible from the controller 200 via the API (step 900.15). The API is typically accessed through the in-band management connection 270. However, the in-band management connection 270 may be selectively enabled and / or disabled with a virtual keyboard. And the controller may use the out-of-band management connection 260 and the virtual keyboard and monitor to communicate through the out-of-band connection 260 (step 900.16). Here, at step 900.17, the controller becomes able to utilize the new resource as a virtual computing resource.

[0193] Exemplary multi-controller system Referring to FIG. 10, a system 100 is shown having computing resources 300, 310 described herein with respect to FIGS. 1-10 of the present specification, including a plurality of physical computing nodes 311, 312, 313, storage resources 400, 410 described herein in the form of a plurality of storage nodes 411, 412 and a JBOD 413, a plurality of controllers 200a, 200b configured as the controller 200 described herein, including components 205, 210, 220, 230 (FIGS. 1-9C), networking resources 600, 610 described herein, including a plurality of fabrics 611, 612, 613, and an application network 390.

[0194] FIG. 10 shows a possible arrangement of the components of the exemplary system 100, but does not limit the possible arrangements of the components of the system 100.

[0195] The user interface or application 110 communicates with the API application 120, and the API application 120 communicates with either or both of the controllers 200a or 200b. The controllers 200a, 200b may be coupled to the out-of-band management connection 260, the in-band management connection 270, the SAN 280 or the network in-band management connection 290. As described with reference to FIGS. 1-9C herein, the controllers 200a, 200b are coupled to the compute nodes 311, 312, 313, the storage 411, 412 including the JBOD 413, and the networking resources 610 via the connections 260, 270, 280 and optionally 290. The application network 390 is coupled to the compute nodes 311, 312, 313, the storage resources 411, 412, 413 and the networking resources 610.

[0196] Controllers 200a and 200b may operate in parallel. Either of controllers 200a or 200b may initially operate as master controller 200 as described with respect to FIGS. 1 through 9C of this specification. Controllers 200a, 200b (s) may be configured to power off the entire system 100 from a powered-off state. One of controllers 200a, 200b may further populate the system state 220 from an existing configuration by searching for other controllers through either out-of-band connection 260 or in-band connection 270. Either of controllers 200a, 200b may access or receive resource states and related information from a resource or other controllers through one or more connections 260, 270. A controller or other resource may update other controllers. Thus, when an additional controller is added to the system, this controller may be configured to return the system 100 to the system state 220. If a failure occurs in one of the controllers or the master controller, another controller may be designated as the master controller. The IT system state 220 may further be reconstructable from state information stored on available or resources. For example, an application may be deployed on a computing resource on which the application is configured to create a virtual computing resource in which the system state is stored or replicated. The global system rules 210, the system state 220, and the template 230 may also be stored or copied to a resource or combination of resources. Thus, when all controllers go offline and a new controller is added, the system may be configured so that the new controller can restore the system state 220.

[0197] The networking resource 610 may include a plurality of network fabrics. For example, as shown in FIG. 10, the plurality of network fabrics may include one or more of an SDN Ethernet switch 611, a ROCE switch 612, an Infiniband switch 613 or other switches, or a fabric 614. A hypervisor with virtual machines on a compute node may utilize one or more of the desired fabrics to connect to a physical switch or a virtual switch. The networking arrangement may permit restrictions on the physical network, for example, through segmented networking, for security or other resource optimization.

[0198] The system 100 through the controller 200 described in FIGS. 1-10 of this specification may automatically set up a service or an application. A user may request to set up a service for the system 100 through the user interface 110 or an application. The service may include, but is not limited to, an email service, a web service, a user management service, a network provider, LDAP, developer tools, VOIP, authentication tools, accounting software. The API application 120 converts a user or application request and sends a message to the controller 200. The service template or image 230 of the controller 200 is used to identify which resources are required for the service. The required resources are identified based on availability according to the system state 220. The controller 200 issues a request to the compute resource 310 or the compute nodes 311, 312 or 313 for the required computing service, issues a request to the storage resource 410 for the required storage resource, and issues a request to the networking resource 610 for the required networking resource. Next, the system state 220 is updated and the resources to be allocated are identified. Next, according to the service template, the service is installed on the allocated resources using the global system rules 210.

[0199] Improvement of System Security Referring to FIG. 13A, an IT system 100 is shown that includes a resource 1310, which can be bare metal or a physical resource. Although FIG. 13A shows only a single resource 1310 connected to the system 100, it should be understood that the system 100 can include multiple resources 1310. The resource(s) 1310 may be or include a bare metal cloud node. The bare metal cloud node may include resources connected to an external network 1380 that enable remote access to a physical host or virtual machine, enable the creation of virtual machines, and enable external users to execute code on the resource(s), but is not limited thereto. The resource(s) 1310 may be directly or indirectly connected to the external network 1380 or the application network 390. The external network 1380 may be the Internet or other resource(s) not managed by the controller 200 or the controller of the IT system 100. The external network 1380 may include the Internet, Internet connection(s), resource(s) not managed by the controller, other wide area networks (e.g., Stratcom, peer-to-peer mesh network, or other external networks that may or may not be publicly accessible), or other networks, but is not limited thereto.

[0200] When physical resource 1310 is added to IT system 100a, it may be coupled to controller 200 and powered off. Resource 1310 is coupled to controller 200a via one or more of out-of-band management (OOBM) connection 260, optionally in-band management (IBM) connection 270, and optionally SAN connection 280. As used herein, SAN 280 may or may not include a configured SAN. The configured SAN may include a SAN used for powering on or configuring the physical resource. The configured SAN may be part of SAN 280 or may be separate from SAN 280. In-band management may also include a configured SAN, which may or may not be SAN 280, as shown herein. When the resource is in use, the configured SAN may also be disabled, disconnected, or unavailable. OOBM connection 260 is not visible to the OS of system 100, but IBM connection 270 and / or the configured SAN may be visible to the OS of system 100. Controller 200 of FIG. 13A may be configured similarly to controller 200 described with reference to FIGS. 1-12B herein. Resource 1310 may include internal storage. In some configurations, controller 200 may populate the storage and may temporarily configure the resource to connect to the SAN to fetch data and / or information. The out-of-band management connection 260 may be coupled to an independent out-of-band management device 315 or to the circuitry of resource 1310 that powers on when resource 1310 is plugged in. Device 315 may enable features including, but not limited to, powering on / off the device, connecting to the console and entering commands, monitoring temperature and other computer health-related elements, and setting BIOS settings and other features outside the scope of the operating system. Controller 200 may reference resource 1310 through out-of-band management network 260. The controller may also identify the type of resource and may identify its configuration using in-band or out-of-band management.Figures 13C to 13E described below show various process flows for adding physical resource 1310 to IT system 100a and / or starting or managing system 100 so as to enhance system security.

[0201] As used herein with reference to a network, networking resource, network device, and / or networking interface, the term "disabled" refers to the action of such network, networking resource, network device, and / or networking interface being powered off (manually or automatically), physically disconnected, and / or disconnected from a network, virtual network (including, but not limited to, VLAN, VXLAN, InfiniBand partition) by virtual or some other means (e.g., by filtering). The term "disabled" also includes one-way or unidirectional restrictions on operability, such as making a resource unable to send or write data to a destination while still having the ability to receive or read data from a source, or making a resource unable to receive or read data from a source while still having the ability to send or write data to a destination. Such network, networking resource, network device, and / or networking interface may be disconnected from an additional network, virtual network, or combination of resources and may remain connected to a previously connected network, virtual network, or combination of resources. Further, such networking resource or device may be switched from one network, virtual network, or combination of resources to another.

[0202] As used herein with reference to a network, networking resource, network device, and / or networking interface, the term "active" refers to the action of such network, networking resource, network device, and / or networking interface being (manually or automatically) powered on, physically connected, and / or connected to a network, virtual network (including, but not limited to, VLAN, VXLAN, InfiniBand partitions), or other means in a virtual or other way. Such network, networking resource, network device, and / or networking interface may be connected to the joining of resources if it is already connected to another system component, an additional network, or a virtual network. Further, such a networking resource or device can be switched from one network, virtual network, or joining of resources to another. The term "active" also includes one-way or unidirectional restrictions on operability, such as enabling a resource to send, write, or receive data to / from a destination (while having the ability to restrict data from the source), and enabling a resource to send, receive, or read data from a source (while having the ability to restrict data from the destination).

[0203] The controller logic 205 is configured to search for added hardware and examine the out-of-band management connection 260 or the in-band management connection 270 and / or the configured SAN 280. If the resource 1310 is detected, the controller logic 205 may determine whether to automatically configure the resource using the global system rules 220 or by interacting with the user. If added automatically, the setup follows the global system rules 210 within the controller 200. If added by the user, the global system rules 210 within the controller 200 may prompt the user to confirm the addition of the resource and what the user wants to do with the resource 1310. The controller 200 may query the API application(s) to confirm that the new resource has been approved, or, in other cases, request it from the user or any program controlling the stack. The approval process may also be automatically and securely completed using encryption to verify the legitimacy of the new resource. Next, the controller logic 205 adds the resource 1310 to the IT system state 220, including the switch or network into which the resource 1310 is plugged.

[0204] If the resource is physical, the controller 200 may turn on the power of the resource through the out-of-band management network 260, and the resource 1310 may boot from the image 350 loaded from the template 230, for example, via the SAN 280, using the global system rules 210 and the controller logic 205. The image may be loaded indirectly through other network connections or via another resource. Once booted, information about the resource 1310 may also be collected and added to the IT system state 220. This may be done through the in-band management connection and / or the configured SAN connection or the out-of-band management connection. The resource 1310 may boot from the image 350 loaded from the template 230, for example, via the SAN 280, using the global system rules 210 and the controller logic 205. The image may be loaded indirectly through other network connections or via another resource. Once booted, information received through the in-band management connection 270 regarding the computing resource 310 may also be collected and added to the IT system state 220. Next, the resource 1310 may be added to the storage resource pool, which becomes a resource managed by the controller 200 and tracked in the IT system state 220.

[0205] In-band management and / or configuration SAN may be used by controller 200 to set up, manage, use, or communicate with resource 1310 and to execute any commands or tasks. However, optionally, the in-band management connection 270 may be configured by controller 200 to be off or disabled at any time, or during the setup, management, use, or operation of system 100 or controller 200. In-band management may be further configured to be on or enabled at any time, or during the setup, management, use, or operation of system 100 or controller 200. Optionally, controller 200 may controllably or switchably disconnect resource 1310 from controller 200(s) from the in-band management connection 270. Such disconnection or disconnection possibility may be physical, for example, using an automatic physical switch, or a switch that turns off the power to the in-band management connection of the resource to the network and / or the configuration SAN. For example, the disconnection may be performed by the network switch cutting off the power to the port to which the in-band management 270 of resource 1310 and / or the configuration SAN 280 is connected. Such disconnection or partial disconnection may be performed using software-defined networking, or may be physically filtered for the controller using software-defined networking. Such disconnection may be performed via the controller, either through in-band management or out-of-band management. According to an exemplary embodiment, in response to a selective control command from controller 200, resource 1310 may be disconnected from the in-band management connection 270 at any point before, during, or after resource 1310 is added to the IT system.

[0206] Using software-defined networking, the in-band management connection 270 and / or the configuration SAN 280 may or may not retain some functions. The in-band management 270 and / or the configuration SAN 280 may be used as a restricted connection for communication between the controller 200 or other resources. The connection 270 may be restricted to prevent an attacker from pivoting to the controller 200, other networks, or other resources. The system may be configured to prevent devices such as the controller 200 and the resource 1310 from communicating openly and to prevent the resource 1310 from being exposed to risks. For example, in the in-band management 270 and / or the configuration SAN 280, only data transmission of the in-band management and / or the configuration SAN may be enabled through software-defined networking or a hardware change method (such as electronic restrictions), and all reception may be prohibited. The in-band management and / or the configuration SAN may be configured to be a unidirectional write component either physically or using software-defined networking that only allows writing from the controller to the resource, or as a unidirectional write connection from the controller 200 to the resource 1310. The unidirectional write nature of the connection may further be controlled or turned on or off according to the desired security situation and various stages or times of the system operation. The system may further be configured to restrict writing or communication from the resource to the controller, for example, to communication of logs or alerts. The interface may further be able to move to other networks, add to or remove from a network by technologies including but not limited to software-defined networking, VLAN, VXLAN, and / or InfiniBand partitioning. For example, the interface can be connected to a setup network, removed from that network, and moved to the network used at runtime. Communication from the controller to the resource may be disconnected or restricted, and as a result, the controller may physically not be able to respond to any data transmitted from the resource 1310.According to one example, when resource 1310 is added and booted, in-band management 270 may either switch or filter the switch, either physically or using software-defined networking. The in-band management may be configured to be able to send data to another resource dedicated to log management.

[0207] The in-band management can be switched on and off using out-of-band management or software-defined networking. When the in-band management is disconnected, there is no need to run the daemon, and the in-band management can be re-enabled using the keyboard function.

[0208] Furthermore, optionally, resource 1310 may not have an in-band management connection, and the resource may be managed through out-of-band management.

[0209] Out-of-band management may be used, alternatively or additionally, to operate various aspects of the system in a way that includes, but is not limited to, for example, a keyboard, virtual keyboard, disk mount console, virtual disk connection, BIOS setting changes, boot parameter and other system aspect changes, execution of existing scripts that may be present on a bootable image or installation CD, or other features of out-of-band management that enable communication between controller 200 and resource 1310 regardless of the availability of the operating system running on resource 1310. For example, controller 200 may send commands using such tools via out-of-band management 260. Controller 200 may further use image recognition to assist in controlling resource 1310. Thus, using the out-of-band management connection, the system may prevent or avoid unwanted operations of resources connected to the system via the out-of-band management connection. The out-of-band management connection may also be configured as a one-way communication system during the operation of the system or at selected times during the operation of the system.

[0210] Furthermore, the out-of-band management connection 260 may be selectively controlled by the controller 200 in the same manner as the in-band management connection, if the implementer desires.

[0211] The controller 200 can automatically turn resources on and off and update the state of the IT system according to global system rules for reasons determined by the IT system user, such as turning resources off to save power, turning resources on to improve application performance, or any other reason the IT system user may have. The controller can further turn the configuration SAN, in-band and out-of-band management connections on and off, or specify such connections as one-way write connections, for various security purposes, at any time during system operation (e.g., disabling the in-band management connection 270 or the configuration SAN 280 while the resource 1310 is connected to the external network 1380 or the internal network 390). One-way in-band management may further be used, for example, to monitor the health of the system and monitor logs and information that may appear to the operating system.

[0212] The resource 1310 may further be coupled to one or more internal networks 390, such as an application network through which services, application users, and / or clients can communicate with each other. Such an application network 390 may further be connected to, or connectable to, the external network 1380. According to exemplary embodiments of the present specification, including but not limited to FIGS. 2A - 12B, in-band management may be disconnected from, or disconnectable from, the resource or the application network 390, or provide a one-way write from the controller, to provide additional security when the resource or the application network is connected to an external network, or when the resource is connected to an application network not connected to an external network.

[0213] The IT system 100 of FIG. 13A can be configured in the same way as the IT system 100 shown in FIG. 3B. The image 350 can be loaded directly or indirectly (through another resource or database) from the template 230 to the resource 1310 for booting the computing resources and / or loading the applications. The image 350 can include a boot file 340 for the resource type and hardware. The boot file 340 may include a kernel 341 corresponding to the resource, application, or service to be deployed. The boot file 340 may further include an initrd or a similar file system used to assist the boot process. The boot system 340 can include a plurality of kernels or initrds configured for different hardware types and resource types. Further, the image 350 can include a file system 351. The file system 351 can include a base image 352 and corresponding file systems, a service image 353 and corresponding file systems, and a volatile image 354 and corresponding file systems. The file systems and data to be loaded may vary depending on the resource type and the application or service to be executed. The base image 352 may include a base operating system file system. The base operating system may be read-only. The base image 352 may further include basic tools for an operating system independent of the one in execution. The base image 352 can include a base directory and operating system tools. The service file system 353 may include configuration files and specifications for the resource, application, or service. The volatile file system 354 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may be configured as variables including, but not limited to, passwords, session keys, and private keys, or may not be configured.The file system may be mounted as a single unified file system using technologies such as overlayFS, enabling some read-only and some read-write file systems to reduce the amount of replicated data used for applications.

[0214] FIG. 13B shows a plurality of resources 1310, each resource comprising one or more hypervisors 1311 that host or include one or more virtual machines. Controller 200a is coupled to resources 1310, each of which comprises bare metal resources. As shown in and described with reference to FIG. 13B, resources 1310 are each coupled to controller 200a. According to an exemplary embodiment of the present specification, the in-band management connection 270, the configuration SAN 280, and / or the out-of-band management connection 260 may be configured as described with respect to FIG. 13A. One or more of the virtual machines or hypervisors may be at risk or may be at risk of being in a compromised state. In a conventional system, if so, other virtual machines on other hypervisors may be at risk of being in a compromised state. For example, this can result from the exploitation of a hypervisor running within a virtual machine. For example, the pivot may move from a compromised hypervisor to controller 200a and then from the compromised controller 200a to other hypervisors coupled to controller 200a. For example, a pivot may occur between a compromised hypervisor and a target hypervisor using a network connected to both. In the configuration of the in-band management 270, the configuration SAN 280, or the out-of-band management 260 of controller 200a and resource 1310 shown in FIG. 13B, some or all may be selectively controlled to disable in-band (or configuration SAN) and / or out-of-band connections on a given link between controller 200a and resource 1310, which may prevent it from being used for a compromised virtual machine to escape from one hypervisor and pivot to other resources.

[0215] As described above with respect to FIGS. 1 to 12, the in-band management connection 270 and the out-of-band management connection 260 may be further configured as described with respect to FIGS. 13A and 13B.

[0216] FIG. 13C shows an exemplary process flow for adding or managing physical resources such as bare metal nodes to system 100. The resources 1310 shown in FIGS. 13A and 13B of this specification or shown with respect to FIGS. 1 to 12 may be connected via the out-of-band management connection 260 and the in-band management connection 270 to the controller of system 100 and / or via the SAN.

[0217] After an instance of the resource connection, at step 1370, the external network and / or the application network is disabled. As described above, any of a variety of techniques can be used for this disabling. For example, before system setup, resource addition, system testing, system update, or execution of other tasks or commands, as described with respect to FIGS. 13A and 13B, use the in-band management connection or the configured SAN to disable, disconnect, or filter components of system 100 (or only those vulnerable to attacks) from the external network or the application network.

[0218] After step 1370, then in step 1371, the in-band management connection and / or the configured SAN is enabled. Thus, the combination of steps 1370 and 1371 separates resources from the external network and / or the application network while keeping the in-band management and / or the SAN connection operational. Next, commands can be executed on the resources under the control of controller 200 via the in-band management connection (see step 1372). For example, setup and configuration steps including but not limited to those described herein with respect to FIGS. 1-13B may be executed in step 1372 using the in-band management and / or the configured SAN. Alternatively or additionally, using the in-band management and / or the configured SAN, in step 1372, operations, updates or management of the system (including but not limited to any change management or system updates), testing, updates, data transfer, collection of information regarding performance and health (including but not limited to errors, CPU usage, network usage, file system information and storage usage), and collection of logs, as well as other tasks including but not limited to other commands that may be used to manage system 100 described in FIGS. 1 through 13B herein may be executed.

[0219] As described herein with respect to FIGS. 13A and 13B, after the addition of resources, system setup, and execution of such tasks or commands, the in-band management connection 270 and / or the configuration SAN 280 between the resource and the controller or other components of the system may be disabled in one or more directions at step 1373. Such disabling may employ disconnection, filtering, etc., as described above. After step 1373, at step 1374, the connection to the external network and / or the application network may be restored. For example, the controller may notify the networking resources so that the resource 1310 can connect to the application network or the Internet. The same steps may be followed when testing or updating the system, i.e., disconnecting or filtering the in-band management connection to the external network and / or the application network, and then the in-band management connection to the resource may be enabled or connected (in one or both directions). Thus, while the resource is connected to the external network and / or the application network, steps 1373 and 1374 operate simultaneously to separate the resource from the connection to the controller through the in-band management connection and / or the configuration SAN.

[0220] Out-of-band management may be used for the management of a system or resource, the setup of a system or resource, or the configuration, boot, or addition of a system or resource. When used in any of the embodiments of this specification, out-of-band management may send commands to a machine using a virtual keyboard to change settings before booting, and may further send commands to an operating system by entering them on the virtual keyboard. If the machine is not logged in, out-of-band management may use the virtual keyboard to enter a username and password, use image recognition to verify the logon, verify the entered commands, and check whether they were executed. If there is only a graphical console for the physical resource, a virtual mouse may also be used and out-of-band management may make changes by image recognition.

[0221] Figure 13D is another exemplary process flow for adding or managing physical resources such as bare metal nodes to system 100. At step 1380, the resources shown in FIGS. 13A and 13B or FIGS. 1 - 12 herein may be connected to the system or resource via out - of - band management 260. By providing access to a disk image (e.g., an ISO image) through out - of - band management facilitated by a controller, the disk may be virtually connected (see step 1381). Next, the resource or system may be booted from the disk image (step 1382), and then files may be copied from the disk image to a bootable disk (see step 1383). This may also be used for booting a system in which resources are set up in this way using out - of - band management. This may also be used to configure and / or boot multiple resources (including but not limited to networking resources) that may be coupled together, whether or not the multiple resources further comprise a controller or constitute a system. Thus, using a virtual disk, the controller may be enabled to connect a disk image to a resource as if the virtual disk were connected to the resource. Out - of - band management may further be used to send files to the resource. Data may be copied from the virtual disk to a local disk at step 1383. The disk image may contain files that the resource can copy and use during operation. The files may be copied or used either through a scheduled program or instructions from out - of - band management. The controller may log on to the resource using a virtual keyboard through out - of - band management, enter commands, and copy files from the virtual disk to the controller's own disk or other storage accessible to the resource. At step 1384, the system or resource is configured to boot by setting the BIOS, EFI, or boot order settings, thereby booting from the bootable disk.In the boot configuration, the EFI manager of the operating system such as efibootmgr may be used, which may be executed directly from out-of-band management or included in the installer script (for example, a script using efibootmgr is automatically executed when the resource boots). Further, boot options or other BIOS changes may be set through an out-of-band management tool such as the Supermicro Boot Manager, which uses either the boot order command or uploads the BIOS configuration (such as the XML BIOS configuration supported by the Supermicro Update Manager). The BIOS may further be configured to set appropriate BIOS settings including the boot order using image recognition from the keyboard and console. The installer may be executed against the loaded pre-configured image. The configuration may be tested by looking at the screen and using image recognition. After configuration, the resources can be enabled (for example, power on, boot, connect to the application network, or combinations thereof) (step 1385).

[0222] Figure 13E is another exemplary process flow for adding or managing physical resources such as bare metal nodes to system 100, in which case PXE, Flexboot or similar network boot is used. At step 1390, the resources 1310 shown in FIGS. 13A and 13B herein, or shown with respect to FIGS. 1 - 12, can be connected to the controller of system 100 via (1) an in - band management connection 270 and / or a SAN, and (2) an out - of - band management connection 260. Next, the external network and / or application network connections may be disabled (e.g., physically or virtually, wholly or in part, filtered or severed using SDN) at step 1391 (similar to that described above with respect to step 1370). For example, before system setup, resource addition, system testing, system updates or execution of other tasks or commands, as described with respect to FIGS. 13A and 13B, the in - band management connection or SAN is used to disable, sever or filter components of system 100 (or only those vulnerable to attack) from the external network or application network.

[0223] In step 1392, the type of the resource is determined. For example, information about the resource can be collected from the mac address by using an out-of-band management tool or by temporarily booting an operating system having a tool that can be used to identify the resource information by connecting a disk image (e.g., an ISO image) to the resource as if the disk were connected to the resource. Next, in step 1393, the resource is identified as being configured or pre-configured for PXE or flexboot or the like. Next, in step 1394, the power of the resource is turned on and PXE, Flexboot or a similar boot is performed (or if the resource is temporarily booted and then powered on again). Next, in step 1395, the resource boots from an in-band management connection or a SAN. In step 1396, the data is copied to a disk accessible by the resource in a manner similar to that described with reference to step 1383 of FIG. 13D. In step 1397, the resource is configured to boot from the disk(s) in a manner similar to that described above with respect to step 1384 of FIG. 13D. If the resource is identified as being pre-configured for PXE, flexboot, etc., the file may be copied at any step from 1393 to 1396. If in-band management is enabled, it may be disabled in step 1398 and the application network or external network may be reconnected or enabled in step 1399.

[0224] Furthermore, it should be understood that technologies other than OOBM can be used to remotely enable resources (such as turning on the power) and verify that the resources have been booted. For example, the system can prompt the user to press the power button and manually communicate to the controller that the system has booted (or use a console connection to the keyboard / controller). Additionally, the system can ping the controller through IBM (e.g., via ssh, telnet, or another method on the network) when the system has booted, the controller has logged on, and is instructed to reboot. For example, the controller can send a reboot command using ssh. If PXE is being used and there is no OOBM, in any case, the system should have a way to instruct the resources to power on automatically or instruct the user to manually power on.

[0225] Deployment of the Controller and / or Environment In an exemplary embodiment, the controller may be deployed into the system from an original controller 200 (such an original controller 200 can be referred to as the "main controller"). Thus, the main controller can set up a system or environment that can be a separate or separable IT system or environment.

[0226] The environments described in this specification refer to a set of resources within a computer system that can interoperate with each other. The computer system may include multiple environments therein, although this is not essential. The resources of an environment may include one or more instances, applications, or sub-applications that are executed in that environment. Further, an environment may include one or more environments or sub-environments. An environment may or may not include a controller, and an environment may operate one or more applications. Such resources of an environment may include, for example, networking resources, computing resources, storage resources, and / or application networks used to execute a particular environment that includes an application within the environment. Therefore, it should be understood that an environment may provide the functions of one or more applications. In some examples, the environments described in this specification may be physically or virtually separated from, or separable from, other environments. Further, in other examples, an environment may have a network connection to other environments, and such a connection may be disabled or enabled as needed.

[0227] Further, the main controller may set up, deploy, and / or manage one or more additional controllers in various environments or as a separate system. Such additional controllers may be independent of or become independent from the main controller. Such additional controllers, even if independent of or pseudo-independent from the main controller, may receive commands from or send information to the main controller (or a separate monitor or environment via a monitoring application) at various points during operation. Environments may be configured for security (e.g., by making environments separable from each other and / or from the main controller) and / or for various administrative purposes. One environment may be connected to an external network, while another related environment may or may not be connected to the external network.

[0228] The main controller may manage an environment or application, regardless of whether the environment or application is a separate system and whether they include a controller or a sub - controller. The main controller may also manage a global configuration file or other shared storage of data. The main controller may further analyze global system rules (e.g., system rule 210), or subsets thereof, for different controllers according to their functions. Each new controller (which can be called a "sub - controller") may receive new configuration rules that can be a subset of the main controller's configuration rules. The subset of global configuration rules deployed to a controller may depend on or correspond to the type of IT system being set up. The main controller may set up or deploy new controllers or separate IT systems, which are then permanently separated from the main controller, for example, for shipping or delivery or otherwise. Global configuration rules (or subsets thereof) may define a framework for setting up applications or sub - applications in various environments and how they can interact with each other. Such applications or environments may run on sub - controllers that include a subset of the global configuration rules deployed by the main controller. In some examples, such applications or environments can be managed by the main controller. However, in other examples, such applications or environments are not managed by the main controller. When a new controller is generated from the main controller to manage an application or environment, an application dependency check can be performed across multiple applications to facilitate control by the new controller.

[0229] Thus, in an exemplary embodiment, the system may include a main controller configured to deploy another controller, or an IT system including such other controllers. Such an implementation system may be configured to be completely disconnected from the main controller. Once independent, such a system may be configured to operate as a stand-alone system, or may be controlled or monitored by another controller (or an environment with applications), such as the main controller, at various discrete or continuous times during operation.

[0230] FIG. 14A shows an exemplary system in which a main controller 1401 has controllers 1401a and 1401b respectively deployed on different systems 1400a and 1400b (where 1400a and 1400b may be referred to as subsystems. However, it should be understood that subsystems 1400a and 1400b may also function as environments). The main controller 1401 can be configured in the same manner as the controller 200 described above. Thus, it may include controller logic 205, global system rules 210, system state 220, and template 230.

[0231] Systems 1400a and 1400b each include controllers 1401a and 1401b respectively coupled to resources 1420a and 1420b. The main controller 1401 may be coupled to one or more other controllers, such as controller 1401a of subsystem 1400a and controller 1401b of subsystem 1400b. The global rules 210 of the main controller 1400 may include rules that can manage and control other controllers. The main controller 1401 may use such global rules 210 together with the controller logic 205, system state 220, and template 230 to set up, provision, and deploy subsystems 1400a and 1400b through controllers 1401a and 1401b in the same manner as described with reference to FIGS. 1-13E herein.

[0232] For example, the main controller 1401 may load the global rule 210 (or a subset thereof) as rules 1410a and 1410b into the subsystems 1400a and 1400b, respectively, such that the global rule 210 (or a subset thereof) instructs the operations of the controllers 1401a, 1401b and their subsystems 1400a, 1400b. Each of the controllers 1401a, 1401b may have rules 1410a, 1410b which may be the same subset or different subsets of the global rule 210. For example, which subset of the global rule 210 is provisioned to a given subsystem may depend on the type of the subsystem being deployed. Further, the controller 1401 may load or instruct the loading of data to be loaded into the system resources 1420a, 1420b or the controllers 1401a, 1401b.

[0233] The main controller 1401 may be connected to the other controllers 1401a, 1401b through the in-band management connection 270 (s), and / or the out-of-band management connection 260 (s) or the SAN connection 280 which may be enabled or disabled at various stages of deployment or management in a manner as described herein with reference to the deployment and management of resources described in FIGS. 13A - 13E. By using the selective enabling and disabling of the in-band management connection 270 or the out-of-band management connection 260, the subsystems 1400a, 1400b may be deployed in a manner such that the subsystems 1400a, 1400b have no (or have limited, controlled or prohibited) knowledge of the main system 100 or the controller 1401 or of each other at various times.

[0234] In an exemplary embodiment, the main controller 1401 may operate a centralized IT system having local controllers 1401a, 1401b deployed and configured by the main controller 1401 such that the main controller 1401 can deploy and / or execute a plurality of IT systems. Such IT systems may or may not be independent of each other. The main controller 1401 may be separated from the IT systems it creates or set up monitoring as a separate application that is air-gapped. A separate console for monitoring may comprise connections between the main controller and the local controller(s) and / or connections between environments that can be selectively enabled or disabled. The controller 1401 may deploy separate systems for various applications including, but not limited to, for example, business, manufacturing systems with data storage, data centers, and various other functional nodes, each having different controllers when each is stopped or put at risk. Such separation may be complete or permanent, or may be pseudo-separated, for example, depending on time, tasks, communication direction, or other parameters, temporarily or otherwise. For example, the main controller 1401 may be configured to provide the system with instructions that may or may not be limited to some defined situations, while the subsystem may have a limited or no ability to communicate with the main controller. Thus, such a subsystem may not be able to expose the main controller 1401 to risks. The main controller 1401 and the sub-controllers 1401a, 1401b may be separated from each other as described herein (in a specific example described below) by, for example, disabling in-band management 270, by one-way writing, and / or by restricting communication to out-of-band management 260. For example, in the event of an intrusion, one or more controllers may disable the in-band management connection 270 to one or more other controllers to prevent the spread of the intrusion or access.The system section can be turned off or separated.

[0235] Subsystems 1400a, 1400b may further share resources with or be connected to another environment or system through in-band management 270 or out-of-band management 260.

[0236] Figures 14B and 14C are exemplary flows showing possible steps for provisioning a controller with a main controller.

[0237] In FIG. 14B, at step 1460, the main controller provisions or sets up resources such as resource 1420a or 1420b. At step 1461, the main controller provisions or sets up the sub - controller. The main controller can use the techniques described above to set up resources within the system and execute steps 1460 and 1461. Further, although FIG. 14B shows step 1460 being executed before step 1461, it should be understood that this is not essential. The main controller 1401 can use its system rules 210 to determine which resources are needed and place the resources on the system or network. At step 1461, the main controller can set up or deploy the sub - controller by loading the system rules 210 into the system (or by providing instructions to the sub - controller on how to set up and obtain its own system rules). These instructions can include, but are not limited to, resource configuration, application configuration, global system rules for creating an IT system to be executed by the sub - controller, instructions to reconnect to the main controller to collect new or changed rules, and instructions to disconnect from the application network to create room for a new production environment. After provisioning the resources, at step 1463, the main controller can allocate resources to the sub - controller via an update to the system rules 210 and / or the system state 220.

[0238] FIG. 14C shows an alternative process flow for deployment. In the example of FIG. 14C, the main controller deploys the sub - controller at step 1470 (which can proceed as described for step 1461). Next, at step 1475, the sub - controller provisions resources using techniques such as those shown in FIGS. 3C and 7B.

[0239] FIG. 15A shows an exemplary system in which the main controller 1501 of the system 100 generates environments 1502, 1503, and 1504. Environment 1502 includes resource 1522, environment 1503 includes resource 1523, and environment 1504 includes resource 1524. Further, environments 1502, 1503, 1504 may share access to a pool of shared resources 1525. Such shared resources may include, but are not limited to, for example, a shared dataset, an API, or running applications that need to communicate with each other.

[0240] In the example of FIG. 15A, each environment 1502, 1503, 1504 shares the main controller 1501. The global system rules 210 of the main controller 1501 may include rules for deploying and managing the environments. Resources 1522, 1523, and / or 1524 may be required by their respective environments 1501, 1502, 1503 to manage one or more applications. Configuration rules for such applications may be implemented by the main controller (or, if present, a local controller within the environment) to define how each such environment operates and how it interacts with other applications and environments. The main controller 1401 may use the global rules 210 along with the controller logic 205, the system state 220, and the template 230 to set up, provision, and deploy the environments in a manner similar to the deployment of the resources and systems described with reference to FIGS. 1 - 14C herein. If the environment includes a local controller, the main controller 1501 may load the global rules 210 (or a subset thereof) into the local controller or associated storage such that the global rules (or a subset thereof) define the operation of that environment.

[0241] The controller 1501 may deploy and configure the respective resources 1522, 1523, 1524 of the environments 1502, 1503, 1504 and / or the shared resource 1525 using configuration rules having system rules 210. The controller 1501 may further monitor the environments or configure the resources 1522, 1523, 1524 (or the shared resource 1525) to enable monitoring of each of the environments 1502, 1503, 1504. Such monitoring may be via a connection to a separate monitoring console that can be enabled or disabled, or through the main controller. The main controller 1501 may be connected to one or more of the environments 1502, 1503, 1504 through in-band management connections 270 (multiple possible) and / or out-of-band management connections 260 (multiple possible) or SAN connections 280 that can be enabled or disabled in various stages of deployment or management in the manner described herein with reference to the deployment and management of the resources of FIGS. 13A - 13E and FIG. 14A. Using the enabling and disabling of the in-band management connection 270 or the out-of-band management connection 260 or the SAN connection 280, the environments 1502, 1503, 1504 may be deployed in such a way that they have no knowledge, limited knowledge, controlled knowledge of the main system 100 or the controller 1501 at various times, or have no knowledge, limited knowledge, controlled knowledge of their connections to each other.

[0242] The environment may be coupled to an external network 1580 that connects to an external environment that is external, or may comprise one or more resources that interact with other resources. The environment may be physical or non-physical. In this context, "non-physical" means that the environments share the same physical host(s) but are virtually separated from each other. The environments and systems may be deployed on the same hardware, similar but different hardware, or non-identical hardware. In some examples, the environments 1502, 1503, 1504 may be active copies of each other, while in other examples, the environments 1502, 1503, 1504 may provide different functionality from each other. As an example, the resources of the environment may be servers.

[0243] In accordance with the techniques described herein, separating systems and resources into separate environments or subsystems can enable the separation of applications for security and / or performance. Separating the environments can also mitigate the impact of resources exposed to risks. For example, one environment may contain sensitive data and may be configured to reduce exposure to the Internet, while another environment may host Internet-facing applications.

[0244] FIG. 15B shows an exemplary process flow in which the controller shown in FIG. 15A sets up an environment. In such an example, the system may be tasked with creating and setting up a new environment. This may be triggered by a user request or by system rules that are executed when engaging in a particular task or series of tasks. As described below, FIGS. 17A-18B show examples of specific change management tasks or series of tasks in which the system creates a new environment. However, there may be numerous situations in which a controller can create and set up a new environment.

[0245] Thus, referring to FIG. 15B, when setting up a new environment, the controller selects an environment rule (step 1500.1). In accordance with the environment rule, using the global system rule 210 and the template 230, the controller finds resources for the environment (step 1500.2). The rule may have a hierarchy of suitable resource selections that the controller goes through until the resources required for the environment are found. At step 1500.3, the controller allocates the resources found at step 1500.2 to the environment using, for example, the techniques described in FIG. 3C or FIG. 7B. Next, the controller configures the system's networking resources with respect to the new environment to ensure an efficient and compatible connection between the new environment and other system components (step 1500.4). The system state is updated at step 1500.5 each time each resource becomes active and each template is processed. Next, the controller sets up and enables the integration and interoperability of the environment's resources and turns on the power of any applications to deploy the new environment (step 1500.6). The system state is updated again at step 1500.7 when the environment becomes available.

[0246] FIG. 15C shows an exemplary process flow in which the controller shown in FIG. 15A sets up multiple environments. When setting up multiple environments, for each environment, the techniques described in FIG. 15B may be used to set up the environments in parallel. However, it should be understood that the environments can be set up in a sequential order or continuously, as described in FIG. 15C. Referring to FIG. 15C, at step 1500.10, the controller sets up and deploys a first new environment (which can be executed as described with respect to step 1500.1 in FIG. 15B). Different types of environments and different ways of interoperating between environments may have different environment rules. At step 1500.11, the controller selects the environment rules for the next environment. At step 1500.12, the controller finds resources according to the priority that can be defined by system rule 210. At step 1500.13, the controller allocates the resources found at step 1500.12 to the next environment. The environments may or may not share resources. At step 1500.14, the controller uses system rule 210 to configure the system's networking resources for the next environment and between environments having dependencies. The system state is updated at step 1500.15 each time each resource is valid, the template is processed, and the networking resources are configured including the dependencies of the environments. Next, the controller sets up and enables the integration and interoperability of the next environment and the resources between the environments, and turns on the power of any application to deploy the new environment (step 1500.16). The system state is updated at step 1500.17 when the next environment becomes available.

[0247] Unidirectional communication to support monitoring FIG. 16A shows an exemplary embodiment in which a first controller 1601 operates as a main controller for setting up one or more controllers such as 1601a, 1601b, and / or 1601c. The main controller 1601 may use the techniques described above with respect to controllers such as controller 200 / 1401 / 1501 to generate a plurality of cloud hosts, systems, and / or applications as environments 1602, 1603, 1604 that may or may not be operationally dependent on each other. As shown in FIG. 16A, an IT system, environment, cloud, and / or any combination thereof may be generated as environments 1602, 1603, 1604. Environment 1602 includes a second controller 1601a, environment 1603 includes a third controller 1601b, and environment 1604 includes a fourth controller 1601c. Environments 1602, 1603, 1604 may each include one or more resources 1642, 1643, 1644, respectively. The resources may include one or more applications 1642, 1643, 1644 that may be running thereon. These applications may connect to the allocated resources, whether or not they are shared. These or other applications may execute on the Internet or on one or more shared resources within a pool 1660 that may include a shared application or application network. The applications may provide services to one or more of a user or an environment or cloud. Environments 1602, 1603, 1604 may share resources or databases and / or may include or use resources within a pool 1660 specifically allocated to a particular environment. The various components of the system including the main controller 1601 and / or one or more environments may further be connectable to an external network 1615 such as an application network or the Internet.

[0248] There may be a connection configured to be selectively enabled and / or disabled in the manner described with respect to FIGS. 13A - 13E herein between any resource, environment, or controller and another resource, environment, controller, or external connection. For example, any resource, controller, environment, or external connection may be disabled or disconnected from the controller 1601, environment 1602, environment 1603, and / or environment 1604, resource, or application via an in - band management connection 270, an out - of - band management connection 270, or a SAN connection 280, or by physically disconnecting. As an example, the in - band management connection 270 between the controller 1601 and any of the environments 1602, 1603, 1604 may be disabled to protect the controller 1601. As another example, such in - band management connection(s) 270 may be selectively disabled or enabled during the operation of the environments 1602, 1603, 1604. In addition to the security purposes described with respect to FIGS. 13A - 13E herein, disabling or disconnecting the main controller 1601 from the environments 1602, 1603, 1604 may allow the main controller 1601 to generate the environments 1602, 1603, 1604 as a cloud, and the environments 1602, 1603, 1604 may then be separated from the main controller 1601 or other clouds or environments. In this sense, the controller 1601 is configured to generate multiple clouds, hosts, or systems.

[0249] Using the disabling or disconnecting elements described herein, a user may be enabled to have limited access to an environment through the main controller 1601 for a particular purpose. For example, a developer may be provided access to a development environment. As another example, an application administrator may be limited to a particular application or application network. As another example, logs may be made visible through the main controller 1601 to collect data without exposing the main controller to danger by an environment or controller generated by the main controller.

[0250] After the main controller 1601 sets up the environment 1602, the environment 1602 is disconnected from the main controller 1601, and thus, the environment 1602 may operate independently of the main controller 1601 and / or may be selectively monitored and maintained by the main controller 1601 or by other applications associated with or executed by the environment 1602.

[0251] An environment such as environment 1602 may be coupled to a user interface or console 1640 that enables a purchaser or user to access the environment 1602. The environment 1602 may host a user console as an application. The environment 1602 may be remotely accessed by a user. Each of the environments 1602, 1603, 1604 may be accessed by a common or separate user interface or console.

[0252] FIG. 16B shows an exemplary system in which the environments 1602, 1603, 1604 may be configured to write to another environment 1641, where, for example, a console (which may be any console that may be connected either directly or indirectly to the environment 1641) may be used to view the logs. In this way, the environment 1641 can function as a log server to which one or more of the environments 1602, 1603, 1604 write events. Next, the main controller 1601 can access the log server 1641 to monitor events on the environments 1602, 1603, 1604 without maintaining a direct connection to such environments 1602, 1603, 1604, as will be described later. The environment 1641 may further be selectively disconnected from the main controller 1601 and may be configured to be read-only from other environments 1602, 1603, 1604.

[0253] As shown in FIG. 16C, the main controller 1601 may be configured to monitor some or all of the environments 1602, 1603, 1604 even if the main controller 1601 is disconnected from any of its environments 1602, 1603, 1604. FIG. 16C shows that the in-band management connection 270 between the main controller 1601 and the environments 1602, 1603, 1604 is disconnected, which may help protect the main controller 1601 if the environments 1602, 1603, 1604 are put at risk. As shown in FIG. 16C, the out-of-band connection 260 may be maintained between the main controller 1601 and an environment such as 1602 even if the in-band connection 270 between the main controller 1601 and the environment 1602 is disconnected. Further, the environment 1641 may have a connection to the main controller 1601 that can be selectively enabled or disabled. The main controller 1601 may be separated from the environments 1602, 1603, 1604 or set up monitoring as a separate application within the air-gapped environment 1641. The main controller 1601 may use one-way communication for monitoring. For example, logs may be provided through one-way communication from the environments 1602, 1603, 1604 to the environment 1641. Through such one-way writing and via the connection between the environment 1641 and the main controller 1601, the main controller 1601 can collect data through the environment 1641 and monitor the environments 1602, 1603, 1604 even when there is no in-band connection 270 between the main controller 1601 and the environments 1602, 1603, 1604, thereby reducing the risk that the environments 1602, 1603, 1604 expose the main controller 1601 to danger. Access may be filtered or controlled, and / or access may be independent from the Internet. For example, as shown in FIG. 16D, when the in-band connection 270 between the main controller 1601 and the environment 1602 is connected, the main controller 1601 can control the network switch 1650 to disconnect the environment 1602 from an external network 1615 such as the Internet.When the environment 1602 is connected to the main controller 1601 by the in-band connection 270, disconnecting the environment 1602 from the external network 1615 can improve the security of the main controller 1601.

[0254] Therefore, it should be understood that the exemplary embodiments of FIGS. 16B - 16D show a way for the main controller to safely monitor the environments 1602, 1603, 1604 while minimizing the exposure to these environments 1602, 1603, 1604. Therefore, the main controller 1601 can disconnect itself from the environments 1602, 1603, 1604 (or at least disconnect itself from the in-band link) while maintaining a mechanism for monitoring these environments via a log server in the environment 1641 where the environments 1602, 1603, 1604 can have one-way write permissions. Therefore, if the main controller 1601 discovers during the process of reviewing the logs of the environment 1641 that the environment 1602 may be at risk due to malware, the main controller 1601 may use SDN tools to isolate the environment 1602 such that only the out-of-band connection 260 exists (see, for example, FIG. 16C). Additionally, the controller 1601 may send a notification regarding the potential problem to the administrator of the environment 1602. The controller may further isolate the at-risk environment 1602 by selectively disabling any connections (e.g., the in-band management connection 270) between the at-risk environment and any of the other environments 1603, 1604. In another example, the main controller 1601 may discover through the logs that the resources within the environment 1603 are too hot. Thereby, the main controller may intervene and migrate an application or service from the environment 1603 to another environment (either an existing environment or a newly created environment).

[0255] The controller 1601 may further set up one or more similar systems according to the requirements of the purchaser or user. As shown in FIG. 16E, the purchase application 1650 may be provided, for example, on a console or otherwise, whereby the purchaser can purchase a cloud, host, system environment or application and make a request to set them up for the purchaser. The purchase application 1650 may instruct the controller 1601 to set up the environment 1602. The environment 1602 may include, for example, a controller 1601a that deploys or constructs an IT system by allocating or assigning resources to the environment 1602.

[0256] FIG. 16F shows user interfaces 1632, 1633, 1634 that can be used in an environment where the environments 1602, 1603, 1604 each operate as a cloud and may or may not include a controller. The user interfaces 1632, 1633, 1634 (corresponding to the environments 1602, 1603, 1604 respectively) may be connected through a main controller 1601 that manages the connection between the user interface and the environment. Alternatively or additionally, an interface 1640a (which may take the form of a console) may be directly coupled to the environment 1602, an interface 1640b (which may take the form of a console) may be directly coupled to the environment 1603, and an interface 1640c (which may take the form of a console) may be directly coupled to the environment 1604. Regardless of whether the connection to the main controller 1601 is separated, disconnected, or disabled, the user may use one or more of the interfaces to use the environment or cloud.

[0257] System Cloning and Backup for Change Management Support Some of the environments 1602, 1603, 1604 may be clones of typical setup software used by developers. Also, these environments may be clones of the current working environment as a way of expansion, for example, by reducing latency caused by location by cloning the environment of another data center at a different location.

[0258] Therefore, it should be understood that the main controller that sets up the system and resources in individual environments or subsystems may enable cloning or backup of a part of the IT system. This may be used in test and change management as described herein. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, and / or other changes. Global rules may include a subset that includes backup rules that can be used in the change management described in various examples herein. Therefore, it should be understood that backup rules (examples of which are described elsewhere herein) can be used for change management. Examples of systems that implement backup rules are described in more detail with respect to FIGS. 21A - 21J.

[0259] According to an exemplary embodiment, the IT system or controller described herein may be configured to clone one or more environments. The new environment or cloned environment may or may not have the same resources as the original environment. For example, in a new environment or a nearly cloned environment, it may be desirable or necessary to use a completely different combination of physical and / or virtual resources. It may be desirable to clone the environment at different locations or times of day that can manage optimization of usage. It may be desirable to clone the environment into a virtual environment. When cloning an environment, the global system rules 210 and global template 230 of the controller or main controller may include information regarding how to configure and / or execute various types of hardware. The configuration rules within the system rules 210 may direct the placement and use of resources so that the resources and applications are more optimal, taking into account the specific available resources.

[0260] The main controller structure provides the ability to set up systems and resources into separate environments or subsystems, provides a structure for cloned environments, provides a structure for creating development environments, and / or provides a structure for deploying a standardized set of applications and / or resources. Such applications or resources can include, for example, those usable for application development and / or execution, or backup or restoration from backup of a portion of an IT system and other disaster recovery applications (e.g., systems including a LAMP (apache, mysql, php) stack, a web front end, and a server running react / redux, as well as resources running node.js and a mongo database and other standardized "stacks"), but are not limited thereto. Optionally, the main controller may deploy an environment that is a clone of another environment and may derive configuration rules from a subset of the configuration rules used to create the original environment.

[0261] According to an exemplary embodiment, change management of a system or a subset of a system may be achieved by cloning one or more environments and a configuration rule or a subset of configuration rules of such an environment. For example, changes may be needed to make changes to code, configuration rules, security patches, changes to templates, hardware changes, addition / removal of components and dependent applications, and other changes.

[0262] According to an exemplary embodiment, such changes to a system may be automated to avoid errors in direct manual entry of changes. The changes may be tested by a user in a development environment before automatically implementing the changes in the running system. According to an exemplary embodiment, an environment configured using the same configuration rules as the production environment may be cloned from the running production environment by using a controller to automatically power on, provision, and / or configure it. The cloned environment can be executed and run (while the backup environment is preferably left intact for emergencies in case changes need to be rolled back). This may be done using a controller to create, configure, and / or provision a new system or environment as described with reference to FIGS. 1-16F above using system rules 210, templates 230, and / or system state 220. The new environment may be used as a development environment to test changes to be implemented in the production environment later. The controller may generate the infrastructure of such an environment from a software-defined structure to the development environment.

[0263] As defined herein, the production environment means an environment used to operate a system, in contrast to an environment dedicated to development and testing, i.e., a development environment.

[0264] When the production environment is cloned, the infrastructure or the clone development environment is configured and generated by a controller according to the global system rule 210, similar to the production environment. Changes to the development environment may be made to the code, template 230 (either changes to existing templates or changes related to creating new templates), security, and / or application or infrastructure configuration. When new changes implemented in the development environment are prepared as needed through development and / or testing, the system automatically makes the changes to the development environment and then becomes operational or is deployed as the production environment. Next, the new system rule 210 is uploaded to either the environment controller or the main controller that applies changes to the system rules for a specific environment. The system state 220 is updated by the controller, and the added or modified template 230 may be implemented. Thus, complete system knowledge of the infrastructure may be maintained by the development environment and / or the main controller, along with the ability to recreate it. Complete system knowledge as used herein may include, but is not limited to, system knowledge of the state of resources, resource availability, and system configuration. Complete system knowledge may be collected by the controller by querying the resources using the system rule 210, system state 220, and / or in-band management connection 270 (s), out-of-band management connection 260 (s), and SAN connection 280 (s). Resources may be queried in particular to determine the usage rate, configuration state, or availability of resources, networks, or applications.

[0265] The cloned infrastructure or environment may be software defined via system rules 210, but this is not required. The cloned infrastructure or environment may or may not generally include a front end or user interface and one or more allocated resources. The allocated resources may or may not include compute resources, networking resources, storage resources, and / or application networking resources. This environment may or may not be configured as a front end, middleware, and database. The service or development environment may be booted with the system rules 210 of the production environment. The infrastructure or environment allocated for use by the controller may be software defined, particularly for cloning. Thus, the environment may be deployable by system rules 210 and may be cloneable by similar means. The clone environment or development environment may be automatically set up by the local or main controller using system rules 210 before or when changes are desired.

[0266] The data of the production environment may be written to read-only data storage until the development environment is separated from the production environment, after which it will be used by the development environment in the development and test processes.

[0267] The user or client may make and test changes in the development environment while the production environment is online. The data in the data storage may be changed during development and the changes are tested in the development environment. In a volatile or writable system, hot synchronization with the data of the production environment may also be used after the development environment is set up or deployed. Desired changes to the system, application, and / or environment may be made to the development environment and tested in the development environment. Next, the desired changes are made to the script of system rules 210 and a new version is created for the environment or the system as a whole and the main controller.

[0268] According to another exemplary embodiment, the newly developed environment may then be automatically implemented as the new production environment while the previous production environment is maintained or fully functional, and thus it is possible to revert to the previous state of the production environment without losing a large amount of data. Next, the development environment is booted with the new configuration rules within the system rules 210, the database is synchronized with the production database, and the writable database is switched. Thereafter, the original production database may be switched to a read-only database. If it is desirable to return to the previous production environment, the previous production environment is maintained as a copy of the previous production environment for the required period of time.

[0269] The environment may be configured as a single server or instance that may include physical and / or virtual hosts, networks, and other resources. In another exemplary embodiment, the environment may be a plurality of servers that include physical and / or virtual hosts, networks, and other resources. For example, there may be a plurality of servers that form a load-balanced Internet-facing application, and those servers may be connected to a plurality of API / middleware applications (which may be hosted on one or more servers). The database of the environment may include one or more databases through which the API communicates queries within the environment. The environment may be constructed from the system rules 210 in a static or volatile form. The environment or instance may be virtual, physical, or a combination of each.

[0270] The configuration rules of the application or the system within the system rules 210 may specify various compute backends (e.g., bare metal, AMD epyc servers, Intel Haswell on qemu / kvm) and may include rules regarding how to execute the application or service on the new compute backend. Thus, for example, if there is a situation where the availability of resources for testing decreases, the application may be virtualized.

[0271] Using the examples described herein and in accordance therewith, the test environment may be deployed on virtual resources where the original environment uses physical resources. Using the controllers described herein with reference to FIGS. 1-18B, and as further described herein, a system or environment may be cloned from a physical environment to an environment that may or may not include virtual resources, either in whole or in part.

[0272] FIG. 17A shows an exemplary embodiment in which system 100 includes controller 1701 and one or more environments, such as 1702, 1703, 1704. System 100 may be a static system, i.e., a system in which active user data does not constantly change the state of the system or frequently manipulate data, such as a system that hosts only static web pages. The system may be coupled to a user (or application) interface 110.

[0273] Controller 1701 may be configured in a manner similar to controllers 200 / 1401 / 1501 / 1601 described herein, and may similarly include global system rules 210, controller logic 205, templates 230, and system state elements 220. Controller 1701 may be coupled to one or more other controllers or environments in a manner as described with reference to FIGS. 14A-16F herein. The global rules 210 of controller 1701 may include rules that may manage and control other controllers and / or environments. Using such global rules 210, controller logic 205, system state 220, and templates 230, a system or environment may be set up, provisioned, and deployed through controller 1701 in a manner similar to that described with reference to FIGS. 1-16F herein. Each environment may be configured using a subset of the global system rules 210 that define the operation of the environment, including with respect to other environments.

[0274] Global system rule 210 may also include change management rule 1711. Change management rule 1711 includes a set of rules and / or instructions that may be used when changes to system 100, global system rule 210, and / or controller logic 205 are desired. Change management rule 1711 may be configured to enable a user or developer to develop a change, test the change in a test environment, and then implement the change by automatically converting the change into a new set of configuration rules within system rule 210. Change management rule 1711 may be a subset of global system rule 210 (as shown in FIG. 17A), or may be separate from global system rule 210. The change management rule may use a subset of global system rule 210. For example, global system rule 210 may include a subset of environment creation rules configured to create a new environment. Change management rule 1711 may be configured to set up and use a system or environment configured and set up by controller 1701 to copy and clone some or all aspects of system 100. Change management rule 1711 may be configured to permit testing of new changes proposed to the system prior to implementation by using a clone of the system for testing and implementation. Change management 1711 may include, or may use, backup rules described below.

[0275] Clone 1705, as shown in FIG. 17A, may include rules, logic, applications, and / or resources of a particular environment or part of system 100. Clone 1705 may be equipped with the same or different hardware as system 100, and may or may not use virtual resources. Clone 1705 may be set up as an application. Clone 1705 may be set up and configured using configuration rules within system rules 210 of system 100 or controller 1701. Clone 1705 may or may not include a controller. Clone 1705 may include allocated networking resources, computing resources, application networks, and / or data storage resources, as described in more detail above. Such resources may be allocated using change management rules 1711 controlled by controller 1701. Clone 1705 may be coupled to a user interface that enables a user to make changes to clone 1705. The user interface may be the same as or different from user interface 110 of system 100. Clone 1705 may be used for the entire system 100 or for a part of system 100, such as one or more environments and / or controllers. Clone 1705 may or may not be a complete copy of system 100. Clone 1705 may be selectively enabled and / or completely disabled, and / or may be converted to a one-way read and / or write connection, and may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280. Thus, when the clone environment 1705 is separated from the production environment during testing, or until the clone environment 1705 is ready to go online as a new production environment, the connection to the data within the clone environment 1705 may be changed to make the clone data read-only. For example, if clone 1705 has a data connection to environment 1702, this data connection may be made read-only for separation.

[0276] Any backup 1706 may or may not be used for the entire system or for a part of the system such as one or more environments and / or controllers. When performing the change management function, individual services may also be backed up. The backup of the service may be executed using backup rules as described below, for example, with reference to FIGS. 21A - 21J. Backup 1706 may include networking resources, computing resources, application network resources, and / or data storage resources as described in more detail above. Backup 1706 may or may not include a controller. Backup 1706 may be a complete copy of system 100. Backup 1706 may be set up as an application or using hardware that is the same as or different from system 100. Backup 1706 may be selectively enabled and / or completely disabled and / or may be coupled to system 100 via in - band management connection 270, out - of - band management connection 260, and / or SAN connection 280 that can be converted to a single - direction read and / or write connection.

[0277] FIG. 17B shows an exemplary process flow for using the clone and backup systems of FIG. 17A in system change management. At step 1785, a user or management application initiates a change to the system. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, changes to templates, hardware changes, addition / removal of components and / or dependent applications, and other changes. At step 1786, controller 1701 sets up the environment in the manner described with respect to FIGS. 14A - 16F to become clone environment 1705 (where the clone environment may have its own new controller or may use the same controller as the original environment).

[0278] In step 1787, the controller 1701 may use the global rule 210 including the change management rule 1711 to clone all or part of one or more environments of the system (e.g., the "production environment") to the clone environment 1705 (e.g., where the clone environment 1705 may function as a "development environment" here). The controller 1701 may extract data using the backup rule 2104, and the extracted data may be restored later using the backup rule as described with reference to FIGS. 21A - 21J herein. Thus, the controller 1701 identifies and allocates resources, uses the system rule 210 to set up and allocate clone resources, and copies any of data, configuration, code, executable files, and other information required for application startup from the environment to the clone. In step 1788, the controller 1701 optionally backs up the system by setting up another environment that functions as a backup 1706 using the configuration rules within the system rule 210 (regardless of the presence or absence of the controller), and copies the template 230, the controller logic 205, and the global rule 210.

[0279] After clone 1705 is created from the production environment, clone 1705 may be used as a development environment in which changes to the clone's code, configuration rules, security patches, templates, and other changes can be made. At step 1789, the changes to the development environment may be tested before implementation. During testing, clone 1706 can be isolated from the production environment (system 100) or other components of the system. This can be achieved by selectively disabling one or more of the connections between system 100 and clone 1706 in controller 1701 (e.g., by disabling in-band management connection 270 and / or by disabling the application network connection). At step 1790, it is determined whether the modified development environment is ready. At step 1709, if it is determined that the development environment is not yet ready (which is usually a decision made by the developer), the process flow returns to step 1789 for further changes to the clone environment 1705. At step 1790, if it is determined that the development environment is ready, at step 1791, the development environment and the production environment can be switched. That is, the controller may change development environment 1705 to the new production environment and maintain the previous production environment until the migration to the development / new production environment is complete and in a good state.

[0280] FIG. 18A shows another exemplary embodiment of system 100 that can be set up and used in the change management of a system. In the example of FIG. 18A, system 100 includes a controller 1801 and one or more environments 1802, 1803, 1804, 1805. The system is shown with a clone environment 1807 and a backup system 1808. Backup and data recovery can be performed using the backup rules described elsewhere in this specification. Examples of backup system management are further described herein with reference to FIGS. 21A - 21J.

[0281] Controller 1801 is configured in a manner similar to the controllers 200 / 1401 / 1501 / 1601 / 1701 described herein and may include elements of the global system rules 210, controller logic 205, templates 230, and system state 220. Controller 1801 may be coupled to one or more other controllers or environments in a manner as described with reference to FIGS. 14A-16F herein. The global rules 210 of controller 1801 may include rules that may manage and control other controllers and / or environments. Using such global rules 210, controller logic 205, system state 220, and templates 230, a system or environment may be set up, provisioned, and deployed through controller 1801 in a manner similar to that described with reference to FIGS. 1-17B herein. Each environment may be configured using a subset of the global rules 210 that define the actions of the environment, including actions regarding other environments.

[0282] Global rule 210 may also include change management rule 1811. The change management rule 1811 may include a set of rules and / or instructions that may be used when changes to the system, global rules, and / or logic are desired. The change management rule may be configured to enable a user or developer to develop a change, test the change in a test environment, and then implement the change by automatically converting the change into a new set of configuration rules within the system rule 210. The change management rule 1811 may be a subset of the global system rule 210 (as shown in FIG. 18A), or may be separate from the global system rule 210. The change management rule 1711 may use a subset of the global system rule 210. For example, the global system rule 210 may include a subset of environment creation rules configured to create a new environment. The change management rule 1811 may be configured to set up and use a system or environment set up and deployed by the controller 1801 to copy and clone some or all aspects of the system 100. The change management rule 1811 may be configured to permit testing of new proposed changes to the system before implementing them by using a clone of the system for testing and implementation. The change management rule 1811 may include or use backup rules described elsewhere in this specification. The backup rule may extract data using the backup rule 2104, and the extracted data may later be restored using the backup rules described with reference to FIGS. 21A-21J of this specification.

[0283] As shown in FIG. 18A, the clone environment 1807 may include a controller 1807a having rules, controller logic, templates, and system state data, and an assigned resource 1820 that can be assigned and set up according to the global system rules 210 and change management rules 1811 of the controller 1801 for one or more environments. The backup system 1808 may further include a controller 1808a having rules, controller logic, templates, and system state data, and an assigned resource 1821 that can be assigned and set up according to the global system rules 210 and change management rules 1811 of the controller 1801 for one or more environments. The system may be coupled to the user (or application) interface 110 or another user interface.

[0284] The clone environment 1807 may include the rules, logic, templates, system state, applications, and / or resources of a particular environment or part of a system. The clone 1807 may be equipped with the same or different hardware as the system 100, and the clone 1807 may or may not use virtual resources. The clone 1807 may be set up as an application. The clone 1807 may be set up and configured using the configuration rules within the system rules 210 of the system 100 or the controller 1801 for the environment. The clone 1807 may or may not include a controller and may share the controller with the production environment. The clone 1807 may include the assigned networking resources, computing resources, application network resources, and / or data storage resources as described in more detail above. Such resources may be assigned using the change management rules 1811 controlled by the controller 1801. The clone 1807 may be coupled to a user interface that allows the user to make changes to the clone 1807. The user interface may be the same as or different from the user interface 110 of the system 100.

[0285] Clone 1807 may be used for the entire system or for a part of the system such as one or more environments and / or controllers. In an exemplary embodiment, Clone 1807 may include a hot standby data resource 1820a coupled to the data resource 1820 of Environment 1802. The hot standby data resource 1820a may be used during the setup of Clone 1807 and during the testing of changes. The hot standby data resource 1820a may be selectively disconnected or separated from the storage resource 1820 during change management, as described herein with respect to FIG. 18B, for example. Clone 1807 may or may not be a complete copy of System 100. Clone 1807 may be selectively enabled and / or completely disabled, and / or may be converted to a one-way read and / or write connection, and may be coupled to System 100 via an in-band management connection 270, an out-of-band management connection 260, and / or a SAN connection 280. Thus, when the clone environment 1807 is separated from the production environment during testing, or until the clone environment is ready to go online as a new production environment, the connection to the volatile data within the clone environment 1807 may be changed to make the clone data read-only.

[0286] When switching from an old production environment to a new production environment, the controller 1801 may instruct the front end, load balancer, or other application or resource to point to the new production environment. Accordingly, users, application resources, and / or other connections may be redirected when the change is made. This may be accomplished, for example, by changing the list of ip / ipoib addresses, Infiniband GUIDs, DNS servers, Infiniband partitions / opensm configurations, or software-defined networking (SDN) configurations that can be achieved by sending commands to networking resources, but is not limited to these methods. The front end, load balancer, or other application and / or resource may refer to systems, environments, and / or other applications including, but not limited to, databases, middleware, and / or other back ends. Such a load balancer may be used for change management to switch from an old production environment to a new environment.

[0287] Clone 1807 and backup 1808 may be set up and used when managing the manner of changes to the system. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, changes to templates, hardware changes, addition / removal of components and / or dependent applications, and other changes. Backup 1808 may be used for the entire system, or for a part of the system such as one or more environments and / or controller 1801. Backup 1808 may include networking resources, computing resources, application networks, and / or data storage resources, as described in more detail above. Backup 1808 may or may not include a controller. Backup 1808 may be a complete copy of system 100. Backup 1808 may include the data necessary to reconstruct the system / environment / application from the configuration rules included in the backup, and may include all application data. Backup 1808 may be set up as an application, or using hardware that is the same as or different from system 100. Backup 1808 may be selectively enabled and / or disabled, and / or may be converted to a one-way read and / or write connection, and may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280.

[0288] Figure 18B is an exemplary process flow showing the use of the system of Figure 18A in change management, and in particular shows the case where the system of Figure 18A includes volatile data, or where the database is writable. Such a database may be part of the storage resources used by an environment within the system. In step 1870, the system is deployed using global system rules (including the production environment).

[0289] Next, at step 1871, the production environment is cloned using the global system rule 210 including the change management rule 1811 and resource allocation by the main controller 1801 or a controller within the clone environment, in order to create a read-only environment in which the clone environment is disabled from writing to the system. The clone environment can then be used as a development environment.

[0290] At step 1872, the hot standby 1820a is started and assigned to the clone environment 1807 to store any volatile data changed within the system 100. The clone data is updated and the new version within the development environment can be tested with the updated data. The hot synchronization data can be turned off at any time. For example, when writing from an old environment or a production environment to the development environment is being tested, the hot synchronization data can be turned off.

[0291] Next, at step 1873, the user can make changes using the clone environment 1807 as a development environment. Next, at step 1874, the changes to the development environment are tested. At step 1875, it is determined whether the modified development environment is ready (usually, such determination is made by the developer). If it is determined at step 1875 that the changes are not ready, the process flow may return to step 1873 so that the user can go back and make other changes to the development environment. If it is determined at step 1875 that the changes are ready to be deployed, the process flow proceeds to step 1876, where the configuration rules are updated within the system or the controller for a specific environment and used to deploy the new updated environment.

[0292] In step 1877, next, the development environment (or, a new environment) may be changed and redeployed in the desired final configuration with the desired resources and hardware allocation before operation. In the next step 1878, the write function of the original production environment is disabled and the original production environment becomes read-only. While the original production environment is read-only, as part of 1878, any new data from the original production environment (or perhaps, also the new production environment) may be cached and identified as migration data. As an example, the data can be cached in a database server or other appropriate location (e.g., a shared environment). Next, the development environment (or, the new environment) and the old production environment are switched in step 1879, and the development environment (or, the new environment) becomes the production environment.

[0293] After this switch, the new production environment is made writable in step 1880. If in step 1881 it is determined that the new production environment is functioning as the developer determined, all data loss during the switch process (such data was cached in step 1878) may be written to the new environment and verified in step 1884. After such verification, the change is completed (step 1885).

[0294] If as a result of step 1881 it is determined that the new production environment is not functioning (e.g., a problem is identified where the system needs to be reverted to the old system), in step 1882 the environment is reverted and the old production environment becomes the production environment again. As part of step 182, the configuration rules of the target environment of the controller 1801 are reverted to the previous version that was used for the currently reverted production environment.

[0295] In step 1883, database changes may be determined, for example, using cached data, and the data is restored to an old production environment with old configuration rules. To support step 1883, the database can maintain a log of changes made to the database, so step 1883 can determine the changes that may need to be undone. Backup databases that track cached data and record time may be used to cache data as described above, and the clock may be turned back to determine what changes were made. Snapshots and logs may be used for this purpose.

[0296] After restoring the cached data at 1883, if you want to start again, the process may return to step 1871.

[0297] Examples of the change management system described herein may be used, for example, when updating, adding, or removing hardware or software, applying patches to software, when a system failure is detected, when migrating a host during a hardware failure or detection, for dynamic resource migration, for changes to configuration rules or templates, and / or when making changes related to any other system. The controller 1801 or the system 100 may be configured to detect a failure, and upon detecting a failure, the controller may automatically implement change management rules or existing configuration rules on other hardware of the available system. Examples of available failure detection methods include, but are not limited to, pinging the host, querying the application, and executing various tests or test suites. The change management configuration rules described herein may be implemented when a failure is detected. Such rules may trigger the automatic generation of a backup environment and the automatic migration of data or resources being implemented by the controller upon detection of a failure. The selection of backup resources may be based on resource parameters. Such resource parameters may include, but are not limited to, usage information, speed, configuration rules, and data capacity and usage.

[0298] As described herein, whenever a change occurs, the controller always creates a log of the change and what was actually executed. For security or system updates, the controller described herein may be configured to automatically turn on and off according to configuration rules and to update the IT system state. The controller may turn off resources to conserve power. The controller may turn on or migrate resources to vary efficiency by time. When migrating, a backup or copy of the environment or system may be created according to configuration rules. In the event of a security breach, the controller may isolate and block the area under attack.

[0299] Configuration and Control of Service Dependencies FIG. 19A shows an exemplary system 100 described herein with reference to FIGS. 1-18, where the system 100 is enhanced using interrelated services (or applications) represented by corresponding service modules 1901, 1902 on one or more resources 1910. The service modules 1901, 1902 can take the form of computer-executable code that provides services such as authentication, email, webmail, web services, middleware, databases, and / or other services. When referring to each, service 1901 can be referred to as service A and service 1902 can be referred to as service B.

[0300] The system of FIG. 19A may be connected to an external network 1980 and / or an application network 390, and this connection may be made invalid and valid according to the description set forth in FIGS. 13A-13E. Services 1901, 1902 are configured by a controller 200 as resources or applications described in various embodiments herein with reference to FIGS. 1-18.

[0301] Services 1901 and 1902 can be controlled by the controller 200 and may interoperate through the common API 1903. Services 1901 and 1902 can resolve dependencies using the common API 1903. For example, assume a web application requires an http server. A service having apache or nginx may have a "common web server API", and the "common web server API" may serve the content of the webapp to that server and proxy information back to the application. Services 1901, 1902, and API 1903 may be coupled directly through the management network, or through a management connection to the controller 200 (e.g., 260 and / or 270), or through any other network connection between the controller 200, service 1901, common API 1903, and service 1902.

[0302] The common API 1903 may run on or in response to one or more of services 1901, 1902, controller 200, or other resources of the system 100. Service A of service module 1901 may be a dependent service configured to be called through the common API 1903 by a dependent service B on service module 1902 to perform one or more functions. A dependent service is a service that can satisfy the dependencies of other services (in this case, the other services are "dependent services"). A dependent service may also be an optional dependent service.

[0303] Services 1901 and 1902 may be configured or set up by the controller 200 to interoperate securely with each other.

[0304] The example of the interoperability of the service and the controller in FIG. 19A is described with reference to the flow in FIG. 19B. The service may be started, for example, by the controller 200 using the configuration rules described herein (see 19.1). The controller 200 resolves the dependencies described in the figures herein (see 19.2). The service may have a set of dependencies listed in its specification. As an example, this can be done using the json specification of the service. The system may further use a dependency resolution similar to how the package manager functions and provide the user with a way to satisfy the dependencies. As another example, the system may propose to the user the installation of the dependent service or the use / selection of an existing dependent service. The dependent service B calls the dependent service A through the common API 1903 (see 19.3). This call may be a call to configure service A to support service B or a call to use a part of the functionality of service A. The common API 1903 is converted to the dependent service A (1901) and instructs the dependent service A to execute the command(s) (see 19.4). The conversion can be done by one of the services or the controller 200 by making an API call (and here, the function of the API may be to call other API functions on different APIs).

[0305] According to some exemplary embodiments described herein, a system having a controller 200, a dependent service 1901, and a dependent service 1902 can be provided with additional security. This additional security is useful when multiple services are connected to the in-band management connection 270 and can communicate directly with each other. Such additional security can be provided during configuration, reconfiguration, and / or operation based on and / or using the controller's global system rules 210, logic 205, templates 230, and / or system state 220. The additional security may be provided in some exemplary embodiments where the services 1901, 1902 communicate via the in-band management connection 270, or other networks or interconnects. According to some exemplary embodiments, the dependent service 1901 is configured to request the controller to verify the validity of the dependent service 1902. This may include verifying the identity of the dependent service 1902, or the identity of a service that executes commands on the API. According to some exemplary embodiments, the dependent service is configured to request the controller 200 for permission to execute its functions, tasks, and / or a plurality of tasks, functions, or combinations thereof for a particular dependent service(s). The dependent service may further or alternatively be provisioned with permission, or a set of permissions, by the controller 200 when the dependent service is configured or reconfigured. The set of permissions for the dependent service may also be updated. For example, the set of permissions may be updated when a dependent service is added.

[0306] Examples of authentication and authorization provided between services are illustrated by the flow shown in FIG. 19C. At step 19.11, the controller 200 provisions a service key or key pair during configuration, thereby enabling authentication between the service and the controller 200. This step can be performed for each service in the system. A dependent service that requires execution by a dependent service is authenticated by the dependent service requesting a validity check from the controller. The identity of the service and the data transmitted and received by the service may be authenticated, for example, by mutual tls authentication, public key authentication, other forms of encryption, any network-based validity check technology (including but not limited to vlans, vxlans, partitions, etc.), and / or combinations thereof. Virtual networks and partitions can be used to divide the network into mini-networks such as Infiniband partitions. As a result, a scenario is possible where, for example, if a port is on partitions 4 and 15, it can only communicate with things on partitions 4 and 15. According to some variations, the controller 200 can be used as a key distribution center while maintaining verification or authentication within service modules 1901, 1902 independent of the controller 200. During service configuration, the controller may provision the service with a key or key pair that may include a public key and / or a private key that directly or indirectly enables authentication between the service and the controller 200. According to one example, the controller 200 can store the public key and delete or invalidate the provisioned private key. As an additional example, the service may generate its own key such that the public key from the service can be identified, recognized, and / or authenticated by the controller 200. In this additional example, since the controller 200 provided the service's initial public key, the service has the ability to send a public key that the controller 200, which does not have knowledge of the current private key, can trust, and the service authenticates with the original key pair of the service and shares the new public key with the controller 200.

[0307] In step 19.12, the dependent service (Service B) calls the dependent service (Service A) to execute a function through the common API 1903. The dependent service (Service A) then authenticates the dependent service (Service B) by the controller 200. According to an example, the dependent service may contact the controller 200, and the controller 200 can authenticate the requesting dependent service using the public key provided by the dependent service. As described above, the dependent service can obtain the public key in step 19.11. The dependent service may also generate a new key pair (public key + private key), and (assuming the controller 200 created the public key and private key), since it is known that the controller 200 trusts the old public key, the old key pair can be used to prove that the new public key is genuine. The dependent service (Service B) may authenticate the dependent service in a similar manner (19.13).

[0308] In step 19.14, the dependent service (Service A) may also establish permission for the dependent service (Service B) to execute a function before executing the function. For example, the permission may be established by the dependent service by asking the controller 200 whether permission is possible. As another example, the permission may be established by the controller 200 through a permission list loaded on the dependent service (Service A).

[0309] Figure 19D shows an example of an enhanced security method used with an interoperability service. At step 19.21, dependent service B may be created using, for example, controller 200 and / or template 230 described herein. At step 19.22, when dependent service B is executed, controller 200 may verify and / or disable connections to external network 1910 and / or application network 390 as described with respect to various embodiments herein (see, e.g., FIGS. 13A-13E). For example, an implementer may desire to disable management connections while the service is open to a network such as the Internet. This provides additional isolation and security as described above. In such a case, an in-band management connection 270 can be switched using a cloud API (or, in other cases, an out-of-band management connection 260). At step 19.23, dependent service B executes an API command to request a service or function from dependent service A. This step may also be completed by dependent service B requesting it from controller 200 and controller 200 executing the command through a common API (see 1903). At step 19.24, dependent service A verifies the identity of dependent service B and the validity of the permission to execute the service or function for service B. As an example, this step 19.24 can be performed by service A verifying the validity of the permission of service B. As another example, this step 19.24 can be performed by service A verifying that service A has permission to provide a service to service B. In either case, a service modified by another service can verify that the other service is permitted to make those modifications. If authenticated and permitted, dependent service A may execute the service or function specified by the command (see 19.25).In step 19.26, management connections such as out-of-band management 260, in-band management 270, or SAN 280 may optionally be disconnected for additional security as described herein with reference to FIGS. 13A-13E. In step 19.27, connections to external network(s) 1980 and / or application network(s) 390 may be re-enabled if the connections were disabled in 19.22.

[0310] FIG. 19E shows an exemplary system 100 such as the system described with respect to FIGS. 19A-19D, and a set of cleanup rules 1904 is included in controller 200. The cleanup rules 1904 may be embodied as its own set of rules within the controller 200, or may be embodied by global system rules 210, controller logic 205, templates 230, or combinations thereof. The cleanup rules 1904 include a set of instructions and rules to follow when a service is deleted. By way of example, the cleanup rule 204 may be included in a template 230 used for setup of a service, in which case rules related to the service are loaded into the service during setup, or a service-specific cleanup rule is generated using rules related to the service. For example, a dns record may be added to the dns service by a mail service. When that mail service is deleted, the dns service can remove the dns record from that mail service.

[0311] The cleanup rules 1904 may be used to identify modifications made to a dependent service by a dependency service and, when the dependency service is deleted or disabled, enable deletion, removal, and / or undo of these modifications. FIG. 19F shows an exemplary process flow for creating cleanup rules. For example, FIG. 19F shows how modifications can be identified, e.g., logged an...

Claims

1. A controller, A resource for connecting to the controller, the resource includes a first service and a second service, the first and second services have a dependency relationship with each other, the first service includes a dependent service for the second service, and the second service includes a dependent service for the first service, the resource, An application program interface (API) for interfacing the first and second services with each other and with the controller, An information technology (IT) computer system comprising: The controller is configured to manage the interoperability of the first service with respect to the second service. The information technology (IT) computer system.

2. The system according to claim 1, wherein the second service is configured to issue a call so that the second service executes an operation on the first service through the API of the first service.

3. The system according to any one of claims 1 to 2, wherein the first service is configured to request the controller to confirm the validity of the second service in order to execute an operation for the second service.

4. The system according to claim 3, wherein the controller is further configured to confirm the validity of the second service based on at least one of mutual TLS authentication, public key authentication, and / or network-based validity confirmation.

5. The system according to any one of claims 1 to 4, wherein the first service is configured to execute an operation for the second service on condition of permission from the controller.

6. The system according to any one of claims 1 to 5, wherein the controller is further configured to provision encryption keys to the first and second services for use in validating the first and second services.

7. The system according to claim 6, wherein the encryption keys include different key pairs for each of the first and second services, and each key pair includes a public key and a private key.

8. The system according to claim 7, wherein after the controller provisions the private key to the first and second services, the controller is further configured to delete a copy of the private key, and the controller is configured to manage verification of the validity of the first and second services based on the public key.

9. The system according to any one of claims 1 to 8, wherein when the first and / or second service is being configured by the controller, the controller is further configured to disconnect the resource from any external network.

10. The system according to claim 9, wherein after the first and / or second service is configured by the controller, the controller is further configured to reconnect the resource to any disconnected external network.

11. The system according to any one of claims 9 to 10, further comprising at least one of an in-band connection, an out-of-band connection, and / or a storage area network connection between the resource and the controller, and when the first and / or second service is being configured by the controller, the controller is further configured to disconnect the resource from the in-band connection, the out-of-band connection, and / or the storage area network connection.

12. The system according to claim 11, wherein after the first and / or second service is configured by the controller, the controller is further configured to reconnect the resource to the disconnected in-band connection, out-of-band connection, and / or storage area network connection.

13. The system according to any one of claims 1 to 12, wherein the controller is further configured to resolve a dependency between the first service and the second service.

14. An information technology (IT) method used in a computer system including a controller and a resource connected to the controller, the resource including a first service and a second service, the first and second services having a dependency relationship with each other, the first service including a dependent service with respect to the second service, and the second service including a dependent service with respect to the first service, the method comprising: interfacing the first and second services with each other and with the controller via an application program interface; the controller managing the interoperability of the first service with respect to the second service; the method comprising the above.

15. The method according to claim 14, further comprising the second service calling the first service to execute an operation through the API of the first service. The method according to claim 14, further comprising the first service requesting the controller to perform validity verification of the second service in order to execute an operation for the second service.

16. The method according to any one of claims 14 to 15, further comprising the controller performing validity verification of the second service based on at least one of mutual TLS authentication, public key authentication, and / or network-based validity verification.

17. The method according to any one of claims 14 to 17, further comprising the first service operating for the second service on condition of permission from the controller.

18. The method according to any one of claims 14 to 18, further comprising the controller provisioning encryption keys to the first and second services for use in validity verification of the first and second services.

19. The method according to claim 19, wherein the encryption keys include different key pairs for each of the first and second services, and each key pair includes a public key and a private key.

20. The method according to claim 20, further comprising the controller deleting a copy of the private key after provisioning the private key to the first and second services, and the controller managing validity verification of the first and second services based on the public key.

21. The method according to claim 20, further comprising the controller deleting a copy of the private key after provisioning the private key to the first and second services, and the controller managing validity verification of the first and second services based on the public key.

22. The method according to claim 21, further comprising the controller deleting a copy of the private key after provisioning the private key to the first and second services, and the controller managing validity verification of the first and second services based on the public key.

23. The method according to claim 22, further comprising the controller deleting a copy of the private key after provisioning the private key to the first and second services, and the controller managing validity verification of the first and second services based on the public key. The method according to claim 23, further comprising the controller deleting a copy of the private key after provisioning the private key to the first and second services, and the controller managing validity verification of the first and second services based on the public key. The method according to claim 20, further comprising

22. When the first and / or second service is being configured by the controller, the controller disconnects the resource from any external network The method according to any one of claims 14 to 21, further comprising

23. After the first and / or second service has been configured by the controller, the controller reconnects the resource to any disconnected external network The method according to claim 22, further comprising

24. The resource is connected to the controller via at least one of an in-band connection, an out-of-band connection, and / or a storage area network connection between the resource and the controller, and the method comprises When the first and / or second service is being configured by the controller, the controller further disconnects the resource from the in-band connection, the out-of-band connection, and / or the storage area network connection The method according to any one of claims 22 to 23, further comprising

25. After the first and / or second service has been configured by the controller, the controller reconnects the resource to the disconnected in-band connection, out-of-band connection, and / or storage area network connection The method according to claim 24, further comprising

26. The controller resolves a dependency between the first service and the second service The method according to any one of claims 14 to 25, further comprising

27. A controller, and A resource for connecting to the controller An information technology (IT) computer system comprising the resource includes a first service and a second service, the first and second services have a dependency relationship with each other, the first service includes a dependent service with respect to the second service, and the second service includes a dependent service with respect to the first service The controller is configured to maintain a cleanup rule that identifies a modification made to the first service as a dependency of the second service, and the cleanup rule supports deletion, removal, and / or undo of the modification to the first service when the second service is deleted and / or invalidated. The information technology (IT) computer system.

28. The second service is configured to call the first service such that the first service executes an operation, and the modification to the first service is related to the execution of the operation. The controller is configured to identify the modification to the first service as part of the cleanup rule. The system according to claim 27.

29. The system according to claim 28, wherein the cleanup rule associates the modification to the first service with the second service.

30. The controller is further configured to (1) determine that the second service is to be or has been deleted and / or invalidated, and (2) in response to determining that the second service is to be or has been deleted and / or invalidated, apply the cleanup rule to delete, remove, and / or undo the identified modification to the first service. The system according to any one of claims 28 to 29.

31. The resource includes a plurality of additional services that have dependencies on each other, and the cleanup rule includes a plurality of cleanup rules that identify modifications made to the service by a service. The system according to any one of claims 27 to 30.

32. Each service that is a dependent service is associated with a cleanup rule that identifies a modification to its dependent service. The system according to claim 31.

33. An application program interface (API) that interfaces services with each other and with the controller The system according to any one of claims 31 to 32, further comprising.

34. The dependent services are configured to call their dependent services through the API. The controller is configured to identify the modification to the dependent service as part of the cleanup rule. The system according to claim 33.

35. The system according to claim 34, wherein the API is configured to log commands from the service, and the controller is further configured to generate the cleanup rules from the logged API commands.

36. The system according to any one of claims 27 to 35, further comprising a plurality of resources for connecting to the controller, the resources including a plurality of services having dependencies on each other, the resources including a plurality of services having dependencies on each other, and the cleanup rules including a plurality of cleanup rules for identifying modifications made to the service by the service.

37. Executing a first service and a second service on the resources of a computer system including a controller connected to the resources, the first and second services having a dependency relationship with each other, the first service including a dependent service for the second service, and the second service including a dependent service for the first service; executing the first service and the second service; The controller maintaining a cleanup rule that identifies a modification made to the first service as a dependency of the second service, the cleanup rule supporting deletion, removal, and / or undo of the modification to the first service when the second service is deleted and / or invalidated; maintaining the cleanup rule; An information technology (IT) method comprising.

38. The calling, wherein the second service calls the first service to perform an operation, and the modification to the first service is related to the performance of the operation; The controller identifying the modification to the first service as part of the cleanup rule; The method according to claim 37, further comprising.

39. The method according to claim 38, wherein the cleanup rule associates the modification to the first service with the second service.

40. The controller determining that the second service is to be, or should be, deleted and / or invalidated; In response to the determination that the second service has been or is to be deleted and / or disabled, the controller applies the cleanup rule that deletes, removes, and / or undoes the identified modification to the first service The method according to any one of claims 38 to 39, further comprising: **Claim 41** The method according to any one of claims 37 to 40, wherein the resource includes a plurality of additional services having dependencies on each other, and the cleanup rule includes a plurality of cleanup rules that identify modifications made to the service by other services **Claim 42** The method according to claim 41, wherein each service that is a dependent service is associated with a cleanup rule that identifies a modification to its dependent service **Claim 43** An application program interface (API) interfaces the services with each other and with the controller The method according to any one of claims 41 to 42, further comprising: **Claim 44** The dependent services call their dependent services through the API, and the controller identifies the modification to the dependent service as part of the cleanup rule The method according to claim 43, further comprising: **Claim 45** The API logs commands from the service, and the controller generates the cleanup rule from the logged API commands The method according to claim 44, further comprising: **Claim 46** The method according to any one of claims 37 to 45, wherein a plurality of resources are connected to the controller, the resources include a plurality of services having dependencies on each other, and the cleanup rule includes a plurality of cleanup rules that identify modifications made to the service by the service **Claim 47** A controller, a computing resource for connecting to the controller, and a storage resource for use by the computing resource, wherein the controller is configured to provision storage authentication information of the storage resource to the computing resource ​ The computing resource is configured to connect to, log on to, and / or communicate with the storage resource based on the storage authentication information. The information technology (IT) computer system. **Claim 48** The system according to claim 47, wherein the controller is further configured to (1) pair the computing resource and the storage resource, and (2) place the paired computing resource and the storage resource on the same network or fabric. **Claim 49** The system according to claim 48, wherein the controller is further configured to place the paired computing resource and the storage resource on the same network or fabric by (1) disabling a storage area network (SAN) connection between the computing resource and the storage resource, and (2) making the storage resource available to the computing resource on a separate connection network. **Claim 50** The system according to claim 49, wherein the separate connection network includes vlans, vx lans, and / or InfiniBand partitions. **Claim 51** The system according to any one of claims 47 to 50, wherein the storage authentication information includes a password or a passphrase. **Claim 52** The system according to any one of claims 47 to 51, wherein the storage authentication information includes a chap key. **Claim 53** The system according to any one of claims 47 to 52, wherein the storage authentication information includes an encryption key. **Claim 54** The system according to claim 50, wherein the computing resource is configured to calculate login authentication information of the storage resource from the encryption key based on an encryption technique. **Claim 55** The system according to any one of claims 47 to 54, wherein the storage authentication information includes a certificate. **Claim 56** The system according to any one of claims 47 to 55, further comprising a plurality of computing resources and a plurality of storage resources, wherein the controller is configured to provision the storage authentication information of the storage resources to the computing resources such that the plurality of computing resources are paired with different storage resources among the storage resources. **Claim 57** An information technology (IT) method for use with a computer system including a controller, a computing resource connected to the controller, and a storage resource for use by the computing resource, comprising: the controller provisioning storage authentication information of the storage resource to the computing resource; the computing resource connecting to, logging on to, and / or communicating with the storage resource based on the storage authentication information The method as described above.

58. the controller further pairing the computing resource with the storage resource; the controller placing the paired computing resource and the storage resource on the same network or fabric The method according to claim 57, further comprising:

59. the controller further disabling a storage area network (SAN) connection between the computing resource and the storage resource; the placing step includes enabling the computing resource to utilize the storage resource on a separate connection network; The method according to claim 58.

60. The method according to claim 59, wherein the separate connection network includes vlans, vx lans, and / or InfiniBand partitions.

61. The method according to any one of claims 57 to 60, wherein the storage authentication information includes a password or passphrase.

62. The method according to any one of claims 57 to 61, wherein the storage authentication information includes a chap key.

63. The method according to any one of claims 57 to 62, wherein the storage authentication information includes an encryption key.

64. The method according to claim 60, further comprising the computing resource calculating login authentication information of the storage resource from the encryption key based on an encryption technique. The method according to claim 60, further comprising:

65. The method according to any one of claims 57 to 64, wherein the storage authentication information includes a certificate.

66. The computer system further includes a plurality of computing resources and a plurality of storage resources, The provisioning step includes the controller provisioning storage authentication information of the storage resources to the computing resources such that a plurality of the computing resources are paired with different storage resources among the storage resources. The method according to any one of claims 57 to 65.

67. A controller, A resource, An in-band management connection for connecting the resource to the controller, A first connection for connecting the controller to a cloud instance, A second connection for connecting the cloud instance to the in-band management connection, An information technology (IT) computer system comprising: The controller is configured to provision the cloud instance via the first connection. The controller and / or the resource is configured to operably interact with the provisioned cloud instance via the second connection. The information technology (IT) computer system.

68. The system according to claim 67, wherein the cloud instance includes a cloud storage resource.

69. The system according to claim 68, wherein the controller is further configured to (1) maintain system state information, (2) create a storage bucket of the cloud storage resource, and (3) store connection information of the cloud storage resource as part of the system state information.

70. The system according to claim 69, wherein the resource includes a computing resource, and the controller provides the connection information of the cloud storage resource to the computing resource, whereby the computing resource is connected to the cloud storage resource and further configured to be able to use the cloud storage resource.

71. The system according to any one of claims 67 to 70, wherein the cloud instance is part of a pool of cloud instances managed by the controller as a cloud resource pool.

72. The system according to claim 71, wherein the controller is further configured to set up the resource as a cloud resource pool corresponding host.

73. The system according to claim 72, wherein the controller is further configured to (1) maintain system state information and (2) add information regarding cloud instances of the cloud resource pool to the system state information.

74. The system according to claim 73, wherein the controller is further configured to create a VPN connection to the cloud resource pool through the in-band management connection.

75. The system according to any one of claims 67 to 74, wherein the controller is further configured to identify the type of the cloud instance via the first connection.

76. The system according to claim 75, wherein the controller is further configured to identify the configuration of the cloud instance via the first connection.

77. The system according to any one of claims 75 to 76, wherein the controller is further configured to (1) maintain system rules and (2) determine whether to purchase and / or add the cloud instance to the system based on the system rules.

78. The system according to claim 77, wherein in response to a determination to purchase and / or add the cloud instance to the system, the controller is further configured to use a template to add the cloud instance to the system as another resource of the system.

79. The system according to claim 78, wherein the controller is further configured to turn on and / or activate the power of the cloud instance via the first connection.

80. The system according to claim 79, wherein the controller is further configured to find and load a boot image of the cloud instance from the template based on the identified cloud instance type.

81. The system according to claim 80, wherein the cloud instance is further configured to boot based on the boot image.

82. The system according to any one of claims 67 to 81, wherein the first connection connects the controller to the cloud instance via a cloud application programming interface (API).

83. The system according to claim 82, wherein the cloud API enables the controller to purchase and / or connect to the cloud instance. **Claim 84** The system according to any one of claims 67 to 83, wherein the second connection connects the cloud instance to the in-band management connection via a VPN. **Claim 85** An information technology (IT) method for use with a computer system including a controller, a resource, an in-band management connection connecting the resource to the controller, a first connection connecting the cloud instance to the in-band management connection, and a second connection connecting the cloud instance to the in-band management connection, comprising: the controller provisioning the cloud instance via the first connection; the controller and / or the resource operably interacting with the provisioned cloud instance via the second connection The method as described above. **Claim 86** The method according to claim 85, wherein the cloud instance includes a cloud storage resource. **Claim 87** The controller further comprising: (1) maintaining system state information; (2) creating a storage bucket for the cloud storage resource; and (3) storing connection information for the cloud storage resource as part of the system state information The method according to claim 86. **Claim 88** The resource includes a computing resource, and the method further includes: the controller providing the connection information for the cloud storage resource to the computing resource, thereby enabling the computing resource to connect to the cloud storage resource and use the cloud storage resource. The method according to claim 87. **Claim 89** The method according to any one of claims 85 to 88, wherein the cloud instance is part of a pool of cloud instances managed by the controller as a cloud resource pool. **Claim 90** The method according to claim 89, further comprising the controller setting up the resource as a cloud resource pool corresponding host. The method according to claim 89. **Claim 91** The controller further (1) maintains system state information and (2) adds information regarding cloud instances of the cloud resource pool to the system state information The method according to claim 90, further comprising the above **Claim 92** The controller creates a VPN connection to the cloud resource pool through the in-band management connection The method according to claim 91, further comprising the above **Claim 93** The controller identifies the type of the cloud instance via the first connection The method according to any one of claims 85 to 92, further comprising the above **Claim 94** The controller identifies the configuration of the cloud instance via the first connection The method according to claim 93, further comprising the above **Claim 95** The controller (1) maintains system rules and (2) determines whether to purchase and / or add the cloud instance to the system based on the system rules The method according to any one of claims 93 to 94, further comprising the above **Claim 96** In response to determining that the controller purchases and / or adds the cloud instance to the system, the controller adds the cloud instance to the system as another resource of the system using a template The method according to claim 95, further comprising the above **Claim 97** The controller turns on and / or activates the power of the cloud instance via the first connection The method according to claim 96, further comprising the above **Claim 98** The controller finds and loads a boot image of the cloud instance from the template based on the identified cloud instance type The method according to claim 97, further comprising the above **Claim 99** The cloud instance boots based on the boot image The method according to claim 98, further comprising the above **Claim 100** The method according to any one of claims 85 to 99, wherein the first connection connects the controller to the cloud instance via a cloud application programming interface (API) **Claim 101** The cloud API enables the purchase of the cloud instance by the controller and / or connection to the cloud instance The method according to claim 100, further comprising

102. The method according to any one of claims 85 to 101, wherein the second connection connects the cloud instance to the in-band management connection via a VPN

103. A controller, A memory, A plurality of resources for connecting to the controller, A plurality of services to be executed by at least one of the resources An information technology (IT) computer system comprising, the services include a first service and a second service, and the first and second services have a dependency relationship with each other The memory is configured to store a plurality of backup rules, and the backup rules include specifications of backup rules associated with the first service At least one of the controller or the resources is configured to (1) access the backup rules associated with the first service in the memory and (2) execute a backup operation of the first service according to the accessed backup rules associated with the first service. The backup rules associated with the first service define a cooperation with the second service that also backs up data related to the second service in cooperation with backing up data related to the first service The information technology (IT) computer system

104. The system according to claim 103, wherein the memory is further configured to store a plurality of templates, the templates include templates associated with the first service, and the templates associated with the first service include backup rules associated with the first service and / or pointers to backup rules associated with the first service

105. The resources include a plurality of computing resources, and at least one of the computing resources is configured to execute the backup operation The resources further include a plurality of storage resources The controller is further configured to provision at least one internal storage space of the storage resources for archiving the backed-up data therein for the backup operation. The system according to any one of claims 103 to 104. **Claim 106** The system according to claim 105, wherein the controller is further configured to provide access to the provisioned storage space for archiving the backed-up data to the first service according to the backup operation. **Claim 107** The system according to any one of claims 105 to 106, wherein the backup rule associated with the first service includes a plurality of rules defining how to provision the storage resources for the backup operation. **Claim 108** The system according to any one of claims 103 to 107, wherein the controller is further configured to start a recovery operation for the first service based on the backed-up data related to the first and second services. **Claim 109** The system according to any one of claims 103 to 108, wherein the backup rule further includes a specification of a backup rule associated with the second service, and the backup operation includes accessing and executing at least a part of the backup rule associated with the second service. **Claim 110** The system according to claim 109, wherein the backup rule associated with the first service includes a pointer to the backup rule associated with the second service. **Claim 111** The service further includes a third service, and the third service depends on the second service. The backup rule associated with the second service defines the cooperation with the third service for also backing up the data related to the third service in cooperation with backing up the data related to the first and second services. The system according to any one of claims 109 to 110. **Claim 112** The service further includes a third service, and the third service depends on the first service. The backup rule associated with the first service defines the cooperation with the third service that also backs up the data related to the third service in cooperation with backing up the data related to the first and second services. The system according to any one of claims 103 to 111.

113. The system according to any one of claims 103 to 112, wherein the backup rule associated with the first service includes recovery information associated with the first service.

114. An information technology (IT) method used with a computer system including a controller, a memory, a plurality of resources connected to the controller, and a plurality of services to be executed by at least one of the resources, wherein the resources include a first service and a second service, and the first and second services have a dependency relationship with each other. Storing in the memory a plurality of backup rules, including the specification of the backup rule associated with the first service. Accessing the backup rule associated with the first service in the memory. Executing a backup operation of the first service according to the accessed backup rule associated with the first service, wherein the backup rule associated with the first service defines the cooperation with the second service that also backs up the data related to the second service in cooperation with backing up the data related to the first service, and the backup operation is executed by the controller or at least one of the resources. The method.

115. The method according to claim 114, wherein the memory stores a plurality of templates, the templates include a template associated with the first service, the template associated with the first service includes the backup rule associated with the first service and / or a pointer to the backup rule associated with the first service, and the accessing step includes accessing the associated backup rule of the first service via the template associated with the first service.

116. The resource includes a plurality of computing resources, and at least one of the computing resources executes the backup operation. The resource further includes a plurality of storage resources. The method further includes the controller provisioning at least one internal storage space of the storage resources for archiving the backed-up data therein for the backup operation. The method according to any one of claims 114 to 115.

117. The controller provides access to the provisioned storage space to the first service for archiving the backed-up data in the provisioned storage space according to the backup operation. The method according to claim 116, further comprising.

118. The backup rule associated with the first service includes a plurality of rules that define how to provision the backup operation of the storage resources. The method according to any one of claims 116 to 117.

119. The controller starts a recovery operation for the first service based on the backed-up data related to the first and second services. The method according to any one of claims 114 to 118, further comprising.

120. The backup rule further includes a specification of a backup rule associated with the second service, and the backup operation includes accessing and executing at least a part of the backup rule associated with the second service. The method according to any one of claims 114 to 119.

121. The backup rule associated with the first service includes a pointer to the backup rule associated with the second service. The method according to claim 120.

122. The service further includes a third service, and the third service depends on the second service. The backup rule associated with the second service defines the cooperation with the third service that also backs up the data related to the third service in cooperation with backing up the data related to the first and second services. The method according to any one of claims 120 to 121.

123. The service further includes a third service, and the third service depends on the first service. A backup rule associated with the first service defines cooperation with the third service that also backs up data related to the third service in cooperation with backing up data related to the first and second services. The method according to any one of claims 114 to 122.

124. The method according to any one of claims 114 to 123, wherein a backup rule associated with the first service includes recovery information associated with the first service.

125. The system or method according to any one of claims 1 to 102, wherein the controller is further configured to manage backup operations of one or more components of the system, including services having a dependency relationship with each other.

126. The system or method according to claim 125, wherein the controller is further configured to manage backup operations according to any one of claims 103 to 124.

127. A system, apparatus, method, and / or computer program product including any combination of features disclosed herein.

Citation Information

Patent Citations

  • Communication interface device, program thereof, and virtual network construction method

    JP2013207784A

  • Address conversion device and node device

    JP2013223214A

  • Deployment system for multi-node applications

    JP2014514659A

  • Information setting device, information setting method, information setting program, storage medium, and radio communication system

    JP2015154445A

  • Identity services for organizations transparently hosted in the cloud

    JP2015518198A