Automatically deployed information technology (IT) systems and methods

Automatically manage IT systems through automated controllers and global system rules, solving the complexity, security risks and management difficulties of IT systems in the setup, configuration and deployment process, achieving higher security, reliability and management efficiency.

CN119989368APending Publication Date: 2025-05-13NET THUNDER LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510128943.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-07-06
Filing Date
2018-12-07
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing IT systems face complexity, security risks and management challenges during setup, configuration, and deployment, resulting in troubleshooting and recovery difficulties, and increased security measures may cause undesirable downtime.

Method used

By introducing an automation controller, the computer system is automatically managed by using global system rules, templates and IT system status to realize automated IT system settings, configuration, maintenance, testing, change management and upgrades.

Benefits of technology

Improves the security and reliability of IT systems, reduces human errors and configuration errors, simplifies the troubleshooting and recovery process, and reduces downtime and management costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989368A_ABST
    Figure CN119989368A_ABST
Patent Text Reader

Abstract

Systems, methods, and apparatus are disclosed herein in which a controller is capable of automatically managing a physical infrastructure of a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. Techniques are described for automatically adding resources, such as computer resources, storage resources, and / or networking resources, to the computer system. Techniques for automatically deploying applications and services on such resources are also described. The techniques provide a scalable computer system capable of serving as a one-stop scalable private cloud.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application 201880088766.6 filed on December 7, 2018, with the invention name “Automatically deployed information technology (IT) system and method”.

[0002] Cross-references and priority claims to related patent applications

[0003] This patent application claims priority to U.S. Provisional Patent Application No. 62 / 596,355, filed on December 8, 2017 and entitled “Automatically Deployed Information Technology (IT) System and Method,” the entire disclosure of which is incorporated herein by reference.

[0004] This patent application also claims priority to U.S. Provisional Patent Application No. 62 / 694,846, filed on July 6, 2018, and entitled “Automatically Deployed Information Technology (IT) System and Method,” the entire disclosure of which is incorporated herein by reference. Background Art

[0005] In recent decades, the demand, use, and need for computing have grown rapidly. The resulting need for greater storage, speed, computing power, applications, and accessibility has led to a rapidly changing computing landscape, providing tools for entities of all types and sizes. As a result, the use of public virtual computing and cloud computing systems has been developed to provide better computing resources to a large number of users and multiple types of users. This exponential growth is likely to continue. At the same time, greater failure and security risks have made infrastructure setup, management, change management, and updates more complex and expensive. Scalability, or the evolution of systems over time, has also become a major challenge in the field of information technology.

[0006] It may be difficult to diagnose and solve problems in most IT systems, many of which involve performance and security. Constraints on the time and resources allowed to set up, configure and deploy systems may cause errors and lead to future IT problems. Over time, many different administrators may be involved in changing, patching or updating IT systems including users, applications, services, security, software and hardware. Typically, the file records and history of configurations and changes may be insufficient or lost, making it difficult to understand how a particular system has been configured and how it works at a later time. This may make future changes or troubleshooting difficult. When problems or failures occur, it may be difficult to restore and reproduce IT configurations and settings. In addition, system administrators may easily make mistakes, such as incorrect commands or other errors, which in turn may paralyze computers and web databases and services. In addition, although the increased risk of security vulnerabilities has become commonplace, changes, updates, and patches made to avoid security vulnerabilities may cause unexpected downtime.

[0007] Once critical infrastructure is in place, working, and active, the costs or risks often appear to outweigh the benefits of changing the system. Problems involved in making changes to active IT systems or environments can cause extensive and sometimes catastrophic problems for users or entities that rely on these systems. At a minimum, the amount of time spent troubleshooting and fixing failures or problems that arise during change management can require significant time, personnel, and monetary resources. Technical problems that potentially arise when making changes to an active environment can have cascading effects and may not be resolved simply by undoing the changes made. Many of these challenges result in an inability to quickly rebuild systems if there are failures during change management.

[0008] In addition, bare metal cloud nodes or resources within an IT system may be vulnerable to security challenges, compromised, or accessed by rogue users. A hacker, attacker, or rogue user may divert from the node or resource to access or intrude into any other part of the IT system or network coupled to the node. A bare metal cloud node or controller of an IT system may also be attacked because of resources connected to an application network, which may expose the system to security threats or otherwise harm the system. According to various example embodiments disclosed herein, an IT system may be configured to improve the security of bare metal cloud nodes or resources that are connected to the Internet or application network, regardless of whether they are connected to an external network. Summary of the invention

[0009] According to an example embodiment, an IT system includes bare metal cloud nodes or physical resources. When bare metal cloud nodes or physical resources are powered on, configured, managed, or used, if they may be connected to a network where the nodes may be used by other people or customers, in-band management may be omitted, can be switched from a controller, can be disconnected from the controller, or can be filtered from the controller.

[0010] Physical resources including virtual machines or hyper-management systems may also be subject to security challenges, compromised or accessed by rogue users, where the hyper-management system may be used to transfer to another hyper-management system as a shared resource. An attacker may be able to break out of the virtual machine and may have network access to the management and / or administration system through the controller. According to various example embodiments disclosed herein, an IT system may be configured to improve security, where one or more physical resources including virtual resources on a cloud platform are disconnected from the controller via an in-band management connection, can be disconnected from the controller, filtered from the controller, can be filtered from the controller, or are not connected to the controller.

[0011] According to an example embodiment, the physical resources of an IT system may include one or more virtual machines or a hypervisor system, wherein an in-band management connection between a controller and the physical resources may be omitted, disconnected from the resources, capable of being disconnected from the resources, or filtered / capable of being filtered from the resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic diagram of a system according to an example embodiment.

[0013] Figure 2A is used for Figure 1 Schematic diagram of an example controller of a system.

[0014] Figure 2B An example flow of operation for an example set of storage expansion rules is shown.

[0015] Figure 2C and Figure 2D Shown is the method for executing Figure 2B An alternative example of steps 210.1 and 210.2 in .

[0016] Figure 2E An example template is shown.

[0017] Figure 2F An example process flow of controller logic with respect to processing templates is shown.

[0018] Figure 2G and Figure 2H Shown for Figure 2F An example process flow of steps 205.11, 205.12 and 205.13.

[0019] Fig.2I Another example template is shown.

[0020] Figure 2J Another example process flow of controller logic with respect to processing templates is shown.

[0021] Figure 2K An example process flow for managing service dependencies is shown.

[0022] Figure 2L is a schematic diagram of an example image derived from a template according to an example implementation.

[0023] Figure 2M A set of example system rules is shown.

[0024] Figure 2N shows the logic processed by the controller Figure 2M Example process flow for system rules.

[0025] Fig.2O An example process flow for configuring storage resources from a file system blob or other group of files is shown.

[0026] Figure 3A yes Figure 2A Schematic diagram of a controller with added computing resources.

[0027] Figure 3B is a schematic diagram of an example image derived from a template according to an example implementation.

[0028] Figure 3C An example process flow for adding resources, such as computing resources, storage resources, and / or networking resources, to a system is shown.

[0029] Figure 4A yes Figure 2A A diagram of a controller with storage resources added.

[0030] Figure 4B is a schematic diagram of an example image derived from a template according to an example implementation.

[0031] Figure 5A yes Figure 2A Figure 1 shows a diagram of a controller with JBOD and storage resources added.

[0032] Figure 5BAn example process flow for adding storage resources and direct attached storage for the storage resources to a system is shown.

[0033] Fig. 6A yes Figure 2A Figure 1 shows a diagram of a controller with networking resources added.

[0034] Figure 6B is a schematic diagram of an example image derived from a template according to an example implementation.

[0035] Fig. 7A is a schematic diagram of a system in an example physical deployment according to an example embodiment.

[0036] Figure 7B An example process for adding resources to an IT system is shown.

[0037] Figure 7C and Fig.7D An example process flow for deploying an application on multiple computing resources, multiple servers, multiple virtual machines, and / or in multiple sites is shown.

[0038] Fig. 8A is a schematic diagram of a system in an example deployment according to an example embodiment.

[0039] Figure 8B An example process flow for scaling from a single-node system to a multi-node system is shown.

[0040] Figure 8C An example process flow for migrating storage resources to new physical storage resources is shown.

[0041] Fig.8D An example process flow is shown for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for compute and storage.

[0042] Fig. 8E Another example process flow for scaling from a single node to multiple nodes in a system is shown.

[0043] Fig. 9A is a schematic diagram of a system in an example physical deployment according to an example embodiment.

[0044] Fig. 9B is a schematic diagram of an example image derived from a template according to an example implementation.

[0045] Fig. 9C An example of installing an application from an NT package is shown.

[0046] Fig.9Dis a schematic diagram of a system in an example deployment according to an example embodiment.

[0047] Fig.9E An example process flow for adding a virtual computing resource host to an IT system is shown.

[0048] Fig.10 is a schematic diagram of a system in an example deployment according to an example embodiment.

[0049] Fig.11A A system and method of an example embodiment is shown.

[0050] Fig. 11B A system and method of an example embodiment is shown.

[0051] Fig.12 A system and method of an example embodiment is shown.

[0052] Fig.13A is a schematic diagram of a system according to an example embodiment.

[0053] Fig. 13B is another schematic diagram of a system according to an example embodiment.

[0054] FIG. 13C to FIG. 13E An example process flow is shown for a system according to one example embodiment.

[0055] Fig.14A An example system is shown in which a master controller has deployed controllers on different systems.

[0056] Fig. 14B and Fig. 14C An example flow is shown demonstrating possible steps for supplying a controller using a master controller.

[0057] Fig.15A An example system in which a host controller generates an environment is shown.

[0058] Fig. 15B An example process flow for a controller to set up an environment is shown.

[0059] Fig. 15C An example process flow for a controller to set up multiple environments is shown.

[0060] Fig.16A An example embodiment is shown in which the controller operates as a master controller to set up one or more controllers.

[0061] FIG. 16B to FIG. 16D An example system is shown in which an environment can be configured to write to another environment.

[0062] Fig.16E An example system is shown in which a user can purchase new environments generated by a controller.

[0063] Fig.16F An example system is shown in which a user interface is provided for interfacing into an environment generated by a controller.

[0064] FIG. 17A to FIG. 18B An example of a change management task for a new environment is shown. DETAILED DESCRIPTION

[0065] In order to provide technical solutions for the needs in the art as described above, the inventors disclose various inventive embodiments related to systems and methods for information technology, which provide automated IT system setup, configuration, maintenance, testing, change management, and / or upgrades. For example, the inventors disclose a controller configured to automatically manage a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. As another example, the inventors disclose a controller configured to automatically manage a physical infrastructure of a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. Examples of automated management that can be performed by a controller may include: remotely or locally accessing and changing settings or other information on a computer that may be running an application or service; building an IT system; changing an IT system; building a separate stack in an IT system; creating a service or application; loading a service or application; configuring a service or application; migrating a service or application; changing a service or application; removing a service or application; cloning a stack to another stack on a different network; creating, adding, removing, setting, configuring, reconfiguring, and / or changing resources or system components; automatically adding, removing, and / or restoring resources, services, applications, IT systems, and / or IT stacks; configuring interactions between applications, services, stacks, and / or other IT systems; and / or monitoring the health of IT system components. In an example embodiment, a controller may be embodied as a physical or virtual computing resource that may be remote or local. Additional examples of controllers that may be employed include, but are not limited to, one or any combination of the following: processes, virtual machines, containers, remote computing resources, applications deployed by other controllers and / or services. Controllers may be distributed across multiple nodes and / or resources and may be in other locations or networks.

[0066] IT infrastructure is most often made up of discrete hardware and software components. The hardware components used generally include servers, racks, power supply equipment, interconnection devices, display monitors and other communication equipment. The methods and techniques of first selecting and then interconnecting these discrete components are highly complex, because a large number of optional configurations can change with varying degrees of efficiency, cost-effectiveness, performance and safety. The employment and training costs of individual technicians / engineers who are good at connecting these infrastructure components are very expensive. In addition, a large number of possible iterations of hardware and software can produce complexity in maintaining and updating hardware and software. When the individuals and / or engineering companies that initially installed the IT infrastructure were not convenient to perform updates, this would present additional challenges. Software components such as operating systems are generally designed to work on a wide range of hardware, or are completely dedicated to specific components. Complex plans or blueprints will be formulated and executed in most cases. Changes, developments, expansions and other challenges all require updating complex plans.

[0067] While some IT users purchase cloud computing services from industry-developing vendors, this does not solve the problems and challenges of setting up infrastructure, but rather shifts the problems and challenges from IT users to cloud service providers. In addition, large cloud service providers have solved the challenges and problems of setting up infrastructure in a manner that may sacrifice flexibility, customization, scalability, and rapid absorption of new hardware and software technologies. In addition, cloud computing services do not provide pre-set bare metal setup, configuration, deployment, and updates, or allow transitions to, from, or between bare metal and virtual IT infrastructure components. These and other limitations of cloud computing services may result in a number of computing, storage, and networking inefficiencies. For example, speed or latency inefficiencies in computing and networking may be presented by the cloud service or in applications or services that utilize the cloud service.

[0068] The system and method of an example embodiment provide a novel and unique IT infrastructure deployment, use and management. According to an example embodiment, the complexity of resource selection, installation, interconnection, management and update is rooted in the core controller system and its parameter files, templates, rules and IT system status. The system includes a set of self-assembly rules and operation rules, which are configured to make components self-assemble instead of requiring technicians to assemble, connect and manage. In addition, the system and method of an example embodiment allow the use of self-assembly rules to achieve greater customization, scalability and flexibility without the current typical external planning documents. The system and method also allow effective resource use and reuse.

[0069] Systems and methods are provided to improve many of the challenges and problems in current IT systems, whether physical or virtual in whole or in part. The system and method of an example embodiment allows for flexibility, reduces variability and human error, and provides a structure that has the potential to increase system security.

[0070] Although there may be some solutions to one or more problems in current IT systems, such solutions do not comprehensively solve a large number of problems as the example embodiments described herein solve. In addition, such existing solutions may solve specific problems while exacerbating other problems.

[0071] Some of the current challenges addressed include, but are not limited to, difficulties related to setup, configuration, infrastructure deployment, asset tracking, security, application deployment, service deployment, documentation on maintenance and compliance, maintenance, scaling, resource allocation, resource management, load balancing, software failures, updating / patching software and security, testing, recovering IT systems, change management, and hardware updates.

[0072] As used herein, IT systems may include, but are not limited to, servers, virtual and physical hosts, databases and database applications, including, but not limited to, IT services, business computing services, computer applications, customer-facing applications, web applications, mobile applications, back-end, case management, customer tracking, ticketing, business tools, desktop management tools, billing, email, documentation, compliance, data storage, backup and / or network management.

[0073] One problem that users may face before setting up an IT system is predicting infrastructure requirements. Users may not know how much storage, computing power, or other requirements will be needed initially or over time during development or changes. According to an example embodiment, the flexibility allowed by the IT system and infrastructure is that if the system needs to change, the self-deploying infrastructure (physical and / or virtual) of an example embodiment can be used to automatically add, remove, or reallocate from the infrastructure internally at a later time. Therefore, the challenge of predicting future requirements when setting up the system is solved in the following ways: provide the ability to add to the system using the global rules, templates, and system states of the system, and track changes in such rules, templates, and system states.

[0074] Other challenges may also be related to the following: correct configuration, consistency of configuration, interoperability and / or interdependence, which may include, for example, future incompatibility caused by changes to configured system elements or their configurations over time. For example, when an IT system is initially set up, there may be missing elements or some elements may fail to configure. In addition, for example, when an iteration of an element or infrastructure component is set up, consistency may be lacking between iterations. When changes to the system are made, it may be necessary to improve the configuration. In the case of future infrastructure changes, a difficult choice is presented between optimal configuration and flexibility. According to an example embodiment, when the system is first deployed, the configuration is self-deployed from a template to an infrastructure component using global system rules, so that the configuration is consistent, repeatable or predictable, thereby allowing optimal configuration. This initial system deployment can be completed on physical components, while subsequent components can be added or modified, and the subsequent components may or may not be physical. In addition, this initial system deployment can be completed on physical components, while subsequent environments can be cloned from physical structures, and the subsequent environments may or may not be physical. This allows the system configuration to be optimal while allowing minimally disruptive future changes.

[0075] During the deployment phase, there are usually challenges with the interoperability of bare metal and / or software-defined infrastructure. There may also be challenges with the interoperability of software with other applications, tools, or infrastructure. These challenges may include, but are not limited to, challenges caused by the deployed products coming from different vendors. The inventors disclose an IT system that can provide infrastructure interoperability regardless of whether it is bare metal, virtual structure, or any combination thereof. Therefore, interoperability (the ability of components to work together) can be built into the disclosed infrastructure deployment, where the infrastructure is automatically configured and deployed. For example, different applications may depend on each other, and they may exist on separate hosts. In order to allow such applications to interact with each other, the controller logic, templates, system states, and system rules described herein contain information and configuration instructions for configuring the interdependencies of applications and tracking the interdependencies. Therefore, the infrastructure features discussed herein provide a way to manage how each application or service communicates with each other. As an example, ensure that the email service communicates correctly with the authentication service; and / or ensure that the groupware service communicates correctly with the email service. Further, this management can go deep into the infrastructure level to allow tracking, for example, how computing resources communicate with storage resources. Otherwise, the complexity of IT systems will be O(n n ) method rises.

[0076] As disclosed, automatic deployment of resources does not require pre-configuration of operating system software, since controllers can be deployed based on global system rules, templates, and IT system state / system self-awareness. According to an example embodiment, users or IT professionals may not need to know whether the addition, allocation, or reallocation of resources will work together to ensure interoperability. According to an example embodiment, additional resources can be automatically added to the network.

[0077] Using applications requires many different resources, typically including computing, storage, and networking resources. It also requires interoperability of resources and system components, including knowledge of what is in place and running, as well as interoperability with other applications. Applications may need to connect to other services and obtain configuration files, and ensure that each component can work together correctly. Application configuration may therefore be time and resource intensive. If there are interoperability issues with other applications, application configuration may bring cascading effects to the rest of the infrastructure. This may cause operational interruptions or vulnerabilities. The inventors disclose automated application deployment for solving these problems. Therefore, as disclosed by the inventors, applications can use the understanding of what is going on on the system and intelligent configuration to read from IT system status, global system rules and templates to perform self-deployment. In addition, according to an example embodiment, pre-deployment testing of configurations can be performed using the change management features as described herein.

[0078] Another challenge addressed by the example embodiments relates to problems associated with broker configuration that may arise where it is desirable to switch to a different vendor or other tool. According to one aspect of an example embodiment, template conversion is provided between the rules and templates of the controller and application templates from a particular vendor. This allows the system to automatically change vendors of software or other tools.

[0079] Many security challenges are caused by misconfigurations, patch failures, and the inability to test patches before deployment. Security challenges often arise during the configuration phase of a setup. For example, a misconfiguration may expose sensitive applications to the Internet or allow forged emails from an electronic server. The inventors disclose a system setup that automatically configures to prevent attackers, avoid unnecessary exposure to attackers, and provide security engineers and application security architects with more knowledge of the system. Automation reduces security flaws caused by human error or misconfiguration. In addition, the disclosed infrastructure provides introspection between services, and can allow rule-based access and limit communications between services to only those communications that are actually needed. The inventors disclose a system and method that has the ability to safely test patches before deployment, for example, as discussed with respect to change management.

[0080] Documentation is often a problematic area for IT management. During setup and configuration, the main goal may often be to get the components to work together. Typically, this involves troubleshooting and trial and error processes, where sometimes it is difficult to know exactly what makes the system work. Although the exact commands as executed are usually recorded, the troubleshooting or trial and error processes that may have achieved a working system are often not well documented or even completely undocumented. Problems or deficiencies in documentation may create problems with audit trails and auditing. Documentation problems that arise may create problems in showing compliance. Typically, when building a system or its components, compliance challenges may not be well known. Only after setting up and configuring the IT system may the applicable compliance determinations be known. Therefore, documentation is crucial for auditing and compliance. The inventors disclose a system comprising a global system rule database, templates, and an IT system status database, which provides automatically recorded setup and configuration. Any configuration that occurs is recorded in the database. According to an example embodiment, the automatically recorded configuration provides an audit trail and can be used to show compliance. Inventory management can use automatically recorded and tracked information.

[0081] Another challenge caused by IT system setup, configuration and operation involves inventory management of hardware and software. For example, it is often important to know how many servers exist, whether the servers are up and still running, what the capabilities of the servers are, in which rack each server is, which power supplies are connected to which servers, what network cards and what network ports each server is using, in which IT system the components are operating, and many other important considerations. In addition to inventory information, passwords and other sensitive information used for inventory management should also be effectively managed. Especially in larger IT systems, data centers, or data centers where equipment changes frequently, the collection and retention of such information is a time-consuming task that is typically managed manually or using various software tools. Compliance protection of secure passwords is a big risk factor that can be a significant challenge in ensuring a secure computing environment. The inventors disclose an IT system in which the collection and maintenance of inventory and operating conditions of all servers and other components is automatically updated, stored, and protected as part of the controller's IT system status, global system rules, templates, and controller logic.

[0082] In addition to the problems of setting up and configuring IT systems, the inventors have also disclosed an IT system that can also solve problems and difficulties that arise in the maintenance of IT systems. Many problems will occur when a data center is continuously running in the presence of hardware failures, such as power failures, memory failures, network failures, network card failures and / or CPU failures, as well as other failures. Other failures will occur when migrating hosts during hardware failures. Therefore, the inventors disclose dynamic resource migration, such as migrating resources from one resource provider to another when a host goes down. In this case, according to an example embodiment, the IT system can migrate to other servers, nodes or resources, or other IT systems. The controller can report the status of the system. A copy of the data is on another host with a known and automatically set configuration. If a hardware failure is detected, any resource that may have been providing the hardware can be automatically migrated after the failure is automatically detected.

[0083] A major challenge for many IT systems is scalability. Growing businesses or other organizations often add or reconfigure their IT systems as they grow and their needs change. Problems arise when existing IT systems require more resources, such as adding hard drive space, storage space, CPU processing, more network infrastructure; more endpoints, more clients and / or more security measures. Problems also arise in configuration, setup and deployment when different services and applications or changes to the infrastructure are required. According to an example embodiment, a data center can be automatically expanded. Nodes or resources can be dynamically and automatically added to or removed from a resource pool. Resources added and removed from a resource pool can be automatically allocated or reallocated. Services can be quickly supplied and moved to new hosts. A controller can dynamically detect and add more resources to a resource pool and know where to allocate / reallocate resources. A system according to an example embodiment can be expanded from a single-node IT system to an expanded system that requires numerous physical and / or virtual nodes or resources across multiple data centers or IT systems.

[0084] The inventors disclose a system for implementing flexible resource allocation and management. The system includes computing resources, storage resources, and networking resources that may be in a resource pool and can be automatically allocated. A controller can identify new nodes or hosts on the network and then configure the new nodes or hosts so that they can become part of the resource pool. For example, whenever a new server is inserted, the controller will configure the new server as part of the resource pool, and the new server can be added to the resources and can begin to use the new server automatically. Nodes or resources can be detected by the controller and added to different pools. Resource requests can be issued to the controller, for example, through an API request. The controller can then deploy or allocate the required resources from the pool according to the rules. This allows the controller and / or applications to balance the load and dynamically distribute resources based on the needs of the requests through the controller.

[0085] Examples of load balancing include, but are not limited to: deploying new resources when hardware or software failures occur; deploying one or more instances of the same application in response to increased user load; and deploying one or more instances of the same application in response to an imbalance in storage, computing, or networking demand.

[0086] The problems involved in making changes to active IT systems or environments can cause extensive and sometimes catastrophic problems for users or entities that rely on these systems to continue to operate. These outages represent not only a potential loss of system use, but also a loss of data, financial loss due to the large amount of time, personnel and financial resources required to repair the problems. The problems can be exacerbated by the difficulty of rebuilding the system if there are errors in the documentation of the configuration or a lack of understanding of the system. Because of this problem, many IT system users are reluctant to patch IT resources to eliminate known security risks. The resources are therefore more vulnerable to security vulnerabilities.

[0087] Many problems that arise in the maintenance of IT systems are related to software failures due to change management or control that may require configuration. Situations where such failures may occur include, but are not limited to: upgrading to a new software version; migrating to a different piece of software; password or authentication management changes; switching between services or between different providers of services.

[0088] Manually configured and maintained infrastructure is often difficult to recreate. Recreating infrastructure can be important for several reasons, including but not limited to: undoing problematic changes, power outages or other disaster recovery. Problems in manually configured systems are difficult to diagnose. Manually configured and maintained infrastructure is difficult to recreate. In addition, system administrators can easily make mistakes, such as incorrect commands, which have been known to paralyze computer systems.

[0089] Making changes to active IT systems or environments can cause extensive and sometimes catastrophic problems for users or entities that rely on these systems to continue to function. These outages not only represent a potential loss of system use, but such outages can also cause data loss, as well as financial losses due to the significant time, personnel, and financial resources required to repair the problems. The problems can be exacerbated by the difficulty of rebuilding the system in the presence of errors in the documentation of the configuration or a lack of understanding of the system. Additionally, in many cases, it is difficult to restore the system to a previous state after a significant or major change has occurred.

[0090] Furthermore, potential technical problems that arise when changes are made to a live environment may have cascading effects. These cascading effects may make it challenging and sometimes impossible to roll back the state before the change. Thus, even when changes need to be rolled back due to problems with the implemented changes, the state of the system has already been changed. It has recently been pointed out that undoing infrastructure and system management errors and erroneous changes to a production environment is an unresolved problem. Additionally, it is known that testing changes to a system before deployment to a live environment is problematic.

[0091] Therefore, the inventors disclose multiple example embodiments of systems and methods configured to restore changes to an active system back to a state before the change. In addition, the inventors disclose (provide) a system and method configured to achieve substantial restoration of the state of a system or environment that is subject to real-time changes, which can prevent or ameliorate one or more of the problems described above.

[0092] According to a variation of an example embodiment, the IT system has complete system knowledge of global system rules, templates, and IT system states. The infrastructure can be cloned using complete system knowledge. The system or system environment can be cloned as a software-defined infrastructure or environment. The system environment (referred to as the production environment) including the volatile database in use can be written to a non-volatile read-only database for use as a development environment during development and testing. The required changes can be made in the development environment and the required changes can be tested. The user or controller logic can make changes to the global rules to create a new version. The version of the rule can be tracked. According to another aspect of an example embodiment, the newly developed environment can be automatically implemented afterwards. The previous production environment can also be maintained or made fully functional, so that the correction of the production environment in the earlier state is possible without losing data. The development environment can then be started with new specifications, rules, and templates, and the database or system can be synchronized with the production database, and the database or system can be switched to a writable database. The original production database can then be switched to a read-only database, and if recovery is required, the system can be restored to the read-only database.

[0093] With respect to upgrading or patching software, if it is detected that a service needs upgrading or patching, a new host can be deployed. In the event of a failure due to upgrading or patching, a new service can be deployed when change recovery is possible as described above.

[0094] Hardware upgrades are important in many situations, especially where the latest hardware is essential. An example of this type of situation occurs in the high-frequency trading industry, where an IT system with a millisecond speed advantage can enable users to achieve superior trading results and profits. In particular, problems arise in ensuring interoperability with the current infrastructure, so that the new hardware will know how to communicate with the protocol and work with the existing infrastructure. In addition to ensuring interoperability of the components, the components will also need to be integrated with the existing setup.

[0095] refer to Figure 1 , shows an IT system 100 of an example embodiment. System 100 may be one or more types of IT systems, including but not limited to those described herein.

[0096] The user interface (UI) 110 is shown coupled to the controller 200 via an application programming interface (API) application 120 that may or may not reside on a separate physical or virtual server. The controller 200 may be deployed on one or more processors and one or more memories to implement any of the control operations discussed herein. Instructions for execution by one or more processors to implement such control operations may reside on a non-transitory computer-readable storage medium such as a processor memory. The API 120 may include one or more API applications that may be redundant and / or operate in parallel. The API application 120 receives a request to configure system resources, parses the request and passes the request to the controller 200. The API application 120 receives one or more responses from the controller, parses one or more responses and passes the one or more responses to the UI (or application) 110. Alternatively or additionally, an application or service may communicate with the API application 120. The controller 200 is coupled to one or more computing resources 300, one or more storage resources 400, and one or more networking resources 500. The resources 300, 400, 500 may or may not reside on a single node. One or more of the resources 300, 400, 500 may be virtual. The resources 300, 400, 500 may or may not reside on multiple nodes or reside on multiple nodes in various combinations. A physical device may include one or more or each of the resource types, including but not limited to computing resources 300, storage resources 400, and networking resources 500. Resources 300, 400, 500 may also include resource pools, whether or not in different physical locations, and whether or not virtual. Bare metal computing resources may also be used to implement the use of virtual or container computing resources.

[0097] In addition to the known definition of a node, a node as used herein may be any system, device or resource connected to one or more networks or other functional units that perform functions on a standalone device or a network-connected device. A node may also include, but is not limited to, for example, a server, a service / application / multiple services on a physical or virtual host, a virtual server, and / or multiple or single services running on a multi-tenant server or inside a container.

[0098] The controller 200 may include one or more physical or virtual controller servers that may also be redundant and / or operate in parallel. The controller may run on a physical or virtual host used as a computing host. As an example, the controller may include a controller running on a host that is used for other purposes, such as because it has access to sensitive resources. The controller receives requests from the API application 120, parses the request and makes appropriate task assignments to other resources and instructs the other resources; monitors resources and receives information from the resources; maintains the status and change history of the system; and may communicate with other controllers in the IT system. The controller may also include the API application 120.

[0099] Computing resources as defined herein may include a real or virtual single computing node or a resource pool with one or more computing nodes. Computing resources or computing nodes may include one or more physical or virtual machines or container hosts that can host one or more services or run one or more applications. Computing resources may also be on hardware designed for multiple purposes, including but not limited to: computing, storage, caching, networking, specialized computing, including but not limited to: GPU, ASIC, coprocessor, CPU, FPGA and other specialized computing methods. PCI Express switches or similar devices may be added to such devices, and the devices may be dynamically added in this way. Computing resources or computing nodes may include or may run multiple different virtual machines that include running services or applications, or may be one or more super management systems or container hosts for virtual computing resources. Although the focus of computing resources may be on providing computing functions, it may also include data storage and / or networking capabilities.

[0100] Storage resources as defined herein may include storage nodes or storage resource pools. Storage resources may include any data storage media, such as fast, slow, hybrid, cache storage media and / or RAM. Storage resources may include one or more types of networks, machines, devices, nodes, or any combination thereof that may be directly attached or not directly attached to other storage resources. According to aspects of an example embodiment, storage resources may be bare metal or virtual resources, or a combination thereof. Although the focus of storage resources may be on providing storage functionality, it may also include computing and / or networking capabilities.

[0101] One or more networking resources 500 may include a single networking resource, multiple networking resources, or a pool of networking resources. One or more networking resources may include one or more physical or virtual devices, one or more tools, switches, routers, or other interconnectors between system resources, or applications for managing networking. Such system resources may be physical or virtual, and may include computing resources, storage resources, or other networking resources. Networking resources may provide connections between external networks and application networks and may host core network services, including but not limited to: DNS, DHCP, subnet management, third-layer routing, NAT, and other services. Some of these services may be deployed on computing resources, storage resources, or networking resources on physical or virtual machines. Networking resources may utilize one or more architectures or protocols, including but not limited to: InfiniBand, Ethernet, RoCE, Fibre Channel, and / or Omnipath, and may include interconnectors between multiple architectures. Networking resources may or may not have SDN capabilities. The controller 200 may be able to directly change the networking resources 300 using VLANs or the like of SDN to configure the topology of the IT system. Although the focus of the networking resources may be on providing networking functions, it may also include computing and / or storage capabilities.

[0102] As used herein, an application network refers to a networked resource for connecting or coupling applications, resources, services, and / or other networks, or for coupling users and / or clients to applications, resources, and / or services, or any combination thereof. An application network may include a network used by a server to communicate with other application servers (physical or virtual) and to communicate with clients. An application network may communicate with machines or networks external to system 100. For example, an application network may connect a web front end to a database. A user may connect to a web application through the Internet or another network that may or may not be managed by a controller.

[0103] According to an example embodiment, computing resources 300, storage resources 400, and networking resources 500 may each be automatically added, removed, set, allocated, reallocated, configured, reconfigured, and / or deployed by controller 200. According to an example embodiment, additional resources may be added to a resource pool.

[0104] Although a user interface 110 is shown, such as a Web UI or other user interface that a user 105 can utilize to access and interact with the system, alternatively or in addition, an application may communicate or interact with the controller 200 through one or more API applications 120 or in other ways. For example, a user 105 or an application may send a request including, but not limited to: building an IT system; building a separate stack in an IT system; creating a service or application; migrating a service or application; changing a service or application; removing a service or application; cloning a stack to another stack on a different network; creating, adding, removing, setting or configuring, reconfiguring resources or system components.

[0105] Figure 1 The system 100 may include a server having connectors or other communication interfaces to various elements, components or resources that may be physical or virtual, or any combination thereof. According to one variation, Figure 1 The illustrated system 100 may include a bare metal server with connectors.

[0106] As described in more detail herein, the controller 200 may be configured to drive resources or components, automatically set, configure and / or control the activation of resources, add resources, allocate resources, manage resources and update available resources. The drive process may begin with the drive controller so that the order of activating the device can be consistent, rather than depending on the user to drive the device. The process may also involve detection of the resources that have been driven.

[0107] refer to Figures 2A to 10 , showing a controller 200 , a controller logic 205 , a global system rules database 210 , an IT system status 220 , and a template 230 .

[0108] The system 100 includes global system rules 210. The global system rules 210 may, among other things, declare rules for setting up, configuring, starting, allocating, and managing resources that may include computing, storage, and networking. The global system rules 210 include minimum requirements for the system 100 to be in a correct or desired state. The requirements may include IT tasks that are expected to be completed, and an updateable list of expected hardware required to predictably build the required system. The updateable list of expected hardware may allow the controller to verify that the required resources (starting from a start rule or using a template, for example, before starting a rule or using a template) are available. The global rules may include a list of operations required for various tasks and corresponding instructions related to the sequencing of operations and tasks. For example, the rules may specify the order of: driver components; startup resources, applications, and services; dependencies; the order in which different tasks, such as loading, configuring, starting, reloading applications, or updating hardware, are started. The rules 210 may also include one or more of the following: for example, a list of resource allocations required by applications and services; a list of templates that can be used; a list of applications to be loaded and how to configure them; a list of services to be loaded and how to configure them; a list of application networks and which application runs using which network; a list of configuration variables specific to different applications and user-specific application variables; an expected state that allows the controller to check the system state to verify that the state is as expected and the results of each instruction are as expected; and / or a version list that includes a list of rule changes (e.g., snapshots) that can allow tracking of changes to the rules and the ability to test or revert to different rules in different situations. The controller 200 can be configured to apply the global system rules 210 to the IT system 100 on physical resources. The controller 200 can be configured to apply the global system rules 210 to the IT system 100 on virtual resources. The controller 200 can be configured to apply the global system rules 210 to the IT system 100 on a combination of physical and virtual resources.

[0109] Figure 2M A set of example system rules 210 that may take the form of global system rules is shown. Figure 2M The example set of system rules 210 shown may be loaded into the controller 200 or obtained by querying the system status (see 210.1). Figure 2MIn the example of , the system rules 210 include a set of instructions that may take the form of a configuration routine 210.2, and also include data 210.3 for creating and / or recreating an IT system or environment. Configuration rules within the system rules 210 may set forth how to locate templates 230 (wherein the templates 230 may reside in a file system, disk, storage resource, or may be located internal to the system rules) via a required template list 210.7. The controller logic 205 may also locate the templates 230 before processing them and ensure that the templates exist before enabling the system rules 210. The system rules 210 may include system rule subsets 210.15, and these subsets 210.15 may be executed as part of the configuration routine 210.2.

[0110] Additionally, subsystem rules 210.15 can be used, for example, as a tool to build systems with integrated IT applications (which are then processed using system rule execution routines 210.16 and then updated to reflect the addition of 210.15 to the system state and current configuration rules). Subsystem rules 210.15 can also be located elsewhere and loaded into system state 220 through user interaction. For example, you can also treat subsystem rules 210.15 as a playbook, and may be able to obtain and run the playbook (which then causes global system rules 210 to update, so you can replay the playbook if you want to clone a system).

[0111] The configuration routine 210.2 may be a set of instructions for building a system. The configuration routine 210.2 may also include subsystem rules 210.15 or system state pointers 210.8, if desired by the practitioner. When running the configuration routine 210.2, the controller logic 205 may process a series of templates in a specific order (210.9), optionally allowing parallel deployment, but maintaining proper dependency handling (210.12). The configuration routine 210.2 may optionally call API calls 210.10, which may set configuration parameters 210.5 on the application that may be configured by processing the templates according to 210.9. Additionally, required services 210.11 are services that need to be started and running if the system is to make API calls 210.10.

[0112] Routine 210.2 may also include a process, program or method for performing data loading (210.13) with respect to volatile data 210.6, including but not limited to: copying data, transferring a database to a computing resource, pairing a computing resource with a storage resource and / or updating system state 220 with the location of volatile data 210.6. Volatile data pointers (see 210.4) may be maintained with data 210.3 to locate volatile data that may be stored elsewhere. If configuration parameters 210.5 are located at a non-standard data storage location (e.g., contained in a database), data loading routine 210.13 may also be used to load the configuration parameters.

[0113] The system rules 210 may also include a resource list 210.18 that may indicate which components are allocated to which resources and will allow the controller logic 205 to determine if the appropriate resources and / or hardware are available. The system rules 210 may also include an alternative hardware and / or resource list 210.19 for alternative deployments (e.g., for development environments where software engineers may want to perform real-time testing but do not want to allocate an entire data center). The system rules may also include a data backup / standby routine 210.17 that provides instructions on how to back up the system and use spare parts to achieve redundancy.

[0114] After each action is taken, the system state 220 may be updated and the query (which may include a write) saved as system state query 210.14.

[0115] Figure 2N The processing by the controller logic 205 is shown Figure 2M Example process flow of system rule 210 (or subsystem rule 210.15). At step 210.20, controller logic 205 checks to ensure that appropriate resources are available (see Figure 2M Otherwise, alternative configurations may be checked at step 210.21. A third option may include prompting the user to select a configuration that may be affected Figure 2M List 210.7 of the supported alternative configurations for template 230 referenced in the example.

[0116] At step 210.22, the controller logic may then ensure that the computing resource (or any appropriate resource) gains access to the volatile data. This may involve connecting to or adding a storage resource to the system state 220. At step 210.23, the configuration routines are then processed, and as each routine is processed, the system state 220 is updated (step 210.24). The system state 220 may also be queried to check whether certain steps are completed before proceeding (step 210.25).

[0117] The configuration routine processing steps as shown in Figure 210.23 may include any of the procedures (or combinations thereof) of 210.26. The configuration routine processing steps may also include other procedures. For example, the processing at 210.26 may include template processing (210.27), loading configuration data (210.28), loading static data (210.29), loading dynamic volatile data (210.30) and / or coupling services, applications, subsystems and / or environments (210.31). Such procedures within 210.26 may be repeated in a loop, or run in parallel, as some system components may be independent, while other system components may be interdependent. Controller logic, service dependencies and / or system rules may indicate which services may be dependent on each other, and the services may be coupled to further build an IT system from the system rules.

[0118] The global system rules 210 may also include storage expansion rules. Storage expansion rules provide a set of rules for automatically adding storage resources to existing storage resources, such as within the system. In addition, storage expansion rules may provide trigger points by which applications running on one or more computing resources will know when to request storage expansion (or the controller 200 may know when to expand the storage of computing resources or applications). The controller 200 may allocate and manage new storage resources, and may merge or integrate the storage resources with existing storage resources for specific operating resources. Such specific operating resources may be, but are not limited to: computing resources within the system, applications running computer resources within the system, virtual machines, containers, or physical or virtual computing hosts or combinations thereof. The operating resources may signal the controller 200, such as through a storage space query, that the operating resources are running out of storage space. In-band management connection 270, SAN connection 280, or any networking or coupling to the controller 200 may be used in such queries. Out-of-band management connection 260 may also be used. These storage expansion rules (or a subset of these storage expansion rules) may also be used for resources that are not operating.

[0119] Storage expansion rules indicate how to locate, connect, and set up new storage resources within the system. The controller registers the new storage resources in the system state 220 and tells the operating resources where and how to connect to the storage resources. The operating resources use this registration information to connect to the storage resources. The controller 200 can merge the new storage resources with existing storage resources, or the controller can add the new storage resources to the volume group.

[0120] Figure 2BAn example flow of the operation of a set of example storage expansion rules is shown. At step 210.41, the operational resource determines that its storage is low based on a trigger point or other aspects. At step 210.42, the operational resource is connected to the controller 200 via an in-band management connection 270, a SAN connection 280, or another type of connection visible to the operating system. Through this connection, the operational resource can notify the controller 200 that its storage is low. At step 210.43, the controller configures the storage resource to expand the storage capacity for the operational resource. At step 210.44, the controller provides the operational resource with information about where the newly configured storage resource is located. At step 210.45, the operational resource is connected to the newly configured storage resource. At step 210.46, the controller adds a mapping of the new storage resource location to the system state 220. The controller can then add the new storage resource to the volume group assigned to the operational resource (step 210.47), or the controller can add the allocation of the new storage resource to the operational resource to the system state 220 (step 210.48).

[0121] Figure 2C Shown is the method for executing Figure 2B . At step 210.50, the controller sends a critical command via the out-of-band management connection 260 to view storage status updates about the running resource on a monitor or console. For example, the monitor may be an ipmi console that can be utilized to view the screen via the out-of-band connection 260. As an example, the out-of-band connection 260 may be plugged into a USB as a keyboard / mouse and into a VGA monitoring port. At step 210.51, the running resource displays information on the screen. At step 210.52, the controller then reads the information presented on the monitor or console via the out-of-band management connection 260 and screen scraping or similar operation; wherein this read information may indicate a low storage condition based on a trigger point. The process flow may then continue Figure 2B Step 210.43.

[0122] Figure 2D Shown is the method for executing Figure 2B Another alternative example of steps 210.41 and 210.42. At step 210.55, the operating resource automatically displays information on a monitor or console for the controller to read. At step 210.56, the controller automatically, periodically or continuously reads the monitor or console to check the operating resource. In response to this reading, the controller learns that the storage of the operating resource is low (step 210.57). The process flow can then continue Figure 2B Step 210.43.

[0123] The controller 200 also includes a library of templates 230 that may include bare metal and / or service templates. These templates may include, but are not limited to: email, file storage, IP telephony, software billing, software XMPP, wiki, version control, account authentication management, and third-party applications that may be configurable by a user interface. Templates 230 may be associated with a resource, application, or service; and the templates may be used as recipes that define how such resources, applications, or services are integrated into the system.

[0124] Likewise, a template may include a set of information for creating, configuring and / or deploying resources, or for establishing an application or service loaded on a resource. Such information may include, but is not limited to: a kernel, an initrd file, a file system or file system image, a file, a configuration file, a configuration file template, information for determining appropriate settings for different hardware and / or computing backends, and / or other options that can be used to configure resources to drive an application and allow and / or facilitate the creation, startup or running of an operating system image for an application.

[0125] The template may contain information that can be used to deploy the application on a variety of supported hardware types and / or computing backends, including but not limited to: multiple physical server types or components, multiple hypervisors running on multiple hardware types, and container hosts that can be hosted on multiple hardware types.

[0126] The template can obtain the boot image of the application or service running on the computing resource. The template and the image derived from the template can be used to create the application, deploy the application or service and / or arrange the resource for various system functions, which allows and / or helps to create the application. The template may have variable parameters in the file, file system and / or operating system image, which can be overwritten by the configuration options from the default setting or the setting given by the controller. The template may have a configuration script for configuring the application or other resources, and the template may use configuration variables, configuration rules and / or default rules or variables; these scripts, variables and / or rules may contain specific rules, scripts or variables for specific hardware, or other resource-specific parameters, such as super management system (in the case of virtualization), available memory specific parameters. The template may have a file in the form of a binary resource, a compilable source code that brings binary resources or hardware or other resource-specific parameters, multiple sets of specific binary resources, or source code with compilation instructions for specific hardware or other resource-specific parameters, such as super management system (in the case of virtualization), available memory specific parameters. The template may include a set of information that is irrelevant to what is running on a certain resource.

[0127] The template may include a base image. The base image may include a base operating system file system. The base operating system may be read-only. The base image may also include basic tools of the operating system that are independent of what is running. The base image may include a base directory and operating system tools. The template may include a kernel. The kernel or multiple kernels may include initrd, or multiple kernels may be configured for different hardware types and resource types. The image may be derived from a template ad loaded into one or more resources or deployments. The loaded image may also include startup files, such as a kernel or initrd corresponding to the template.

[0128] The image may include template file system information that can be loaded into a resource based on a template. The template file system can configure an application or service. The template file system may include a shared file system that is shared by all resources, or similar resources, for example, to save storage space for storing the file system or to facilitate the use of read-only files. The template file system or image may include a set of files shared by the deployed services. The template file system may be preloaded onto the controller or downloaded. The template file system may be updated. The template file system may allow relatively faster deployment because it does not need to be rebuilt. Sharing the file system with other resources or applications may allow for reduced storage because files are not copied unnecessarily. This may also allow for easier recovery from failures because only files different from the template file system need to be restored.

[0129] The template boot file may include a kernel and / or initrd or similar file system for assisting the boot process. The boot file may boot the operating system and set up the template file system. The initrd may include a small temporary file system that explains how to set up the template so that the template can boot.

[0130] The template may also include template BIOS settings. Template BIOS settings can be used to set optional settings for running applications on the physical host. If used, then as described in this article about Figures 1 to 12The described out-of-band management 260 can be used to start resources or applications. The physical host can use the out-of-band management network 260 or CDROM to start resources or applications. The controller 200 can set application-specific bios settings defined in such a template. The controller 200 can use the out-of-band management system to make direct bios changes through APIs specific to specific resources. The settings can be verified by console and image recognition. Therefore, the controller 200 can use console features and make bios changes using a virtual keyboard and mouse. The controller can also use a UEFI shell and can directly input into the console, and can use image recognition to verify successful results, correctly enter commands and ensure successful setting changes. If there is a bootable operating system available for BIOS changes or updates to a specific BIOS version, the controller 200 can remotely load a disk image or ISO to start an application running an operating system that updates the BIOS and allows configuration changes to be made in a reliable manner.

[0131] A template may also include a list of template-specific supporting resources, or a list of resources required to run a specific application or service.

[0132] The template image or portions of the image or template may be stored on the controller 200 , or the controller 200 may move or copy them to the storage resource 410 .

[0133] Figure 2E An example template 230 is shown. The template contains all the information needed to create an application or service. The template 230 may also contain information, alternative data, files, binaries for different hardware types that provide similar or identical functionality. For example, there may be file system blobs 232 for / usr / bin and / bin and binaries 234 compiled for different architectures. The template 230 may also include a daemon 233 or script 231. The daemon 233 is a binary or script that may run at boot time when the host is driven and ready; and in some cases, the daemon 233 may drive an API that may be accessible by the controller and may allow the controller to change the settings of the host (and the controller may then update the active system rules). The daemons may also be shut down and restarted by out-of-band management 260 or in-band management 270 discussed above and below. These daemons may also drive a generic API to provide dependency services for new services (e.g., a generic web server API that communicates with the API that controls nginx or apache). The script 231 may be an installation script that may be run when or after the image is booted, or after a daemon is started or a service is enabled.

[0134] The template 230 may also include a kernel 235 and a pre-boot file system 236. The template 230 may also include multiple kernels 235 and one or more pre-boot file systems (such as initrd or initramfs for Linux, or ramdisk for bsd) for different hardware and different configurations. Initrd may also be used to mount the file system blob 232 presented as an overlay, and mount the root file system on remote storage by booting into initramfs 236, which may optionally be connected to storage resources via SAN connection 280 as described below.

[0135] File system blob 232 is a file system image that can be divided into separate blobs. Blobs may be interchangeable based on configuration options, hardware type, and other setup differences. A host booted from template 230 can boot from a union file system (such as overlayfs) containing multiple blobs, or from an image created from one or more file system blobs.

[0136] Template 230 may also include or be linked to additional information 237, such as volatile data 238 and / or configuration parameters 239. For example, volatile data 238 may be included in template 230, or the volatile data may be included externally. The volatile data may be in the form of a file system blob 232 or other data store, including but not limited to: a database, a flat file, a file stored in a directory, a file archive, a git or other version control repository. In addition, configuration parameters 239 may be included externally or internally to template 230, and optionally included in system rules and applied to template 230.

[0137] The system 100 also includes an IT system state 220 that tracks, maintains, changes, and updates the status of the system 100, including but not limited to resources. The system state 220 can track available resources, which will tell the controller logic whether resources are available to implement rules and templates, and what resources are available to implement rules and templates. The system state can track the resources used, which allows the controller logic 205 to check efficiency, utilization efficiency, and whether resources need to be switched for upgrades or other reasons, such as to improve efficiency or achieve priority. The system state can track what applications are running. The controller logic 205 can compare the expected application operation with the actual application operation based on the system state, and whether corrections are needed. The system state 220 can also track where the application is running. The controller logic 205 can use this information for the purpose of evaluating efficiency, change management, updating, troubleshooting, or audit trails. The system state can track networking information, such as what network is running or currently running, or track configuration values ​​and history. The system state 220 can track change history. The system state 220 can also track which templates are used in which deployment based on global system rules, which specify which templates are used. The history can be used for auditing, alerting, change management, building reports, tracking versions and configurations or configuration variables associated with hardware and applications. The system state 220 can maintain configuration history for the purpose of auditing, compliance testing, or troubleshooting.

[0138] The controller has logic 205 for managing all information contained in the system state, templates, and global system rules. The controller logic 205, global system rules database 210, IT system state 220, and templates 230 are managed by the controller 200 and may or may not reside on the controller 200. The controller logic or application 205, global system rules database 210, IT system state 220, and templates 230 may be physical or virtual, and may or may not be distributed services, distributed databases, and / or files. The API application 120 may be included in the controller logic / controller application 205.

[0139] The controller 200 may run as a standalone machine and / or may include one or more controllers. The controller 200 may include a controller service or application and may run inside another machine. The controller machine may first start the controller service to ensure an orderly and / or consistent startup of the entire stack or set of stacks.

[0140] Controller 200 may control computing resources, storage resources, and networking resources of one or more stacks. Each stack may or may not be controlled by a different subset of rules within global system rules 210. For example, there may be a pre-production stack, a production stack, a development stack, a test stack, a parallel stack, a backup stack, and / or other stacks with different functions within the system.

[0141] The controller logic 205 may be configured to read and interpret global system rules to achieve the desired IT system state. The controller logic 205 may be configured to use templates to build system components, such as applications or services, according to the global rules, and allocate, add or remove resources to achieve the desired IT system state. The controller logic 205 may read the global system rules, generate a task list to achieve the correct state, and issue instructions for fulfilling the rules based on available operations. The controller logic 205 may include logic for the following: performing operations, such as starting the system, adding, removing, reconfiguring resources; identifying what can be done. The controller logic may check the system state at the start time and at regular time intervals to see if the hardware is available, and if available, the hardware can perform the task. If the necessary hardware is not available, the controller logic 205 uses the global system rules 210, templates 220, and the hardware available according to the system state 230 to present alternative options, and modify the global rules and / or system state 220 accordingly.

[0142] The controller logic 205 may know what variables are needed, what the user needs to enter to proceed, or what the user needs in the system to run. The controller logic can use a list of templates from the global system rules and compare the template list with the templates required in the system state to ensure that the required templates are available. The controller logic 205 can discern from the system state database whether the resources on the template-specific supported resource list are available. The controller logic can allocate resources, update the state and enter the next set of tasks to implement the global rules. The controller logic 205 can start / run the application on the allocated resources as specified in the global rules. The rules can specify how to build an application from a template. The controller logic 205 can grab one or more templates and configure the application based on the variables. The template can tell the controller logic 205 which kernel, which startup files, which file systems, and which supported hardware resources are required. The controller logic 205 can then add information about the application deployment to the system state database. After each instruction, the controller logic 205 can check the system state database against the expected state of the global rules to verify whether the expected operation is completed correctly.

[0143] Controller logic 205 may use versions according to version rules.System state 220 may have a database of which rule versions have been used in different deployments.

[0144] Controller logic 205 may include validation logic for rule optimization and efficient sequencing. Controller logic 205 may be configured to optimize resources. Information in the system state, rules, and templates related to applications that are running or expected to run may be used by the controller logic to achieve efficiency or priority with respect to resources. Controller logic 205 may use information in the "used resources" in system state 220 to determine efficiency, or determine the need to switch resources for upgrades, reuse, or other reasons.

[0145] The controller may check application operation according to the system state 220 and compare the application operation with the expected application operation of the global rules. If the application is not running, the controller may start the application. If the application should not run, the controller may stop the application and reallocate resources when appropriate. The controller logic 205 may include a database of resource (computing, storage, networking resources) specifications. The controller logic may include logic for identifying resource types that can be used for the system. This can be performed using the out-of-band management network 260. The controller logic 205 may be configured to use out-of-band management 260 to identify new hardware. The controller logic 205 may also obtain information about change history, rules used, and versions from the system state 220 for auditing, building reports, and change management purposes.

[0146] Figure 2F An example process flow of controller logic 205 is shown for processing template 230 and obtaining an image to start, drive and / or enable a resource, which for the purposes of this example may be referred to as a host. This process may also include configuring storage resources and coupling storage and computing hosts and / or resources. Controller logic 205 is aware of the hardware resources available in system 100, and system rules 210 may indicate which hardware resources can be utilized. Controller logic 205 parses template 230 at step 205.1, which may include an instruction file that may be executed to cause controller logic to collect information in a file generated by Figure 2E The instruction file may be in json format. At step 205.2, the controller logic collects the required file bucket list. In addition, at step 205.3, the controller logic 205 collects the required hardware-specific files into buckets, which are referenced by the hardware and optionally the hypermanagement system (or container host system, multi-tenant type system). If the hardware is to be run on a virtual machine, it may need to be referenced by the hypermanagement system (or container host system or multi-tenant type system).

[0147] If hardware specific files exist, the controller logic will collect the hardware specific files at step 205.4. In some cases, the file system image may contain a kernel and initramfs as well as a directory containing kernel modules (or the kernel modules are ultimately placed into a directory). The controller logic 205 then selects a compatible, appropriate base image at step 205.5. The base image contains operating system files that may not be specific to the application or image derived from the template 230. Compatibility in this context means that the base image contains the files required to turn the template into a working application. The base image can be managed outside the template as a mechanism for saving space (and typically, the base image may be the same for several applications or services). In addition, at step 205.6, the controller logic 205 selects one or more buckets with executable files, source code, and hardware specific configuration files. The template 230 may reference other files, including but not limited to: configuration files, configuration file templates (the configuration file template is a configuration file containing placeholders or variables, the configuration file is filled with variables in the system rules 210 that may become known in the template 230, so that the controller 200 can convert the configuration template into a configuration file and optionally change the configuration file through an API endpoint), binary and source code (the source code can be compiled when the image is started). At step 205.7, hardware specific instructions corresponding to the elements selected at steps 205.4., 205.5 and 205.6 can be loaded as part of the image being started. The controller logic 205 obtains the image from the selected components. For example, there may be different pre-installed scripts for a physical host compared to a virtual machine, or there may be differences for powerpc compared to x86.

[0148] At step 205.8, the controller logic 205 mounts overlayfs and repackages the subject files into a single file system blob. When multiple file system blobs are used, the image can be created through multiple blobs, decompressing the compressed package and / or obtaining git. If step 205.8 is not performed, the file system blobs can remain independent, and the image is created as a set of file system blobs and mounted using a file system (such as overlayfs) that can mount multiple smaller file systems together. The controller logic 205 can then locate a compatible kernel (or a kernel specified in the system rules 210) at step 205.9 and locate an applicable initrd at step 205.10. A compatible kernel can be a kernel that satisfies the dependencies of the template and the resources used to implement the template. A compatible initrd can be an initrd that loads the template onto the required computing resources. Typically, initird can be used for physical resources so that it can mount storage resources before fully booting (because the root file system may be remote). The kernel and initrd can be packaged into a filesystem blob for direct kernel booting, or for use on a physical host using kexec to change the kernel on the live system after booting the initial operating system.

[0149] The controller then configures one or more storage resources to allow one or more computing resources to drive one or more applications and / or one or more images using any of the techniques shown by 205.11, 205.12, and / or 205.13. Under 205.11, an overlayfs file may be provided as a storage resource. Under 205.12, a file system is presented. For example, a storage resource may present a combined file system, or a computing resource may simultaneously mount multiple file system blobs using a file system similar to overlayfs. Under 205.13, a blob is sent to the storage resource before presenting the file system.

[0150] Figure 2G and Figure 2H Shown for Figure 2F Further, the system may employ processes and rules for connecting computer resources to storage resources, which may be referred to as storage connection processes. Appendix A of the attached document provides a description of such storage connection processes, except for the steps 205.11 and 205.12 of FIG. Figure 2G and Figure 2H Instances other than those shown. Figure 2GAn example process flow for connecting storage resources is shown. Some storage resources may be read-only and other storage resources may be writable. The storage resource may manage its own write locks so that there are no simultaneous writes that would cause a race condition, or the system state 220 may track (see, e.g., step 205.20) which connections may write to the storage resource and / or prevent multiple read-write connections from being connected to the resource (step 205.21). The controller logic or the resource itself may query the controller's system state 220 for the location and transport type (e.g., iscsi, iser, nvmeof, fiber channel, fcoe, nfs, nfs over rdma, afs, cifs, window sharing) of the storage resource (step 205.22). If the computing resource is virtualized, the hypervisor (e.g., implemented via a hypervisor daemon) may handle the connection to the storage resource (step 205.23). This may have desirable security advantages because the virtual machine may not be aware of the SAN 280.

[0151] Referring to step 205.24, the process for connecting computing resources and storage resources may be indicated in the system rules 210. The controller logic then queries the system state 220 to ensure that the resources are available and writable (if necessary) (step 205.22). The system state 220 may be queried via any of a variety of techniques such as SQL queries (or other types of database queries), JSON parsing, etc. The query will return the information required for the computing resource to connect to the storage resource. The controller 200, the system state 220, or the system rules 210 may provide authentication credentials for the computing resource to connect to the system state (step 205.25). The computing resource will then update the system state 220 directly or via the controller (step 205.26).

[0152] Figure 2H An example boot process is shown in which a physical, virtual or other type of computing resource, application, service or host drives a storage resource and connects to the storage resource. The storage resource may optionally utilize a fused file system and / or expandable volumes. In the event that a controller or other system enables a physical host, the physical host may be preloaded with an operating system for configuring the system. Therefore, at step 205.31, the controller may preload a boot disk with initramfs. In addition, the controller 200 may use the out-of-band management connection 260 to network boot a preliminary operating system (step 205.30), and then optionally preload the preliminary operating system to the host (step 205.31). Thereafter, initramfs is loaded at step 205.32, and at step 205.33, the boot disk may be preloaded with the initramfs. Figure 2GThe storage resource may be connected using the method shown in FIG. 205. Then, if there is an expandable volume, the coupled subvolumes or devices are optionally assembled into a volume group at step 205.34 if logical volume management (LVM) is in use. Alternatively, other methods of combining disks may be used at step 205.34 to couple the volumes.

[0153] If the fusion file system is in use, the files may be combined at step 205.36 and the boot process may continue (step 205.46). If overlayfs is used in linux to work around some known issues, the following subprocess may be run. A / data directory may be formed in each mounted file system blob that may be volatile (step 205.37). The new_root directory may then be created at step 205.38 and overlayfs may be mounted into the directory at step 205.39. initramfs then runs exec_root on / new_root (step 205.40).

[0154] If the host is a virtual machine, additional tools such as direct kernel boot may be available. In this case, the hypervisor may connect to the storage resources before starting the VM (step 205.41), or the hypervisor may do so at boot time. The VM may then be directly kernel booted along with loading initramfs (step 205.42). The initramfs is then loaded at step 205.43, and the hypervisor may connect to the storage resources, which may be remote, at this time (step 205.44). To accomplish this, the hypervisor host may need to pass in interfaces (e.g., if Infiniband needs to connect to an iSER target, it may use pci-passhtru to pass in SR-IOV-based virtual functions, or in some cases paravirtualized network interfaces may be used). These connections are available for use by initramfs. If the virtual machine is not yet ready, the virtual machine may then connect to the storage resources at step 205.45. The virtual machine may also receive its storage resources through the hypervisor (optionally through paravirtualized storage). The process may be similar for virtual machines that optionally mount fusion file systems and LVM type disks.

[0155] Fig.2OAn example process flow for configuring storage resources from file system blobs or other file groups as at 205.13 is shown. The blobs are collected at step 205.75; and the blobs may be copied directly to the storage resource host at 205.73 (if the storage resource host is different from the device that holds the file system blob 232). Once the storage resources are in place, the system state is updated at 205.74 using the location of the storage resources and the available transport (e.g., iSER, nvmeof, iSCSI, FcoE, Fibre Channel, nfs, nfs over rdma). Some of these blobs may be read-only, in which case the system state remains unchanged and new computing resources or hosts may be connected to the read-only storage resources (e.g., when connected to a base image). In some cases, it may be desirable to place files into a single file system image as shown by 205.70 to avoid any fusion file system overhead. This can be done by mounting the blobs as a fused file system (step 205.71), then copying the blobs into a new file system or repacking them into a single file system (step 205.72), and then optionally copying the new file system image to the appropriate location where the new file system image will appear as a storage resource. Some fused file systems may allow the merge to be done without first mounting the fused file systems at step 205.71 and merging them in a single step.

[0156] Fig.2I It shows that Figure 2E Another example template 230 is shown. In this example, the controller can be configured to use Fig.2ITemplate 230 with a mediator configuration tool is shown. According to an example embodiment, the mediator configuration tool may include a common API for coupling a new application or service with a dependent application or service. Therefore, the template 230 may additionally include a list of dependencies 244 that may be required for setting up the service of the template. The template 230 may also include connection rules 245, which may include calls to the common API of the dependencies. The template 230 may also include one or more common APIs 243 and a common API and version list 242. The common API 243 may have methods, functions, scripts, or instructions that may (or may not) be called from an application or controller, which allow the controller to configure the dependent application or service so that the dependent application or service can then be coupled to the new application built by the template 230. The controller may communicate with the common API 243 and / or make API calls to configure the coupling of the new service or application with the dependent service or application. Optionally, the instructions may allow the application or service to communicate directly with the common API 243 on the dependent application or service and / or send calls to the common API. Template 230 connects rules 245, which is a set of rules and / or instructions that may include API calls for connecting a new service or application with dependent services or applications.

[0157] The system state 220 may also include a list of running services 246. The list of running services 246 may be queried by the controller logic 205 to try to satisfy the dependencies 244 from the templates 230. The controller may also include a list 247 of different common APIs that may be used for a particular service / application or a type of service / application, and may also include templates containing common APIs. The list may reside in the controller logic 205, the system rules 210, the system state 220, or in a template store accessible to the controller. The controller also maintains a common API index 248 compiled from all existing or loaded templates.

[0158] Figure 2J The controller logic 205 is shown as follows: Figure 2F The processing template 230 is shown but there is an example process flow for step 255 of managing service dependencies by the controller. Figure 2K Shown for Figure 2J255. At step 255.1, the controller collects the dependency list 244 from the template. The controller also collects the common API list 243 from the template. (A). At step 255.2, the controller narrows down the list of possible dependency applications or services by comparing the common API list 243 from the template to the common API index 248 and based on the type of application or service trying to satisfy the dependency. At step 255.3, the controller determines whether the system rule 210 specifies a way to satisfy the dependency.

[0159] If the determination is yes at step 255.3, the controller determines whether the dependency service or application is running by querying the running template list (step 255.4). If the determination is no at step 255.4, the service application is run (and / or configured first, then run), which may include the template of the dependency service / application being processed by the controller logic (step 255.5). If the dependency service or application is found to be running at step 255.4, the process flow proceeds to step 255.6. At step 255.6, the controller uses the template to couple the new service or application being built to the dependency service or application. In the process of coupling the new service or application with the dependency application / service, the controller will complete the template it is processing and will run the connection rules 245. The controller sends commands to the common API 243 based on the connection rules 245 on how to satisfy the dependencies 244 and / or couple the applications / services. The common API 243 translates the instructions from the controller to connect the new service or application with the dependent application or service, which may include but is not limited to: calling the service's API function, changing the configuration, running a script, calling other programs. After step 255.6, the process flow proceeds to Figure 2J Step 205.2.

[0160] If step 255.3 determines that the system rules 210 do not specify a way to satisfy the dependencies, the controller will query the system state 220 at step 255.7 to see if there are appropriate dependency applications or services running. At step 255.8, the controller makes its determination as to whether there are appropriate dependency applications or services running based on the query. If the determination at step 255.8 is no, the controller may notify the administrator or user to take action (step 255.9). If the determination at step 255.8 is yes, the process flow proceeds to step 255.6, which may be operated as described above. The user may optionally be asked whether the new application should be connected to the running dependency application, in which case the controller may couple the new application or service to the dependency application or service at step 255.6 as follows: The controller will complete the template 230 it is processing and will run the connection rules 245. The controller then sends a command to the common API 243 based on the connection rules 245 as to how to satisfy the dependencies 244. The common API 243 translates instructions from the controller to connect the new service or application with the dependent applications or services.

[0161] A user communicates with the controller 200 via an external user interface or Web UI, or application through the API application 120 , which may also be incorporated into the controller application or logic 205 .

[0162] Controller 200 communicates with the stack or resources through one or more of a plurality of networks, interconnects, or other connections that the controller can utilize to direct the operation of computing resources, storage resources, and networked resources. Such connections may include: out-of-band management connection 260; in-band management connection 270; SAN connection 280; and optional networking in-band management connection 290.

[0163] Out-of-band management can be used by the controller 200 to detect, configure, and manage components of the system 100 through the controller 200. The out-of-band management connection 260 enables the controller 200 to detect resources that are plugged in and available, but not turned on. The resources can be added to the IT system status 220 when they are plugged in. The out-of-band management can be configured to load a boot image, configure and monitor resources that are subordinate to the system 100. Out-of-band management can also start a temporary image for diagnosing the operating system. Out-of-band management can be used to change BIOS settings, and console tools can also be used to run commands on the running operating system. The settings can also be changed by the controller using the console, keyboard, and image recognition of video signals from physical or virtual monitoring ports on the hardware resources, such as VGA, DVI or HDMI ports, and / or using APIs provided by out-of-band management, such as Redfish.

[0164] Out-of-band management as used herein may include, but is not limited to, a management system capable of connecting to resources or nodes independent of the operating system and the main motherboard. Out-of-band management connection 260 may include a network, or multiple types of direct or indirect connectors or interconnects. Examples of out-of-band management connection types include, but are not limited to, IPMI, Redfish, SSH, telnet, other management tools, keyboard, display and mouse (KVM) or KVM over IP, serial console, or USB. Out-of-band management is a tool that can be used over a network, can drive and cut off nodes or resources, monitor temperature and other system data; make BIOS changes and other low-level changes that may be outside the control of the operating system; connect to a console and send commands; control inputs including but not limited to a keyboard, mouse, and monitor. Out-of-band management can be coupled to out-of-band management circuits in physical resources. Out-of-band management can connect to a disk image as a disk that can be used to boot the installation media.

[0165] A management network or in-band management connection 270 may allow the controller to collect information about computing resources, storage resources, networking resources, or other resources to be communicated directly to the operating system on which the resource is running. The storage resources, computing resources, or networking resources may include a management interface that interfaces with connection 260 and / or 270, whereby the resource may communicate with the controller 200 and inform the controller what is running and what is available for the resource, and receive commands from the controller. As used herein, an in-band management network includes a management network that is capable of communicating with a resource, directly leading to the operating system of the resource. Examples of in-band management connections may include, but are not limited to: SSH, telnet, other management tools, serial console, or USB.

[0166] Although out-of-band management is described herein as a network that is physically or virtually separate from the in-band management network, they may be combined or may work in conjunction with each other for efficiency purposes as described in more detail herein. Additionally and accordingly, out-of-band management and in-band management or aspects thereof may communicate through the same port of the controller or be coupled using a combined interconnect. Optionally, one or more of the connections 260, 270, 280, 290 may be separate or combined with other networks in such networks and may or may not include the same architecture.

[0167] In addition, computing resources, storage resources, and controllers may or may not be coupled to a storage network (SAN) 280 in a manner that enables the controller 200 to use the storage network to boot each resource. The controller 200 may send a boot image or other template to a separate storage resource or other resource or other resource so that the other resource can be booted from the storage resource or other resource. The controller may indicate where to boot from in this case. The controller may drive the resource, indicating where the resource is booted from and how to configure itself. The controller 200 indicates how the resource is booted, what image to use, and where the image is located if the image is on another resource. Resources for the BIOS may be preconfigured. The controller may also or optionally configure the BIOS through out-of-band management so that they will boot from a storage area network. The controller 200 may also be configured to boot an operating system from an ISO and enable resources to copy data to a local disk. The local disk may then be used for booting. The controller may configure other resources including other controllers in a manner so that the resources can boot. Some resources may include applications that provide computing, storage, or networking functions. In addition, it is possible for the controller to boot a storage resource and then have the storage resource be responsible for supplying a boot image for subsequent resources or services. Storage can also be managed over a different network that is used for another purpose.

[0168] Optionally, one or more of the resources may be coupled to a networking in-band management connection 290. Connection 290 may include one or more types of in-band management as described with respect to in-band management connection 270. Connection 290 may connect the controller to an application network to utilize the network, or to manage the network through the in-band management network.

[0169] Figure 2LAn image 250 is shown, which can be loaded directly or indirectly (through another resource or database) from a template 230 to a resource to start the resource, or an application or service loaded on the resource. The image 250 may include a boot file 240 for the resource type and hardware. The boot file 240 may include a kernel 241 corresponding to the resource, application or service to be deployed. The boot file 240 may also include an initrd or a similar file system for assisting the boot process. The boot system 240 may include multiple kernels or initrds configured for different hardware types and resource types. In addition, the image 250 may include a file system 251. The file system 251 may include a base image 252 and a corresponding file system, as well as a service image 253 and a corresponding file system, and a volatile image 254 and a corresponding file system. The loaded file system and data may vary according to the resource type and the application or service to be run. The base image 252 may include a base operating system file system. The base operating system may be read-only. The base image 252 may also include basic tools of the operating system that are independent of what is running. The base image 252 may include a base directory and operating system tools. The service file system 253 may include configuration files and specifications for resources, applications, or services. The volatile file system 254 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables, including but not limited to: passwords, session keys, and private keys. The file system can be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.

[0170] As described above, the controller 200 may be used to add resources such as computing resources, storage resources, and / or networking resources to the system. Fig.11AAn example method for adding a physical resource, such as a bare metal node, to a system 100 is shown. A resource, i.e., a computing resource, a storage resource, or a networking resource, is inserted into a controller via a network connection 1110. The network connection may include an out-of-band management connection. The controller recognizes that the resource is inserted via the out-of-band management connection 1111. The controller identifies information related to the resource, which may include, but is not limited to, the type, capabilities, and / or attributes of the resource 1112. The controller adds the resource and / or information related to the resource to its system state 1113. An image derived from a template is loaded onto a physical component of the system, which may include, but is not limited to, another resource such as a storage resource, or a resource on a controller 1114. The image includes one or more file systems that may include configuration files. Such configurations may include BIOS and boot parameters. The controller instructs the physical resource to boot using the image's file system 1115. Additional resources or multiple different types of bare metal or physical resources may be added in this manner using an image of a template, or at least a portion thereof.

[0171] Fig. 11B An example method of automatically allocating resources using global system rules and templates for an example implementation is shown. A request is made to a system that needs resource allocation to satisfy the request 1120. The controller is aware of its resource pool based on its system state database 1121. The controller uses templates to determine the required resources 1122. The controller dispatches the resources and stores the information in the system state 1123. The controller deploys the resources using the template 1124.

[0172] refer to Fig.12 , an example method for automatically deploying an application or service is illustrated using the system 100 described herein. A user or application issues a request for a service 1210. The request is translated to an API application 1220. The API application delivers the request to a controller 1230. The controller interprets the request 1240. The controller takes into account the state of the system and its resources 1250. The controller uses its rules and templates to deploy the service 1260. The controller 1270 sends the request to the resource 1270 and deploys the image 1280 derived from the template and updates the IT system state.

[0173] Other more detailed examples of operations such as adding resources, allocating resources, and deploying applications or services are discussed in more detail below.

[0174] Adding compute resources to the system

[0175] refer to Figure 3A, showing the addition of a computing resource 310 to the system 100. When adding a computing resource 310, the computing resource is coupled to the controller 200 and the computing resource can be shut down. It should be noted that if the computing resource 310 is pre-loaded with an image, alternative steps can be followed, where any network connection can be used to communicate with the resource, start the resource and add information to the system state. If the computing resource and the controller are on the same node, the service running the computing resource is shut down.

[0176] like Figure 3A As shown, the computing resource 310 is coupled to the controller via the following networks: out-of-band management connection 260, in-band management connection 270, and optionally SAN 280. The computing resource 310 is also coupled to one or more application networks 390, where services, application users and / or clients can communicate with each other. The out-of-band management connection 260 can be coupled to a separate out-of-band management device 315 or circuit of the computing resource 310 that is turned on when the computing resource 310 is inserted. The device 315 can allow features including, but not limited to: driving / disconnecting devices, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings and other features outside the scope of the operating system. The controller 200 can view the computing resource 310 through the out-of-band management network 260. The controller can also distinguish the type of computing resource and use in-band management or out-of-band management to distinguish the configuration of the computing resource. The controller logic 205 is configured to carefully check the added hardware in the out-of-band management 260 or in-band management 270. If a computing resource 310 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource will be configured automatically or through interaction with the user. If the resource is added automatically, the settings will follow the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the computing resource. The controller 200 can query the API application or otherwise request the user or any program in the control stack to confirm that the new resource has been authorized. The authorization process can also be automatically and securely completed using cryptography to confirm the legitimacy of the new resource. The controller logic 205 adds the computing resource 310 to the IT system state 220, which includes the switch or network into which the computing resource 310 is inserted.

[0177] If the computing resource is physical, the controller 200 can drive the computing resource through the out-of-band management network 260, and the computing resource 310 can be started from the image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, such as through the SAN 280. The image can be loaded through other network connections or indirectly through another resource. Once started, information related to the computing resource 310 received through the in-band management connection 270 can also be collected and added to the IT system state 220. The computing resource 310 can then be added to the storage resource pool, and the computing resource will become a resource managed by the controller 200 and tracked in the IT system state 220.

[0178] If the computing resource is virtual, the controller 200 can drive the computing resource through the in-band management network 270 or through the out-of-band management 260. The computing resource 310 can be started from the image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, such as through the SAN 280. The image can be loaded through other network connections or indirectly through another resource. Once started, information related to the computing resource 310 received through the in-band management connection 270 can also be collected and added to the IT system state 220. The computing resource 310 can then be added to the storage resource pool, and the computing resource will become a resource managed by the controller 200 and tracked in the IT system state 220.

[0179] The controller 200 may be able to automatically turn resources on and off based on global system rules and update the IT system status for reasons determined by the IT system user, such as turning resources off to save power, or turning resources on to improve application performance, or any other reason the IT system user may have.

[0180] Figure 3BIt is an image 350 that is loaded directly or indirectly (through another resource or database) from the template 230 to the computing resource 310 to start the computing resource and / or load the application. The image 350 may include a startup file 340 for the resource type and hardware. The startup file 340 may include a kernel 341 corresponding to the resource, application or service to be deployed. The startup file 340 may also include initrd or a similar file system for assisting the startup process. The startup system 340 may include multiple kernels or initrd configured for different hardware types and resource types. In addition, the image 350 may include a file system 351. The file system 351 may include a base image 352 and a corresponding file system, a service image 353 and a corresponding file system, and a volatile image 354 and a corresponding file system. The loaded file system and data may vary according to the resource type and the application or service to be run. The base image 352 may include a base operating system file system. The base operating system may be read-only. The base image 352 may also include basic tools of the operating system that are independent of what is running. The base image 352 may include a base directory and an operating system tool. The service file system 353 may include configuration files and specifications for resources, applications, or services. The volatile file system 354 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to: passwords, session keys, and private keys. The file system may be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.

[0181] Figure 3C 1 shows an example process flow for adding a resource, such as computing resource 310, to system 100. Although in this example, the subject resource will be described as computing resource 310, it should be understood that Figure 3C The subject resource of the process flow may also be a storage resource 410 and / or a networked resource 510. Figure 3C In the example of , the added resource 310 is not on the same node as the controller 200. At step 300.1, the resource 310 is coupled to the controller 200 in a disconnected state. Figure 3CIn the example of FIG. 300 , the out-of-band management connection 260 is used to connect to the resource 310. However, it should be understood that other network connections may be used if desired by the practitioner. At steps 300.2 and 300.3, the controller logic 205 goes through the system's out-of-band management connection and uses the out-of-band management connection 260 to identify and discern the type of resource 310 being added and its configuration. For example, the controller logic may look at the BIOS or other information (such as serial number information) of the resource as a reference to obtain type and configuration information.

[0182] At step 300.4, the controller uses the global system rules to determine whether a particular resource 310 should be automatically added. If not, the controller will wait until its use is authorized (step 300.5). For example, at step 300.4, a user may respond to the query that the user does not want to use a particular resource 310, or that the particular resource may be automatically put on hold until it is to be used. If step 300.4 determines that a resource 310 should be automatically added, the controller will use its rules to perform automatic setup (step 300.6) and proceed to step 300.7.

[0183] At step 300.7, the controller selects and uses a template 230 associated with the resource to add the resource to the system state 220. In some cases, a template 230 may be specific to a particular resource. However, some templates 230 may cover multiple resource types. For example, some templates 230 may be cross-hardware. At step 300.8, the controller drives the resource 310 through its out-of-band management connection 260 according to the global system rules 210. At step 300.9, using the global system rules 210, the controller finds and loads the boot image of the resource from one or more selected templates. The resource 310 is then started from the image derived from the subject template 230 (step 300.10). Then, after starting the resource 310, additional information about the resource 310 may be received from the resource 310 through the in-band management connection 270 (step 300.11). Such information may include, for example, firmware version, network card, any other device to which the resource may be connected. The new information may be added to the system state 220 at step 300.12. The resource 310 may then be considered to have been added to the resource pool and is ready for allocation (step 300.13).

[0184] about Figure 3C, if the resource and the controller are on the same node, it should be understood that the service running the resource may be remote from the node. In this case, the controller can use an inter-process communication technology for the resource, such as unix sockets, loopback adapters or other inter-process communication technologies to communicate with the resource. According to system rules, the controller can install a virtual host, or a hypervisor or container host to run the application using a template known from the controller. The resource application information can then be added to the system state 220, and the resource will be ready for allocation.

[0185] Add storage resources to the system:

[0186] Figure 4A The addition of storage resource 410 to system 100 is shown. In one example embodiment, the Figure 3C 200 to add a storage resource 410 to the system 100, where the added storage resource 410 is not on the same node as the controller 200. Additionally, it should be noted that if the storage resource 410 is preloaded with an image, alternative steps may be followed, where any network connection may be used to communicate with the storage resource 410, start the storage resource 410, and add information to the system state 220.

[0187] When a storage resource 410 is added, the storage resource is coupled to the controller 200 and the storage resource can be disconnected. The storage resource 410 is coupled to the controller through the following networks: an out-of-band management network 260, an in-band management connection 270, a SAN 280, and optionally a connection 290. The storage resource 410 may also be coupled or uncoupled to one or more application networks 390, where services, application users, and / or clients can communicate with each other. An application or client may achieve direct or indirect access to the storage of the resource via the application, whereby the storage of the resource is not accessed through the SAN. The application network may have built-in storage, or may be accessed and identified as a storage resource in the IT system state. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 415 or circuitry of the storage resource 410 that is turned on when the storage resource 410 is inserted. The device 415 may allow features including, but not limited to: driving / disconnecting the device, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings and other features outside the scope of the operating system. The controller 200 can view the storage resource 410 through the out-of-band management network 260. The controller can also identify the type of storage resource and use in-band management or out-of-band management to identify the configuration of the storage resource. The controller logic 205 is configured to carefully check the added hardware in the out-of-band management 260 or the in-band management 270. If the storage resource 410 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource 410 will be configured automatically or configured by interacting with the user. If the resource is added automatically, the setting will follow the global system rules 210 in the controller 200. If the resource is added by the user, the global system rules 210 in the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the storage resource. The controller 200 can query one or more API applications or otherwise request the user or any program of the control stack to confirm that the new resource has been authorized. The authorization process can also be automatically and securely completed using cryptography to confirm the legitimacy of the new resource. The controller logic 205 adds the storage resource 410 to the IT system state 220, which includes the switch or network into which the storage resource 410 is plugged.

[0188] The controller 200 may drive the storage resource 410 through the out-of-band management network 260, and the storage resource 410 may be booted from an image 450 loaded from the template 230, for example, through the SAN 280, using the global system rules 210 and the controller logic 205. The image may also be loaded through other network connections or indirectly through another resource. Once booted, information related to the storage resource 410 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 410 is now added to the storage resource pool, and the storage resource will become a resource managed by the controller 200 and tracked in the IT system state 220.

[0189] The storage resources may include a pool of storage resources, or multiple pools of storage resources that the IT system can use or access independently or simultaneously. When storage resources are added, the storage resources may provide one storage pool, multiple storage pools, a portion of a storage pool, and / or multiple portions of multiple storage pools to the IT system state. The controller and / or storage resources may manage the various storage resources of the pool, or the grouping of such resources within the pool. The storage pool may include multiple storage pools running on multiple storage resources. For example, a flash disk or array, a cache disk or array, or a storage pool on a dedicated computing node coupled with a pool on a dedicated storage node is used to optimize bandwidth and latency simultaneously.

[0190] Figure 4BAn image 450 is shown loaded directly or indirectly (via another resource or database) from template 230 to storage resource 410 to start the storage resource and / or load the application. Image 450 may include a boot file 440 for resource type and hardware. Boot file 440 may include a kernel 441 corresponding to the resource, application or service to be deployed. Boot file 440 may also include initrd or a similar file system for assisting the boot process. Boot system 440 may include multiple kernels or initrds configured for different hardware types and resource types. In addition, image 450 may include file system 451. File system 451 may include base image 452 and corresponding file system, as well as service image 453 and corresponding file system, and volatile image 454 and corresponding file system. The loaded file system and data may vary according to the resource type and the application or service to be run. Base image 452 may include a base operating system file system. The base operating system may be read-only. Base image 452 may also include basic tools of the operating system that are independent of what is running. Base image 452 may include a base directory and operating system tools. The service file system 453 may include configuration files and specifications for resources, applications, or services. The volatile file system 454 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to: passwords, session keys, and private keys. The file system may be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.

[0191] Figure 5A An example is shown in which another storage resource, namely direct attached storage 510 (which may take the form of a node with JBOD or other type of direct attached storage) is coupled to storage resource 410 as an additional storage resource for the system. JBOD is an external disk array that is typically connected to the node providing the storage resource, and the JBOD will be used as Figure 5A , but it should be understood that other types of direct-attached storage may also be adopted as 510.

[0192] For example, regarding Figure 5AAs described, the controller 200 can add storage resources 410 and JBODs 510 to its system. The JBODs 510 are coupled to the controller 200 via an out-of-band management connection 260. The storage resources 410 are coupled to the following networks: the out-of-band management connection 260, the in-band management connection 270, the SAN 280, and optionally the connection 290. The storage nodes 410 communicate with the storage of the JBODs 510 via a SAS or other disk drive architecture 520. The JBODs 510 may also include an out-of-band management device 515 that communicates with the controller via the out-of-band management connection 260. Through the out-of-band management 260, the controller 200 can detect the JBODs 510 and the storage resources 410. The controller 200 may also detect other parameters that are not controlled by the operating system, for example, as described herein with respect to various out-of-band management circuits. The global system rules 210 of the controller 200 provide configuration startup rules for starting or starting JBODs and storage nodes that have not yet been added. The order in which storage resources are enabled may be controlled by the controller logic 205 using global rules 220. According to one set of global system rules 220, the controller may first drive the JBOD 510, and the controller 200 may then use the loaded image 450 to drive the storage resource 410 in a manner similar to that described with respect to FIG. 4. In another set of global system rules, the controller 200 may first enable the storage resource 410 and then enable the JBOD 510. In other global system rules, timing or delays between driving various devices may be specified. Through the controller logic 205, the global system rules 210 and / or the templates 230, detection of the readiness or operational status of various resources may be determined and / or used by the controller 200 for device allocation management. The IT system state 220 may be updated by communicating with the storage resource 410. The storage node 410 learns the storage parameters and configuration of the JBOD 510 by accessing the JBOD via the disk architecture 520. The storage resource 410 provides information to the controller 200, which then updates the IT system state 220 with information about the amount of storage available and other attributes. The controller updates the IT system state 220 when the storage resource 410 is started and the storage resource 410 is identified as part of the pool of storage resources 400 of the system 100. The storage node processes the logic for controlling the JBOD storage resource using the configuration set by the controller 200. For example, the controller may instruct the storage node to configure the JBOD to create a pool from a RAID 10 or other configuration.

[0193] Figure 5BAn example process flow is shown for adding a storage resource 410 and a direct attached storage 510 for the storage resource 410 to the system 100. At step 500.1, the direct attached storage 510 is coupled to the controller 200 in a disconnected state via the out-of-band management connection 260. At step 500.2, the storage resource 410 is coupled to the controller 200 in a disconnected state via the out-of-band management connection 260 and the in-band management connection 270, while the storage resource 410 is coupled to the direct attached storage 510, for example, via a SAS 520, such as a disk drive fabric.

[0194] The controller logic 205 may then traverse the out-of-band management connection 260 to detect the storage resource 410 and the directly attached storage 510 (step 500.3). While any network connection may be used, in this example, out-of-band management is available for use by the controller logic to identify and discern the type of resource being added (in this case, the storage resource 410 and the directly attached storage 510) and its configuration (step 500.4).

[0195] At step 500.5, the controller 200 selects and uses a template 230 for a specific storage type for each type of storage device to add the resources 410 and 510 to the system state 220. At step 500.6, the controller starts the direct storage and storage nodes in this order through the out-of-band management connection 260 according to the global system rules 210 (which may specify the boot order, drive order, and start the direct storage and storage nodes in this order (500.6). Using the global system rules 210, the controller finds and loads the boot image of the storage resource 410 from the template 230 selected for the storage resource 410, and then starts the storage resource from the image (step 500.7). The storage resource 410 learns the storage parameters and configuration of the direct attached storage 510 by accessing the direct attached storage 510 via the disk architecture 520. The storage resource 410 can then be connected to the storage resource 520. The in-band management connection 270 provides additional information about the storage resource 410 and / or the direct attached storage 510 to the controller (step 500.8). At step 500.9, the controller updates the system state 220 with the information obtained at step 500.8. At step 500.10, the controller processes the direct attached storage 510 and sets a configuration for the storage resource 410 and how the direct attached storage is configured. At step 500.11, a new resource comprising a combination of the storage resource 410 and the direct attached storage 510 may then be added to the resource pool and made ready for allocation within the system.

[0196] According to another aspect of an example embodiment, the controller may use out-of-band management to identify other devices in the stack that may not be participating in the computation or service. For example, such devices may include, but are not limited to: cooling towers / air conditioners, lights, temperature devices, sound devices, alarm devices, power systems, or any other devices associated with the system.

[0197] Add networking resources to the system:

[0198] Fig. 6A The addition of a networked resource 610 to the system 100 is shown. In one example embodiment, the Figure 3C 200 to add a networked resource 610 to the system 100, where the added networked resource 610 is not on the same node as the controller 200. Additionally, it should be noted that if the networked resource 610 is preloaded with an image, alternative steps may be followed where any network connection may be used to communicate with the network resource 610, start the network resource 610, and add information to the system state 220.

[0199] When a networked resource 610 is added, the networked resource is coupled to the controller 200 and the networked resource can be disconnected. The networked resource 610 can be coupled to the controller 200 through the following connections: an out-of-band management connection 260 and / or an in-band management connection 270. The networked resource is optionally inserted into the SAN 280 and / or the connection 290. The networked resource 610 may also be coupled or uncoupled to one or more application networks 390, where services, application users and / or clients can communicate with each other. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 615 or circuitry of the networked resource 610 that is turned on when the networked resource 610 is inserted. The device 615 may allow features including, but not limited to: driving / disconnecting the device, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings and other features outside the scope of the operating system. The controller 200 can view the networked resource 610 through the out-of-band management connection 260. The controller can also identify the type of networked resource and / or network architecture and use in-band management or out-of-band management to identify the configuration. The controller logic 205 is configured to carefully check the added hardware in the out-of-band management 260 or the in-band management 270. If a networked resource 610 is detected, the controller logic 205 can use the global system rules 220 to determine whether the networked resource 610 will be configured automatically or by interacting with the user. If the resource is added automatically, the settings will follow the global system rules 210 within the controller 200. If added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the resource. The controller 200 can query one or more API applications or otherwise request the user or any program of the control stack to confirm that the new resource has been authorized. The authorization process can also be automatically and securely completed using cryptography to confirm the legitimacy of the new resource. The controller logic 205 can then add the networked resource 610 to the IT system state 220. For switches that cannot identify themselves to the controller, the user can manually add the switch to the system state.

[0200] If the networking resource is physical, the controller 200 can drive the networking resource 610 through the out-of-band management connection 260, and the networking resource 610 can be started from the image 605 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, through the SAN 280. The image can also be loaded through other network connections or indirectly through other resources. Once started, information related to the networking resource 610 received through the in-band management connection 270 can also be collected and added to the IT system state 220. The networking resource 610 can then be added to the storage resource pool, and the networking resource will become a resource managed by the controller 200 and tracked in the IT system state 220. Optionally, some networking resource switches can be controlled through a console port connected to the out-of-band management 260 and can be configured when driven, or can have a switch operating system installed through a boot loader, such as through ONIE.

[0201] If the networking resource is virtual, the controller 200 may drive the networking resource through the in-band management network 270 or through out-of-band management 260. The networking resource 610 may be started from the image 650 loaded from the template 230 through the SAN 280 using the global system rules 210 and the controller logic 205. Once started, information related to the networking resource 610 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The networking resource 610 may then be added to a storage resource pool and the networking resource will become a resource managed by the controller 200 and tracked in the IT system state 220.

[0202] The controller 200 may instruct a networked resource (whether physical or virtual) to assign, reassign, or move a port to connect to a different physical or virtual resource, i.e., a connection, storage, or computing resource as defined herein. This may be accomplished using technologies including, but not limited to, SDN, InfiniBand zoning, VLAN, vXLAN. The controller 200 may instruct a virtual switch to move or assign a virtual interface to a network or interconnector that communicates with a virtual switch or a resource hosting a virtual switch. Some physical or virtual switches may be controlled by an API coupled to the controller.

[0203] Controller 200 may also instruct computing resources, storage resources, or networking resources to change fabric types if such a change is possible. A port may be configured to switch to a different fabric, such as a hybrid Infiniband / Ethernet interface.

[0204] The controller 200 may give instructions to networking resources that may include switches or other networking resources that switch multiple application networks. The switches or network devices may include different architectures, or, for example, they may be plugged into InfiniBand switches, ROCE switches, and / or other switches that preferably have SDN capabilities and multiple architectures.

[0205] Figure 6B An image 650 is shown loaded directly or indirectly (e.g., through another resource or database) from template 230 to networked resources 610 to start networked resources and / or load applications. Image 650 may include boot files 640 for resource types and hardware. Boot files 640 may include a kernel 641 corresponding to the resource, application, or service to be deployed. Boot files 640 may also include initrd or a similar file system for assisting the boot process. Boot system 640 may include multiple kernels or initrds configured for different hardware types and resource types. In addition, image 650 may include file system 651. File system 651 may include base image 652 and corresponding file system, as well as service image 653 and corresponding file system, and volatile image 654 and corresponding file system. The loaded file system and data may vary according to the resource type and the application or service to be run. Base image 652 may include a base operating system file system. The base operating system may be read-only. Base image 652 may also include basic tools of the operating system that are independent of what is running. Base image 652 may include a base directory and operating system tools. The service file system 653 may include configuration files and specifications for resources, applications, or services. The volatile file system 654 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to: passwords, session keys, and private keys. The file system may be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.

[0206] Deploy an application or service on a resource:

[0207] Fig. 7A The system 100 is shown, and the system 100 includes: a controller 200; physical and virtual computing resources, the physical and virtual computing resources including a first computing node 311, a second computing node 312, and a third computing node 313; storage resources 410; and network resources 610. The resources are shown as described herein with respect to Figures 1 to 6B The described approach sets and adds to the IT system state 220 .

[0208] Although multiple computing nodes are shown in this figure, a single computing node may also be used according to an example embodiment. A computing node may host physical or virtual computing resources, and applications may be run on a physical or virtual computing node. Similarly, although a single network provider node and storage node are shown, it is contemplated that multiple resource nodes of these types may or may not be used in a system of an example embodiment.

[0209] A service or application may be deployed in any system according to an example embodiment. An instance of a service deployed on a computing node may be associated with Fig. 7A , but may similarly be used with different arrangements of system 100. For example, Fig. 7A The controller 200 in the example of FIG. 200 may automatically configure computing resources 310 in the form of computing nodes 311, 312, 313 according to the global system rules 210. The computing resources may then also be added to the IT system state 220. The controller 200 may thus identify the computing resources 311, 312, 313 (which may or may not be cut off) and any possible physical or virtual applications running on the computing resources or nodes. The controller 200 may also automatically configure one or more storage resources 410 and one or more networking resources 610 according to the global system rules 210 and the templates 230 and add them to the IT system state 220. The controller 200 may identify the storage resources 410 and the networking resources 610 that may or may not be in a cut off state to begin with.

[0210] Figure 7B An example process for adding a resource to an IT system 100 is shown. At step 700.1, a new physical resource is coupled to the system. At step 700.2, the controller becomes aware of the new resource. The resource may be connected to remote storage (step 700.4). At step 700.3, the controller configures a method for starting the new resource. All connections made to the resource may be logged to the system state 220 (step 700.5). Figure 3C Provides information such as Figure 7B More details of an example embodiment of the process flow shown.

[0211] Figure 7C and Fig.7DAn example process flow for deploying an application on multiple computing resources, multiple servers, multiple virtual machines, and / or in multiple sites is shown. The process of this example differs from a standard template deployment in the fact that the IT system 100 will require components to couple redundant and related applications and / or services. The controller logic may process a meta-template at step 700.11, where the meta-template may contain multiple templates 230, file system blobs 232, and other components required to configure a multi-hosted service (which may be in the form of other templates 230).

[0212] At step 700.12, the controller logic 205 checks the available resources in the system state 220; however, if there are not enough resources, the controller logic may cut back on the number of redundant services that may be deployed (see 700.16, where the number of redundant services is identified). At step 700.13, the controller logic 205 configures the networking resources and interconnects needed to connect the services together. If the service or application is deployed across multiple sites, the meta-template may include (or the controller logic 205 may configure) services that are optionally configured from a template that allows data synchronization and interoperability across sites (see 700.15).

[0213] At step 700.16, the controller logic 205 may determine from the system rules the meta template data, resource availability, and the number of redundant services (if there will be redundant services on multiple hosts). At 700.17, there is coupling to other redundant services and coupling to the motherboard. If there are multiple redundant hosts, the controller logic 205 or logic within the template (which may contain a binary 234 of a configuration file set in the boot operating system, a daemon 232, or a file system blob) may prevent network address and host name conflicts. Optionally, the controller logic will provide the network address (see 700.18) and register each redundant service in the DNS (700.19) and the system state 220 (700.18). The system state 220 will track redundant services, and if the controller logic 205 notices that a redundant service has conflicting parameters, such as host name, dns name, network address, already in the system state 220, the controller logic will not allow duplicate registrations.

[0214] Depend on Fig.7DThe configuration routine shown will process one or more templates in the meta template. The configuration routine will process all redundant services, deploy multi-host or cluster services to multiple hosts, and deploy services to couple the hosts. Any process that can deploy an IT system from a system rule can run a configuration routine. In the case of multi-host services, the example routine may process service templates as at 700.32, provision storage resources as at 700.33, drive hosts as at 700.35, couple hosts / computing resources with storage resources as at 700.36 (and register in the system state 220) (then repeat for the number of redundant services (700.38); each time register in the system state 220 (see 700.20) and use controller logic to record information to track individual services and prevent conflicts (see 700.31).

[0215] Some of the service templates may contain services and tools that can couple multi-host services. Some of these services may be considered dependencies (700.39), and then the coupling routine at 700.40 may be used to couple the services and register the coupling in the system state 220. In addition, one of the service templates may be a master template, and then the dependent service template at 700.39 will be a slave or secondary service; and the coupling routine at 700.40 will connect the services. The routines may be defined in a meta template; for example, for a redundant dns configuration, the coupling routine at 700.40 may include a connection from the dns to the primary dns and configuration for zone transfers along with dnssec. Some services may use physical storage (see 700.34) to improve performance, and the physical storage may be loaded with Figure 5B The preliminary OS disclosed in . Tools for coupling services can be included in the template itself, and configuration between services can be done using APIs accessible by the controller and / or other hosts in a multi-node application / service.

[0216] The controller 200 may allow a user or controller to determine an appropriate computing backend for an application. The controller 200 may allow a user or controller to optimally place an application on an appropriate physical or virtual computing resource by determining resource usage. When a super management system or other computing backend is deployed to a computing node, the super management system or other computing backend may report resource utilization statistics back to the controller via an in-band management connection 270. When the controller decides to create an application on a virtual computing resource based on its own logic and global system rules, or based on user input, the controller may automatically select the super management system on the optimal host and drive the virtual computing resources on the host.

[0217] For example, the controller 200 deploys an application or service to one or more computing resources using one or more templates 230. Such an application or service may be, for example, a virtual machine running the application or service. In one example, Fig. 7A The deployment of multiple virtual machines (VMs) on multiple computing nodes is shown, as shown, the controller 200 can identify that there are multiple computing resources 310 in the form of computing nodes 311, 312, 313 in its computing resource pool. The computing nodes can be deployed, for example, using a hypervisor system or optionally on bare metal, where the use of virtual machines may be undesirable due to speed reasons. In this example, the computing resource 310 is loaded with a hypervisor application and has VM (1) 321 and VM (2) 322 configured and deployed on the computing node 311. If, for example, the computing node 311 does not have resources for additional VMs, or if other resources are preferred for a particular service, the controller 200 can identify that there are no available resources on the computing node 311 based on the stack state 220, or it is preferred to set up a new VM in a different resource. It may also be recognized that the hypervisor system is loaded on, for example, computing resource 312, but not on resource 313, which may be a bare metal computing node used for other purposes. Therefore, based on the requirements of the service or application template being installed, and the status of the system state 220, the controller in this example may select a compute node 313 for deploying the next required resource VM(3) 323.

[0218] The computing resources of the system may be configured to share storage on the storage resources of the storage nodes.

[0219] The user can request to set up services for the system 100 through the user interface 110 or the application. The services may include but are not limited to: email services; web services; user management services; network providers; LDAP; Dev tools; VOIP; authentication tools; billing.

[0220] The API application 120 translates the user or application request and sends a message to the controller 200. The service template or image 230 of the controller 200 is used to identify which resources are required for the service. The resources to be used are then identified based on availability according to the IT system state 220. The controller 200 makes a request to one or more of the computing nodes 311, 312, or 313 for the required computing service, a request to the storage resource 410 for the required storage resource, and a request to the network resource 610 for the required networking resource. The IT system state 220 is then updated to identify the resources to be allocated. The service is then installed to the allocated resources using the global system rules 210 according to the template 230 of the service or application.

[0221] According to an example embodiment, multiple computing nodes may be used by the same service or different services, and, for example, a storage service and / or a pool of network providers may be shared among the computing nodes.

[0222] refer to Fig. 8A , shows a system 100 in which a controller 200, as well as computing resources 300, storage resources 400, and networking resources 600 are on the same or shared physical hardware, such as a single node. Figures 1 to 10 The various features described and shown in the figure can be incorporated into a single node. When the node is driven, the controller image is loaded on the node. The computing resources 300, storage resources 400, and networking resources 600 utilize the template 230 and use the global system rules 210 to configure. The controller 200 can be configured to load the computing backends 318, 319 as computing resources, and the computing backends may be added or not added to the node or one or more different nodes. Such backends 318, 319 may include, but are not limited to: virtualization technology, containers, and multi-tenant processes for creating virtual computing resources, networking resources, and storage resources.

[0223] Applications or services 725, such as web, email, core network services (DHCP, DNS, etc.), collaboration tools can be installed on virtual resources on nodes / devices that are shared with the controller 200. These applications or services can be moved to physical or virtual resources independently of the controller 200. Applications can run on virtual machines on a single node.

[0224] Figure 8B A method for expanding from a single-node system to a multi-node system (such as a Fig. 8A Node 318 and / or 319) of the example process flow shown. Therefore, reference Fig. 8A and Figure 8B , we consider an IT system with a controller 200 running on a single server; wherein it is desired to scale out the IT system to a multi-node IT system. Thus, before the expansion, the IT system is in a single-node state. Fig. 8A As shown, the controller 200 runs on a multi-tenant single-node system to drive various IT system management applications and / or resources, which may include but are not limited to: storage resources, computing resources, hypervisor systems and / or container hosts.

[0225] At step 800.2, the new physical resource is coupled to the single-node system by connecting the new physical resource via the out-of-band management connection 260, the in-band management connection 270, the SAN 280, and / or the network 290. For the purposes of this example, this new physical resource may also be referred to as hardware or a host. The controller 200 may detect the new resource on the management network and then query the device. Optionally, the new device may broadcast a message announcing itself to the controller 200. For example, the new device may be identified by MAC address, out-of-band management, and / or booting into a preliminary OS and using in-band management to identify the new device, thereby identifying the hardware type. In either event, at step 800.3, the new device provides information about its node type and its currently available hardware and software resources to the controller. The controller 200 then learns about the new device and its capabilities.

[0226] At step 800.4, tasks assigned to the system running the controller 200 may be assigned to the new host. For example, if the host is preloaded with an operating system (such as a storage host operating system or a hypervisor), the controller 200 assigns new hardware resources and / or capabilities. The controller may then provide the image and provision the new hardware, or the new hardware may request the image from the controller and configure itself using the methods disclosed above and below. If the new host is unable to host storage resources or virtual computing resources, the new resources may be made available to the controller 200. The controller 200 may then move and / or assign existing applications to the new resources, or use the new resources for newly created applications or applications created later.

[0227] At step 800.5, the IT system may keep its current applications running on the controller or migrate the current applications to the new hardware. If virtual computing resources are migrated, VM migration techniques (such as migration tools such as qemu+kvm) may be used and the system state and new system rules may be updated. The change management techniques discussed below may be used to make these changes reliably and securely. Since more applications may be added to the system, the controller may use any of a variety of techniques to determine how to allocate the system's resources, including but not limited to polling techniques, weighted polling techniques, least utilization techniques, weighted least utilization techniques, prediction techniques based on utilization and auxiliary training, scheduling techniques, expected capacity techniques, and capacity capping techniques.

[0228] Figure 8CAn example process flow for migrating a storage resource to a new physical storage resource is shown. The storage resource may then be mirrored, migrated, or a combination thereof (e.g., storage may be mirrored and then the original storage resource disconnected). At step 820, the storage resource is coupled to the system by having the new storage resource contact the controller or having the controller discover the new storage resource. This may be accomplished using an out-of-band management connection 260, an in-band management connection 270, a SAN network 280, or a flat network that may be being used by the application network, or a combination thereof. Under in-band management, the operating system may be pre-booted and the new resource may be connected to the controller.

[0229] At step 822, a new storage target is created on the new storage resource; and at step 824, this may be recorded in a database. In one instance, the storage target may be created by copying a file. In another instance, the storage target may be created by creating a block device and copying data (the data may be in the form of one or more file system blobs). In another instance, the storage target may be created by mirroring 2 or more storage resources between block devices (e.g., creating a raid) and optionally connecting via one or more remote storage transports, including but not limited to: iscsi, iser, nvmeof, nfs, nfs over rdma, fc, fcoe, srp, etc. The database input at step 824 may include information about the computing resources (or other types of resources and / or hosts) connected to the new storage resource remotely or locally (if the storage resource is on the same device as the other resources or hosts).

[0230] At step 826, the storage resources are synchronized. For example, the storage may be mirrored. As another example, the storage may be taken offline and synchronized. At step 826, techniques such as raid 1 (or other types of raid - but typically raid 1 or raid 0, and if desired raid 110 (mirrored raid 10) (mdadm, zfs, btrfs, hardware raid)) may be employed.

[0231] Then, after the database record at step 828, the data from the old storage resource is optionally connected (if the operation occurs later, the database may contain information about the status of copying the data, if such data must be recorded). If the storage target is being migrated away from the original host (e.g., as previously described according to Fig. 8A and Figure 8BIf the system is moving from a single-node system to a multi-node system and / or a distributed IT system as described above), the new storage resource may be designated as the primary storage resource by the controller, system state, computing resource, or a combination thereof at step 830. This may be done as a step to remove the old storage resource. In some cases, the physical or virtual hosts connected to the resource may then need to be updated, and in some cases the physical or virtual hosts may be shut down (and subsequently restarted) during the transition at step 832 (which may be driven by the techniques disclosed herein for the physical or virtual hosts).

[0232] Fig.8D An example process flow for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for computing and storage is shown. At step 850, the controller 200 creates a virtual machine, container, and / or process that may be on the new node (e.g., see Fig. 8A 8 and 8. The new storage resources on the nodes 318 and 319 in (in) are moved to the new storage resources. At step 852, the old application host can then be cut off. Then, at step 854, the data is copied or synchronized. By shutting down the host at step 852 first and then copying / synchronizing at step 854, the migration will be safer in the case where it involves migrating the VM from a single node. Cutting off will also be beneficial for moving from VM to physical resources. Step 854 can also be completed via a data pre-synchronization step 862 before shutting down, which can help minimize the associated downtime. In addition, the host may not be shut down as at step 852, in which case the old host remains online until the new host is ready (or the new storage resources are ready). The technology for avoiding cutting off step 852 will be discussed in more detail below. At step 854, unless the storage resources are mirrored or synchronized using hot standby, the data can be optionally synchronized.

[0233] The new storage resource is now operational and may be recorded in the database at step 856, enabling the controller 200 to connect the new host to the new storage resource at step 858. When migrating from a single node with multiple virtual hosts, this process may need to be repeated for multiple hosts (step 860). The boot order may be determined by the controller logic using the dependencies of the applications (if they are tracked).

[0234] Fig. 8EAnother example process flow for expanding from a single node to multiple nodes in a system is shown. At step 870, a new resource is coupled to the single-node system. The controller may have a set of system rules and / or expansion rules for the system (or the controller may obtain expansion rules based on service operations, templates of the services, and dependencies of services on each other). At step 872, the controller checks such rules for use to facilitate expansion.

[0235] If the new physical resources include storage resources, the storage resources may be moved from a single node or other form of simpler IT system at step 874 (or the storage resources may be mirrored). If the storage resources are moved, the computing resources or running resources may be reloaded or restarted at step 876 after the storage resources are moved. In another example, the computing resources may be connected to the mirrored storage resources at step 876 and the computing resources may be kept running, while the old storage resources on the single node system or the hardware resources of the previous system may be disconnected or disabled. For example, a running service may be coupled to two mirrored block devices - one mirrored block device on a single node server (e.g., using mdadm raid 1) and the other mirrored block device on the storage resource; and once the data is synchronized, the drive on the single node server may be disconnected. The previous hardware may still include multiple parts of the IT system and the system may be run in a hybrid mode with the controller on the same node (step 878). The system may continue to iterate through this migration process until the original node drives only the controller, whereupon the system is distributed (step 880). In addition, in Fig. 8E At each step of the process flow, the controller may update the system status 220 and record any changes to the system in the database (step 882).

[0236] refer to Fig. 9A , application 910 is installed on resource 900. Resource 900 may be as described herein. Figures 1 to 10 The computing resource 310, storage resource 410, or networking resource 610 described herein. Resource 900 may be a physical resource. A physical resource may include a physical machine or a physical IT system component. Resource 900 may be, for example, a physical computing resource, a storage resource, or a networking resource. Resource 900 may be similar to the one described herein. Figures 2A to 10 Other computing resources, networking resources, or storage resources described are coupled together to the controller 200 in the system 100 .

[0237] The resource 900 may be initially powered off. The resource 900 may be coupled to the controller via the following networks: an out-of-band management connection 260, an in-band management connection 270, a SAN 280, and / or a network 290. The resource 900 may also be coupled to one or more application networks 390 where services, application users, and / or clients may communicate with each other. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 915 or circuitry of the resource 900 that is turned on when the resource 900 is plugged in. The device may allow features including, but not limited to: powering on / off devices, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings 195 and other features outside the scope of the operating system.

[0238] The controller 200 can detect the resource 900 through the out-of-band management network 260. The controller can also identify the type of resource and use in-band management or out-of-band management to identify the configuration of the resource. The controller logic 205 can be configured to carefully check the attached hardware in the out-of-band management 260 or the in-band management 270. If the resource 900 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource 900 will be configured automatically or by interacting with the user. If the resource is added automatically, the setting will follow the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the computing resource. The controller 200 can query the API application or otherwise request the user or any program of the control stack to confirm that the new resource has been authorized. The authorization process can also be automatically and securely completed using cryptography to confirm the legitimacy of the new resource. The resource 900 is then added to the IT system state 220, which includes the switch or network into which the resource 900 is inserted.

[0239] The controller 200 may drive the resource through the out-of-band management network 260. The controller 200 may use the out-of-band management connection 260 to drive the physical resource and configure the BIOS 195. The controller 200 may automatically use the console 190 and select the desired BIOS options, which may be accomplished by the controller 200 reading the console image using image recognition and controlling the console 190 through out-of-band management. The boot state may be determined by the console of the resource 900 through image recognition, or by out-of-band management querying the services being listened on the resource using a virtual keyboard, or querying the services of the application 910. Some applications may have a process that allows the controller 200 to monitor settings in the application 910, or in some cases change the settings using in-band management 270.

[0240] On physical resource 900 (or as described herein with respect to Figures 1 to 10The application 910 of the described resources 300, 310, 311, 312, 313, 400, 410, 411, 412, 600, 610) can be started through SAN 280 or another network using BIOS boot options or configuring remote boot, such as other methods of enabling PXE boot or Flex boot. Additionally or optionally, the controller 200 can use out-of-band management 260 and / or in-band management connection 270 to instruct the physical resource 900 to start the application image in the image 950. The controller can configure the boot options for the resource, or can use existing enabled remote boot methods, such as PXE boot or Flex boot. The controller 200 can optionally or optionally use out-of-band management 260 to boot from the ISO image, configure the local disk, and then instruct the resource to boot from one or more local disks 920. One or more local disks can load boot files. This can be done using out-of-band management 260, image identification, and a virtual keyboard. The resource may also be installed with a boot file and / or a boot loader. The resource 900 and the application can be started from the image 950 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, via the SAN 280. The global system rules 220 can specify a boot order. For example, the global system rules 220 may require that the resource 900 be started first, and then the application 910 be started. Once the resource 900 is started using the image 950, information related to the resource 900 received through the in-band management connection 270 can also be collected and added to the IT system state 220. The resource 900 can be added to a storage resource pool, and the resource will become a resource managed by the controller 200 and tracked in the IT system state 220. The application 910 can also be started using the image 950 or the application image 956 loaded on the resource 900 in the order specified by the global system rules 220.

[0241] The controller 200 may configure the networked resources 610 using the out-of-band management connection 260 or another connection to connect the application 910 to the application network 390. The physical resource 900 may be connected to a remote storage, such as a block storage resource, such as but not limited to: ISER (ISCSI over RDMA), NVMEOF FCOE, FC or ISCSI, or another storage backend, such as SWIFT, GFUSTER or CEPHFS. The IT system state 220 may be updated using the out-of-band management connection 260 and / or the in-band management connection 270 when the service or application is up and running. The controller 200 may use the out-of-band management connection 260 or the in-band management connection 270 to determine the power state of the physical resource 900, i.e., whether it is turned on or off. The controller 200 may use the out-of-band management connection 260 or the in-band management connection 270 to determine whether the service or application is in a running state or in a startup state. The controller may take other actions based on the information it receives and the global system rules 210.

[0242] Fig. 9B An image 950 is shown loaded directly or indirectly (eg, via another resource or database) from template 230 to a computing node to launch application 910. Image 950 may include a custom kernel 941 for application 910.

[0243] Image 950 may include boot files 940 for resource types and hardware. Boot files 940 may include kernels 941 corresponding to resources, applications, or services to be deployed. Boot files 940 may also include initrd or similar file systems for assisting the boot process. Boot system 940 may include multiple kernels or initrds configured for different hardware types and resource types. In addition, image 450 may include file systems 951. File systems 951 may include base images 952 and corresponding file systems, as well as service images 953 and corresponding file systems, and volatile images 954 and corresponding file systems. The loaded file systems and data may vary according to resource types and applications or services to be run. Base images 952 may include base operating system file systems. The base operating system may be read-only. Base images 952 may also include basic tools of the operating system that are independent of what is running. Base images 952 may include base directories and operating system tools. Service file systems 953 may include configuration files and specifications for resources, applications, or services. The volatile file system 594 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configurable as variables, including, but not limited to: passwords, session keys, and private keys. The file system may be mounted as a separate file system using a technology such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.

[0244] Fig. 9C An example of installing an application from an NT package, which may be a type of template 230, is shown. At step 900.1, the controller determines that a package blob needs to be installed. At step 900.2, the controller creates a storage resource on a default data store for the blob type (block, file, file system). At step 900.3, the controller connects to the storage resource via a storage transport available for the storage resource type. At step 900.4, the controller copies the package blob to the connected storage resource. The controller then disconnects from the storage resource (step 900.5) and sets the storage resource to read-only (step 900.6). The package blob is then successfully installed (step 900.7).

[0245] In another example, the attached Appendix B describes example details about how the system connects computing resources to overlayfs. Such techniques can be used to facilitate the following operations: Fig. 9A Install the application on the resource, or Figure 2F Step 205.11 activates computing resources from storage resources.

[0246] Fig.9D An application 910 deployed on a resource 900 is shown. The resource 900 may include a computing node, which may include a virtual computing resource, for example, the virtual computing resource may include a hypervisor 920, one or more virtual machines 921, 922, and / or containers. The resource 900 may be related to Figures 1 to 10A similar approach as described is configured using an image 950 loaded on a resource 900. In this example, the resource 920 is shown as being managed by a hypermanagement system as a virtual machine 921, 922. The controller 200 can use in-band management 270 to communicate with the resource 900 hosting the hypermanagement system 920 to create resources and configure the resources and allocate appropriate hardware resources, including but not limited to: CPU RAM, GPU, remote GPU (which can use RDMA to remotely connect to another host), network connection, network fabric connection and / or virtual and physical connection to partition and / or segmented network. The controller 200 can use a virtual console 190 (for example, including but not limited to SPICE or VNC) and image identification to control the resource 900 and the hypermanagement system 920. Additionally or alternatively, or the controller 200 can use out-of-band management 260 or in-band management connection 270 to instruct the hypermanagement system 920 to start the application image 950 from the template 230 using the global system rule 210. The image 950 may be stored on the controller 200, or the controller 200 may move or copy the image to the storage resource 410. The boot image for the VM 921, 922 may be stored locally as a file, for example, on the image 950, or a block device, or on a remote host, and shared using an image type such as qcow2 or a raw image through a file share, such as NFS over RDMA / NFS, or the boot image may use a remote block device using ISCSI, ISER, NVMEOF, FC, FCOE. Portions of the image 950 may be stored on the storage resource 410 or the compute node 310. The controller 200 may use global rules and / or templates to appropriately configure the networked resources 610 to support the application via the out-of-band management connection 260 or another connection. The application 910 on the resource 900 can be started by using an image 950 loaded via a SAN 280 or another network, using a BIOS boot option or allowing a hypermanagement system 920 on the resource 900 to connect to a block storage resource, such as, but not limited to, ISER (ISCSI over RDMA), NVMEOFFCOE, FC or ISCSI, or another storage backend, such as SWIFT, GFUSTER or CEPHFS. Storage resources can be copied from a template target for a storage resource. The IT system state 220 can be updated by querying the hypermanagement system 920 for information. The in-band management connection 270 can communicate with the hypermanagement system 920 and can be used to determine the power state of the resource, i.e., whether it is connected or disconnected or to determine the startup state. The hypermanagement system 920 can also use a virtual in-band connection 923 connected to the virtualized application 910 and use the hypermanagement system 920 to implement functions similar to out-of-band management.Depending on whether the service or application is driven or started, this information may indicate whether the service or application is started and running.

[0247] The startup state can be determined by mirroring identification via the console 190 of the resource 900, or by out-of-band management 260 using a virtual keyboard to query the services that are listening on the resource, or querying the services of the application 910 itself. Some applications may have processes that allow the controller 200 to monitor settings in the application 910, or in some cases change the settings using in-band management 270. Some applications may be on virtual resources, and the controller 200 may monitor by communicating with the hypervisor 920 using in-band management 270 (or out-of-band management 260). The application 910 may not have such a process for monitoring (or such a process may be turned off to save resources) and / or a process for adding input; in this case, the controller 200 may use the out-of-band management connection 260 and use the mirroring process and / or virtual keyboard to log in to the system to make changes and / or open a management process. Similar to virtual computing resources, a virtual machine console 190 may be used.

[0248] Fig.9E An example process flow for adding a virtual computing resource host to the IT system 100 is shown. At step 900.11, a host capable of being a virtual computing resource is added to the system. The controller may Fig. 15B The process flow may configure a bare metal server (step 900.12); or an operating system may be preloaded and / or the host may be preconfigured (step 900.13). The resource is then added to the system state 220 as a virtual computing resource pool (step 900.14), and the resource becomes accessible to the controller 200 through an API (step 900.15). The API is typically accessed through the in-band management connection 270; however, the in-band management connection 270 may be selectively enabled and / or disabled using a virtual keyboard; and the controller may communicate through the out-of-band connection 260 using an out-of-band management connection 260 and a virtual keyboard and monitor (step 900.16). At step 900.17, the controller may now utilize the new resource as a virtual computing resource.

[0249] Example multi-controller system:

[0250] refer to Fig.10 , shows a system 100, the system 100 having: as described herein with respect to Figures 1 to 10The computing resources 300, 310 described herein include multiple physical computing nodes 311, 312, 313; the storage resources 400, 410 described herein are in the form of multiple storage nodes 411, 412 and JBOD 413; multiple controllers 200a, 200b, the multiple controllers 200a, 200b include components 205, 210, 220, 230 ( Figures 1 to 9C ) and configured like the controller 200 described herein; networked resources 600, 610 as described herein, the networked resources 600, 610 comprising multiple frameworks 611, 612, 613; and an application network 390.

[0251] Fig.10 A possible arrangement of components of system 100 is shown for one example embodiment without limiting the possible arrangements of components of system 100 .

[0252] The user interface or application 110 communicates with an API application 120, which communicates with either or both of the controllers 200a or 200b. The controllers 200a, 200b may be coupled to an out-of-band management connection 260, an in-band management connection 270, a SAN 280, or a networked in-band management connection 290. Figures 1 to 9C As depicted, controllers 200a, 200b are coupled to compute nodes 311, 312, 313, storage 411, 412 (including JBOD 413), and networking resources 610 via connections 260, 270, 280, and optionally 290. Application network 390 is coupled to compute nodes 311, 312, 313, storage resources 411, 412, 413, and networking resources 610.

[0253] Controllers 200a, 200b may operate in parallel. Either controller 200a or 200b may initially be as described herein with respect to Figures 1 to 9CThe controller 200a, 200b is described as operating as the main controller 200. The controllers 200a, 200b can be arranged to configure the entire system 100 from a cut-off state. One of the controllers 200a, 200b can also populate the system state 220 from an existing configuration by probing other controllers via an out-of-band connection 260 and an in-band connection 270. Any of the controllers 200a, 200b can access or receive resource status and related information from resources or other controllers via one or more connections 260, 270. The controller or other resource can update the other controller. Therefore, when an additional controller is added to the system, the additional controller can be configured to restore the system 100 back to the system state 220. In the event of a failure of one of the controllers or the main controller, the other controller can be designated as the main controller. The IT system state 220 may also be able to be rebuilt from status information available or stored on the resource. For example, an application can be deployed on a computing resource, wherein the application is configured to create a virtual computing resource at which the system state is stored or replicated. Global system rules 210, system states 220, and templates 230 may also be saved or copied on a resource or combination of resources. Thus, if all controllers are forced offline and a new controller is added, the system may be configured to allow the new controller to restore the system state 220.

[0254] Networking resources 610 may include multiple network architectures. Fig.10 As shown, the plurality of network architectures may include one or more of the following: SDN Ethernet switch 611, ROCE switch 612, InfiniBand switch 613, or other switch or architecture 614. A hypervisor including virtual machines on a computing node may utilize one or more of the architectures as required to connect to a physical switch or a virtual switch. The networking arrangement may permit, for example, limiting the physical network by segmenting the networking for security or other resource optimization purposes.

[0255] System 100 may be implemented as described herein. Figures 1 to 10The controller 200 described in the above automatically sets up services or applications. A user may request to set up services for the system 100 through the user interface 110 or the application. The services may include, but are not limited to: email services; web services; user management services; network providers; LDAP; Dev tools; VOIP; authentication tools; billing software. The API application 120 translates the user or application request and sends a message to the controller 200. The service template or image 230 of the controller 200 is used to identify which resources are required for the service. The required resources are identified based on availability according to the system state 220. The controller 200 makes a request to the computing resources 310 or computing nodes 311, 312 or 313 for the required computing services, makes a request to the storage resources 410 for the required storage resources, and makes a request to the network resources 610 for the required networking resources. The system state 220 is then updated to identify the resources to be allocated. The global system rules 210 are then used to install the services to the allocated resources according to the service template.

[0256] Enhanced system security:

[0257] refer to Fig.13A , shows an IT system 100, where the system 100 includes a resource 1310, where the resource 1310 can be a bare metal or physical resource. Fig.13A Only a single resource 1310 connected to the system 100 is shown, but it should be understood that the system 100 may include multiple resources 1310. One or more resources 1310 may be or may include bare metal cloud nodes. Bare metal cloud nodes may include, but are not limited to, resources connected to an external network 1380 that allow remote access to a physical host or virtual machine, allow the creation of virtual machines, and allow external users to execute code on one or more resources. One or more resources 1310 may be directly or indirectly connected to the external network 1380 or the application network 390. The external network 1380 may be the Internet or one or more other resources that are not managed by the controller 200 or multiple controllers of the IT system 100. The external network 1380 may include, but is not limited to: the Internet, one or more Internet connections, one or more resources not managed by the controller, other wide area networks (e.g., Stratcom, peer-to-peer mesh networks, or other external networks that may or may not be publicly accessible), or other networks.

[0258] When a physical resource 1310 is added to the IT system 100a, the physical resource is coupled to the controller 200 and the physical resource can be cut off. The resource 1310 is coupled to the controller 200a through one or more of the following networks: an out-of-band management (OOBM) connection 260, an optional in-band management (IBM) connection 270, and an optional SAN connection 280. As used herein, the SAN 280 may or may not include a configuration SAN. The configuration SAN may include a SAN for driving or configuring physical resources. The configuration SAN may be part of the SAN 280 or may be separate from the SAN 280. In-band management may also include a configuration SAN that may or may not be a SAN 280 as described herein. The configuration SAN may also be disabled, disconnected, or unavailable when the resource is used. Although the OOBM connection 260 is not visible to the OS of the system 100, the IBM connection 270 and / or the configuration SAN may be visible to the OS of the system 100. Fig.13A The controller 200 can be used with reference to this article Figures 1 to 12 B. The resource 1310 may include internal storage. In some configurations, the controller 200 may populate the storage and may temporarily configure the resource to connect to a SAN to obtain data and / or information. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or circuitry of the resource 1310 that is turned on when the resource 1310 is inserted. The device 315 may allow features including, but not limited to: powering on / off the device, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings and other features outside the scope of the operating system. The controller 200 can view the resource 1310 through the out-of-band management network 260. The controller can also identify the type of resource and use in-band management or out-of-band management to identify the configuration of the resource. The following discussion FIG. 13C to FIG. 13E Various process flows are shown for adding physical resources 1310 to IT system 100a and / or starting or managing system 100 in a manner that enhances system security.

[0259] The term "disable" as used herein with reference to a network, networking resource, network device, and / or networking interface refers to the action by which such network, networking resource, network device, and / or networking interface achieves the following operations: cut off (manually or automatically), physically disconnected and / or virtually or in some other way (e.g., filtered) from a network, i.e., a virtual network (including but not limited to: VLAN, VXLAN, InfiniBand partitions). The term "disable" also encompasses one-way or unilateral restrictions on operability, such as preventing a resource from sending or writing data to a destination (while still having the ability to receive or read data from the resource), preventing a resource from receiving or reading data from a source (while still having the ability to send or write data to a destination). Such a network, networking resource, network device, and / or networking interface may be disconnected from an additional network, virtual network, or from the coupling of a resource, and remain connected to a previously connected network, virtual network, or resource. In addition, such a networking resource or device may switch from the coupling of one network, virtual network, or resource to another.

[0260] The term "enable" as used herein with reference to a network, networking resource, network device and / or networking interface refers to the action by which such network, networking resource, network device and / or networking interface implements the following operations: driving (manually or automatically), physically connecting and / or virtually or in some other way connecting to a network, i.e., a virtual network (including but not limited to: VLAN, VXLAN, InfiniBand partitioning). Such a network, networking resource, network device and / or networking interface can be connected to the coupling of an additional network, virtual network or resource while already connected to another system component. In addition, such a networking resource or device can switch from the coupling of one network, virtual network or resource to another. The term "enable" also encompasses unidirectional or unilateral permission for operability, such as allowing a resource to send, write data to a destination, or receive data from the destination (while still having the ability to restrict data from a certain source), allowing a resource to send data to a certain source, receive or read data from the source (while still having the ability to restrict data from the destination).

[0261] The controller logic 205 is configured to carefully check the added hardware in the out-of-band management connection 260 or the in-band management connection 270 and / or the configuration SAN 280. If a resource 1310 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource will be configured automatically or by interacting with the user. If the resource is added automatically, the settings will follow the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the resource 1310. The controller 200 can query the API application or otherwise request the user or any program of the control stack to confirm that the new resource has been authorized. The authorization process can also be automatically and securely completed using cryptography to confirm the legitimacy of the new resource. The controller logic 205 then adds the resource 1310 to the IT system state 220, which includes the switch or network into which the resource 1310 is inserted.

[0262] In the case where the resource is physical, the controller 200 can drive the resource through the out-of-band management network 260, and the resource 1310 can be started from the image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, through the SAN 280. The image can be loaded through other network connections or indirectly through another resource. Once started, information related to the resource 1310 can also be collected and added to the IT system state 220. This can be done through in-band management and / or configuring a SAN or out-of-band management connection. The resource 1310 can be started from the image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, through the SAN 280. The image can be loaded through other network connections or indirectly through another resource. Once started, information related to the computing resource 310 received through the in-band management connection 270 can also be collected and added to the IT system state 220. The resource 1310 can then be added to the storage resource pool, and the resource will become a resource managed by the controller 200 and tracked in the IT system state 220.

[0263] In-band management and / or configuration SAN can be used by the controller 200 to set up, manage, use or communicate with the resource 1310 and run any command or task. However, optionally, the in-band management connection 270 can be configured by the controller 200 to be turned off or disabled at any time or during the setting, management, use or operation of the system 100 or the controller 200. In-band management can also be configured to be turned on or enabled at any time or during the setting, management, use or operation of the system 100 or the controller 200. Optionally, the controller 200 can controllably or switchably disconnect the resource 1310 from the in-band management connection 270 connected to one or more controllers 200. This disconnection or disconnectability can be physical, such as using an automated physical switch or some kind of switch to cut off the in-band management connection and / or configuration SAN of the resource to the network. The disconnection can be accomplished, for example, by turning off the power of the port of the in-band management 270 and / or configuration SAN 280 connected to the resource 1310 by a network switch). This disconnection or partial disconnection may also be accomplished using software defined networking, or may be physically filtered out with respect to the controller using software defined networking. This disconnection may be accomplished by the controller via in-band management or out-of-band management. According to an example embodiment, at any point in time before, during, or after the resource 1310 is added to the IT system, the resource 1310 may be disconnected from the in-band management connection 270 in response to a selective control instruction from the controller 200.

[0264] Using software defined networking, the in-band management connection 270 and / or configuration SAN 280 may or may not retain certain functionality. The in-band management 270 and / or configuration SAN 280 may be used as a limited connection for communication to or from the controller 200 or to other resources. The connection 270 may be restricted to prevent attackers from pivoting to the controller 200, other networks, or other resources. The system may be configured to prevent devices such as the controller 200 and the resource 1310 from communicating openly to avoid damaging the resource 1310. For example, the in-band management 270 and / or configuration SAN 280 may only allow the in-band management and / or configuration SAN to deliver data but not receive anything through software defined networking or hardware change methods (such as electronic restrictions). The in-band management and / or configuration SAN may be physically or using software defined networking that only allows writes from the controller to the resource to be configured as a one-way write component or one-way write connection from the controller 200 to the resource 1310. The one-way write nature of the connection may also be controlled or turned on or off according to the expectations for security and different stages or times of system operation. The system may also be configured so that writing or communication from the resource to the controller is limited to, for example, conveying logs or alarms. Interfaces may also be moved to other networks or added and removed from networks using techniques including, but not limited to, software-defined networking, VLANS, VXLANS, and / or InfiniBand partitioning. For example, an interface may be connected to a setup network, removed from the network, and moved to a network for runtime. Communication from the controller to the resource may be cut off or limited so that the controller may not physically be able to respond to any data sent from the resource 1310. According to one example, once the resource 1310 is added and started, the in-band management 270 may be turned off, or filtered out physically or using software-defined networking. In-band management may be configured so that it can send data to another resource dedicated to log management.

[0265] In-band management can be turned on and off using out-of-band management or software-defined networking. In the event that in-band management is disconnected, the daemon may not need to be running and keyboard functionality can be used to re-enable in-band management.

[0266] Additionally, optionally, resource 1310 may not have an in-band management connection, and the resource may be managed via out-of-band management.

[0267] Alternatively or in addition, out-of-band management can be used to manipulate various aspects of the system through methods including, but not limited to, for example, keyboards, virtual keyboards, disk mount consoles, attaching virtual disks, changing bios settings, changing boot parameters and other aspects of the system, running existing scripts that may exist on a bootable image or installation CD, or other features of out-of-band management that allow the controller 200 and the resource 1310 to communicate with or without exposing the operating system running on the resource 1310. For example, the controller 200 can send commands using such tools through out-of-band management 260. The controller 200 can also use image identification to assist in controlling the resource 1310. Thus, using an out-of-band management connection, the system can prevent or avoid undesired manipulation of the resource connected to the system through the out-of-band management connection. The out-of-band management connection can also be configured as a one-way communication system during operation of the system or at selected times during operation of the system.

[0268] Additionally, if desired by a practitioner, the out-of-band management connection 260 may also be selectively controlled by the controller 200 in the same manner as the in-band management connection.

[0269] The controller 200 may be able to automatically turn resources on and off according to global system rules and update the IT system status for reasons determined by the IT system user, such as turning off resources to save power, or turning on resources to improve application performance, or any other reason the IT system user may have. The controller may also be able to turn the configuration SAN, in-band management connections, and out-of-band management connections on and off, or designate such connections as unidirectional write connections (e.g., disabling the in-band management connection 270 or the configuration SAN 280 when the resource 1310 is connected to the external network 1380 or the internal network 390) at any time during system operation and for various security purposes. One-way in-band management may also be used, for example, to monitor the health of the system, monitoring logs and information that may be visible to the operating system.

[0270] Resources 1310 may also be coupled to one or more internal networks 390, such as application networks, where services, application users, and / or clients may communicate with each other. Such application networks 390 may also be connected to or capable of connecting to external networks 1380. Figures 2A to 12 In an example embodiment of B, in-band management can be disconnected, can be disconnected from the resource or application network 390, or can provide one-way writes from the controller to provide additional security in the event that the resource or application network is connected to an external network, or in the event that the resource is connected to an application network that is not connected to an external network.

[0271] Fig.13A The IT system 100 may be configured similarly to Figure 3B The IT system 100 shown; the image 350 can be loaded directly or indirectly (through another resource or database) from the template 230 to the resource 1310 to start the computing resource and / or load the application. The image 350 may include a boot file 340 for the resource type and hardware. The boot file 340 may include a kernel 341 corresponding to the resource, application or service to be deployed. The boot file 340 may also include initrd or a similar file system for assisting the boot process. The boot system 340 may include multiple kernels or initrds configured for different hardware types and resource types. In addition, the image 350 may include a file system 351. The file system 351 may include a base image 352 and a corresponding file system, a service image 353 and a corresponding file system, and a volatile image 354 and a corresponding file system. The loaded file system and data may vary according to the resource type and the application or service to be run. The base image 352 may include a base operating system file system. The base operating system may be read-only. The base image 352 may also include basic tools of the operating system that are independent of what is running. The base image 352 may include a base directory and operating system tools. The service file system 353 may include configuration files and specifications for resources, applications, or services. The volatile file system 354 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to: passwords, session keys, and private keys. The file system may be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.

[0272] Fig. 13B A plurality of resources 1310 are shown, each of which includes one or more hypervisor management systems 1311 hosting or including one or more virtual machines. The controller 200a is coupled to the resources 1310, each of which includes bare metal resources. Fig. 13B As shown and described, resources 1310 are each coupled to controller 200a. According to example embodiments herein, in-band management connection 270, configuration SAN 280, and / or out-of-band management connection 260 may be similar to those described with respect to Fig.13A200a, and from the damaged controller 200a to other hypermanagements coupled to the controller 200a. For example, the transfer may occur between the damaged hypermanagement system and the target hypermanagement system using a network connected to both. Fig. 13B The illustrated arrangement of in-band management 270, configuration SAN 280, or out-of-band management 260 of the controller 200a and resource 1310, wherein either or both of the controller and resource can be selectively controlled to disable in-band connections (or configuration SAN) and / or out-of-band connections in a given link between the controller 200a and the resource 1310, can prevent a damaged virtual machine being used from escaping one hypervisor and being transferred to other resources.

[0273] The above article about Figures 1 to 12 The in-band management connection 270 and out-of-band management connection 260 described above may also be used in conjunction with Fig.13A and Fig. 13B Configure in a similar manner as described.

[0274] Fig. 13C 2 shows an example process flow for adding physical resources such as bare metal nodes to the system 100, or managing the physical resources. The management system 100 described herein can be connected via an out-of-band management connection 260 and an in-band management connection 270 and / or a SAN. Fig.13A and Fig. 13B As shown or as about Figures 1 to 12 Resource 1310 is shown connected to a controller of system 100 .

[0275] After connecting to the instance of the resource, the external network and / or application network is disabled at step 1370. As described above, any of a variety of techniques may be used for such disabling. For example, before using an in-band management connection or configuring a SAN to set up the system, add the resource, test the system, update the system, or perform other tasks or commands, such as with respect to Fig.13A and Fig. 13B As described, components of system 100 (or only those components that are vulnerable) are disabled, disconnected from any external network or application network, or filtered out of any external network or application network.

[0276] After step 1370, the in-band management connection and / or configuration SAN is then enabled at step 1371. The combination of steps 1370 and 1371 thus isolates the resource from the external network and / or application network while the in-band management and / or SAN connection is active. Commands may then be run on the resource via the in-band management connection under the control of controller 200 (see step 1372). For example, setup and configuration steps (such as, but not limited to, those described herein with respect to Figures 1 to 13B Alternatively or in addition, in-band management and / or configuration of the SAN may be used at step 1372 to perform other tasks including, but not limited to: operating, updating or managing the system (which may include, but is not limited to, any change management or system updates), testing, updating, transferring data, collecting information about performance and health (including, but not limited to, errors, cpu usage, network usage, file system information, and storage usage), and collecting logs and other information that may be used as described herein. Figures 1 to 13B Other commands for managing system 100 are described in .

[0277] After adding the resources, setting up the system and / or executing such tasks or commands, the system may be started at step 1373 as described herein. Fig.13A and Fig. 13B The described in-band management connection 270 and / or configuration SAN 280 between the resource and the controller or other components of the system is disabled in one or more directions. This disabling may be achieved by disconnection, filtering, etc. as described above. After step 1373, the connection to the external network and / or application network may be restored at step 1374. For example, the controller may inform the networked resource that the resource 1310 is allowed to connect to the application network or the Internet. The same steps may be followed in the case of testing or updating the system, that is, the in-band management connection, or the external network and / or application network may be disconnected or filtered before enabling the in-band management connection or connecting the in-band management connection (one-way or two-way) to the resource. Therefore, steps 1373 and 1374 operate together to prevent the resource from connecting to the controller via the in-band management connection and / or configuration SAN when the resource is connected to the external network and / or application network.

[0278] Out-of-band management can be used to manage systems or resources, set up systems or resources, configure, start or add systems or resources. Out-of-band management can use a virtual keyboard to send commands to the machine to change settings before startup in any embodiment of the present invention, and can also send commands to the operating system by inputting to the virtual keyboard; if the machine is not logged in, the out-of-band management can use the virtual keyboard to enter the username and password and can use image recognition to verify the login, and verify the commands it enters and check to see if the commands are executed. If the physical resource only has a graphical console, a virtual mouse can also be used and image recognition will allow out-of-band management to make changes.

[0279] Fig.13D Another example process flow for adding physical resources such as bare metal nodes to the system 100 or managing the physical resources is shown in FIG. 1380. At step 1380, the physical resources described herein may be managed by out-of-band management 260. Fig.13A and Fig. 13B As shown or Figures 1 to 12The resource shown is connected to a certain system or resource. The disk can be virtually connected by providing access to the disk image (e.g., ISO image) via out-of-band management with the help of a controller (see step 1381). The resource or the system can then be started from the disk image (step 1382), and then the file is copied from the disk image to the bootable disk (see step 1383). This can also be used to start a system in which resources are set up in this way using out-of-band management. This can also be used to configure and / or start multiple resources that may be coupled together (including but not limited to coupling using networked resources), regardless of whether the multiple resources also include a controller or constitute a system. Therefore, the virtual disk can be used to allow the controller to connect the disk image to the resource, just like attaching the virtual disk to the resource. Out-of-band management can also be used to send files to the resource. _Data can be copied from the virtual disk to the local disk at step 1383. The disk image may contain files that the resource can copy and use in its operation. Files can be copied or used by scheduled programs or instructions from out-of-band management. The controller can log on to the resource using a virtual keyboard through out-of-band management, and enter commands to copy files from the virtual disk to its own disk or other storage accessible to the resource. At step 1384, the system or resource is configured to start by setting the bios, efi or boot order settings, so the system or resource will be started from the bootable disk. The boot configuration can use an EFI manager such as efibootmgr in the operating system, which can be directly run through out-of-band management or by including it in the installer script (for example, when the resource starts, the resource will automatically run the script using efibootmgr). In addition, the boot options and any other bios changes can be set by an out-of-band management tool such as Supermicro Boot Manager using a boot order command or uploading a bios configuration (such as an XML BIOS configuration supported by Supermicro Update Manager). The bios can also be configured from the console using a keyboard and image recognition to set appropriate bios settings including the boot order. The installer can run on the loaded pre-configured image. The configuration can be tested by watching the screen and using image recognition. After configuration, the resource may then be enabled (eg, driven, started, connected to an application network, or a combination thereof) (step 1385).

[0280] Fig.13E Another example process flow for adding physical resources such as bare metal nodes to the system 100, or managing the physical resources, in this case using PXE, Flex boot, or similar network boot, is shown. At step 1390, the management system 100 described herein may be connected to the management system 100 via (1) in-band management connection 270 and / or SAN and (2) out-of-band management connection 260. Fig.13A and Fig. 13BAs shown or as about Figures 1 to 12 The resources 1310 shown are connected to the controller of the system 100. The external network and / or application network connection may then be disabled (e.g., physically, filtered out or disconnected in whole or in part using an SDN or virtually) at step 1391 (similar to what was discussed above with respect to step 1370). For example, before setting up the system, adding resources, testing the system, updating the system, or performing other tasks or commands using an in-band management connection or SAN, as described with respect to Fig.13A and Fig. 13B As described, components of system 100 (or only those components that are vulnerable) are disabled, disconnected from any external network or application network, or filtered out of any external network or application network.

[0281] At step 1392, the type of resource is determined. For example, an out-of-band management tool may be used, or information about the resource may be collected from the mac address by connecting a disk image (e.g., an ISO image) to the resource as if the disk were attached to the resource, to temporarily boot an operating system that has tools that can be used to identify resource information. Then, at step 1393, the resource is configured or identified as pre-configured for PXE or flex boot, etc. Thereafter, at step 1394, the resource is driven to perform a PXE, Flex boot, or similar boot (or is in a situation where it is temporarily started and the resource is driven again). Then, at step 1395, the resource is booted from an in-band management connection or SAN, or the resource is booted from the in-band management connection or SAN. At step 1396, the resource is connected to the reference Fig.13D The data is copied to the disk accessible by the resource in a similar manner as described in step 1383 of the embodiment. Then, in step 1397, the data is copied to the disk accessible by the resource in a similar manner as described in the embodiment of the embodiment of the invention. Fig.13D The resource is configured to boot from one or more disks in a similar manner as described in step 1384 of . In the event that the resource is identified as pre-configured for PXE, flex boot, etc., files may be copied at any step from 1393 to 1396. If in-band management is enabled, it may be disabled at step 1398, and the application network or external network may be reconnected or enabled at step 1399.

[0282] Further, it should be understood that techniques other than OOBM may be used to remotely enable (such as drive) resources and verify that the resources have been started. For example, the system may prompt the user to press a power button and manually inform the controller that the system has been started (or using a keyboard / console connection to the controller). In addition, once the system has been started, the system can check the controller through IBM, and the controller logs in and tells the system to restart (for example, through methods such as ssh, telnet, or another method implemented on the network). For example, the controller can import and send a restart command through ssh. If PXE is being used and OOBM is not present, then in any case, the system should have a way to remotely instruct the resource to drive, or tell the user to manually drive the resource.

[0283] Deploy controllers and / or environments:

[0284] In an example embodiment, controllers may be deployed within a system from an originating controller 200 (where such an originating controller 200 may be referred to as a "master controller"). Thus, the master controller may set up a system or environment that may be an isolated or isolatable IT system or environment.

[0285] An environment as described herein refers to a collection of resources within a computer system that can interoperate with each other. A computer system may include multiple environments within it; however, this need not be the case. One or more resources of an environment may include one or more instances, applications, or sub-applications running on the environment. Further, an environment may include one or more environments or sub-environments. An environment may or may not include a controller, and the environment may operate one or more applications. Such resources of an environment may include, for example, networking resources, computing resources, storage resources, and / or application networks for running a specific environment, including applications in the environment. Therefore, it should be understood that an environment may provide the functionality of one or more applications. In some instances, the environment described herein may or can be physically or virtually separated from other environments. In addition, in other instances, an environment may have a network connection to other environments, wherein such a connection may be disabled or enabled as needed.

[0286] In addition, the main controller can be set, deployed and / or managed in various environments or in a separate system. Such additional controllers can be or become independent of the main controller. Even if independent or quasi-independent of the main controller, such additional controllers can also obtain instructions from the main controller (or a separate monitor or an environment using a monitoring application) or send information to the main controller at various times during operation. The environment can be configured for security purposes (for example, by enabling the environment to be isolated from each other and / or from the main controller) and / or for various management purposes. A certain environment can be connected to an external network, while another related environment can be connected or not connected, or connected or not connected to an external network.

[0287] The master controller can manage an environment or application, regardless of whether the environment or application is a separate system, and regardless of whether the environment or application includes a controller or a sub-controller. The master controller can also manage shared storage of global configuration files or other data. The master controller can also parse global system rules (e.g., system rules 210) or a subset thereof to different controllers according to their functions. Each new controller (the new controller can be referred to as a "sub-controller") can receive new configuration rules that may be a subset of the configuration rules of the master controller. The subset of global configuration rules deployed to the controller can depend on or correspond to the type of IT system being set up. The master controller can set up or deploy a new controller or a separate IT system, which is then permanently separated from the master controller, for example, for shipping or distribution or other reasons. Global configuration rules (or a subset thereof) can define a framework for setting up applications or sub-applications in various environments, and how the applications or sub-applications can interact with each other. Such applications or environments can run on sub-controllers that include a subset of global configuration rules deployed by the master controller. In some instances, such applications or environments can be managed by the master controller. However, in other instances, such applications or environments are not managed by the primary controller. If a new controller is being generated from the primary controller to manage the application or environment, a dependency check of the application may be performed across multiple applications to facilitate control by the new controller.

[0288] Thus, in one example embodiment, a system may include a main controller configured to deploy another controller, or an IT system including such other controller. Such an implemented system may be configured to be completely disconnected from the main controller. Once independent, such a system may be configured to operate as an independent system; or the system may be controlled or monitored by another controller such as the main controller (or an environment with an application) at various discrete or continuous times during operation.

[0289] Fig.14AAn example system is shown in which a master controller 1401 has deployed controllers 1401a and 1401b on different systems 1400a and 1400b, respectively (where 1400a and 1400b may be referred to as subsystems; however, it should be understood that subsystems 1400a and 1400b may also be used as environments). The master controller 1401 may be configured in a similar manner to the controller 200 discussed above. Thus, the master controller may include controller logic 205, global system rules 210, system state 220, and templates 230.

[0290] Systems 1400a and 1400b include controllers 1401a, 1401b coupled to resources 1420a, 1420b, respectively. Master controller 1401 may be coupled to one or more other controllers, such as controller 1401a of subsystem 1400a and controller 1401b of subsystem 1400b. Global rules 210 of master controller 1400 may include rules that may manage and control other controllers. Master controller 1401 may use such global rules 210 in conjunction with controller logic 205, system state 220, and templates 230 to communicate with controllers 1401a, 1401b in accordance with the present disclosure. Figures 1 to 13E Subsystems 1400a, 1400b are set up, provisioned, and deployed in a similar manner as described.

[0291] For example, the master controller 1401 may load the global rules 210 (or a subset thereof) as rules 1410a, 1410b onto the subsystems 1400a, 1400b, respectively, in the following manner: The global rules 210 (or a subset thereof) indicate the operation of the controllers 1401a, 1401b and their subsystems 1400a, 1400b. Each controller 1401a, 1401b may have rules 1410a, 1410b that may be the same or different subsets of the global rules 210. For example, which subset of the global rules 210 is supplied to a given subsystem may depend on the type of subsystem being deployed. The controller 1401 may also load or direct data to be loaded to the system resources 1420a, 1420b or the controllers 1401a, 1401b.

[0292] The master controller 1401 may be connected to the other controllers 1401a, 1401b via one or more in-band management connections 270 and / or one or more out-of-band management connections 260 or SAN connections 280, which may be connected at various stages of deployment or management in the manner described herein, such as with reference to FIG. 13A to FIG. 13EUsing the selective enabling and disabling of the in-band management connection 270 or the out-of-band management connection 260, the subsystems 1400a, 1400b may be deployed in a manner where the subsystems 1400a, 1400b may have no knowledge (or have limited, controlled, or restricted knowledge) of the main system 100 or the controller 1401 or about each other at various times.

[0293] In an example embodiment, the main controller 1401 can operate a centralized IT system having local controllers 1401a, 1401b deployed and configured by the main controller 1401, so that the main controller 1401 can deploy and / or run multiple IT systems. Such IT systems may be independent or not independent of each other. The main controller 1401 can set monitoring to a separate application that is isolated or isolated from the IT system it has created. A separate console for monitoring can be provided to be connected between the main controller and one or more local controllers and / or to be connected between environments that can be selectively enabled or disabled. The controller 1401 can deploy, for example, an isolation system for various purposes, including but not limited to: a commercial system, a manufacturing system with data storage, a data center, and other different functional nodes, each of which has a different controller to prevent interruption of operation or damage. This isolation can be thorough or permanent, or it can be quasi-isolation, for example, temporary, time or task-related, communication direction-related, or other parameter-related. For example, the main controller 1401 may be configured to provide instructions to the system that may be limited or not limited to certain predefined situations, while the subsystem may have limited or no ability to communicate with the main controller. Therefore, such a subsystem may not be able to harm the main controller 1401. The main controller 1401 and the subcontrollers 1401a, 1401b may be separated from each other, for example, as described herein (with specific examples discussed below), by disabling in-band management 270, by one-way writes, and / or by limiting communications to out-of-band management 260. For example, if a vulnerability occurs, one or more controllers may disable the in-band management connection 270 with respect to one or more other controllers to prevent the spread of the vulnerability or access. System segments may be shut down or isolated.

[0294] The subsystems 1400a, 1400b may also share resources with or connect to another environment or system through in-band management 270 or out-of-band management 260.

[0295] Fig. 14B and Fig. 14C is an example flow showing possible steps for supplying a sub-controller by a main controller.

[0296] exist Fig. 14BIn step 1460, the master controller supplies or sets up resources, such as resource 1420a or 1420b. At step 1461, the master controller supplies or sets up the sub-controller. The master controller can perform steps 1460 and 1461 using the techniques discussed above for setting up resources within the system. In addition, although Fig. 14B Step 1460 is shown to be performed before step 1461, but it should be understood that this need not be the case. Using its system rules 210, the main controller 1401 can determine which resources are needed and locate the resources on the system or network. The main controller can set up or deploy the sub-controller by loading the system rules 210 onto the system (or by providing instructions to the sub-controller on how to set up and obtain its own system rules) at step 1461. These instructions may include, but are not limited to: configuring resources, configuring applications, global system rules for creating IT systems run by sub-controllers, instructions for reconnecting to the main controller to collect new or changed rules, and instructions for disconnecting from the application network to make room for a new production environment. After deploying the resources, at step 1463, the main controller can then dispatch the resources to the sub-controllers via updates to the system rules 210 and / or system status 220.

[0297] Fig. 14C An alternative process flow for deployment is shown. Fig. 14C In the example of , the main controller deploys the sub-controller at step 1470 (which can be done as described with respect to step 1461). Then, at step 1475, the sub-controller uses a Figure 3C and Figure 7B The resources are deployed using the techniques shown.

[0298] Fig.15A An example system is shown in which a master controller 1501 for system 100 generates environments 1502, 1503, and 1504. Environment 1502 includes resources 1522, environment 1503 includes resources 1523, and environment 1504 includes resources 1524. In addition, environments 1502, 1503, 1504 may share access to a shared resource pool 1525. Such shared resources may include, but are not limited to, for example, shared data sets, APIs, or application programs running that need to communicate with each other.

[0299] exist Fig.15AIn the example of , each environment 1502, 1503, 1504 shares the master controller 1501. The global system rules 210 of the master controller 1501 may include rules for deploying and managing environments. Resources 1522, 1523 and / or 1524 may be required by their respective environments 1501, 1502, 1503 to manage one or more applications. Configuration rules for such applications may be implemented by the master controller (or by local controllers in the environment, if any) to define how each such environment operates and interacts with other applications and environments. The master controller 1401 may use the global rules 210 in conjunction with the controller logic 205, system state 220, and templates 230 to configure the environment in a manner consistent with the present invention. Figures 1 to 14C The described resources and system deployment are similarly configured, provisioned and deployed to the environment. If the environment includes a local controller, the master controller 1501 can load the global rules 210 (or a subset thereof) onto the local controller or associated storage in a manner that the global rules (or a subset thereof) define the operation of the environment.

[0300] The controller 1501 may deploy and configure the resources 1522, 1523, 1524 and / or shared resources 1525 of the environments 1502, 1503, 1504 using configuration rules and system rules 210 accordingly. The controller 1501 may also monitor the environments, or configure the resources 1522, 1523, 1524 (or shared resources 1525) to allow monitoring of the respective environments 1502, 1503, 1504. Such monitoring may be performed using a connection to a separate monitoring console that may be enabled or disabled, or may be performed through the master controller. The master controller 1501 may be connected to one or more of the environments 1502, 1503, 1504 via one or more in-band management connections 270 and / or one or more out-of-band management connections 260 or SAN connections 280, which may be used at various stages of deployment or management as described herein. FIG. 13A to FIG. 13E and Fig.14A Using the enabling and disabling of the in-band management connection 270 or the out-of-band management connection 260 or the SAN connection 280, the environments 1502, 1503, 1504 may be deployed at various times in a manner that may have no knowledge or limited or controlled knowledge, or no connection or limited or controlled connection, with respect to each other or to the host system 100 or the controller 1501.

[0301] The environment may include a resource or resources that are coupled or interact with other resources, or coupled to an external network 1580 that is connected to an external, external environment. The environment may be physical or non-physical. Non-physical in this context means that the environments share one or more of the same physical hosts, but are virtually separated from each other. The environment and system may be deployed on the same, similar but different, or different hardware. In some instances, environments 1502, 1503, 1504 may be valid copies of each other; however, in other instances, environments 1502, 1503, 1504 may provide functionality that is different from each other. As an example, the resources of an environment may be servers.

[0302] Placing systems and resources in separate environments or subsystems according to the techniques described herein can allow applications to be isolated for security and / or for performance reasons. Separated environments can also mitigate the impact of compromised resources. For example, one environment may contain sensitive data and may be configured for less Internet exposure, while another environment may host Internet-facing applications.

[0303] Fig. 15B It shows that Fig.15A The controller shown sets up an example process flow for an environment. In this example, the system may be tasked with creating and setting up a new environment. This may be triggered by a user request or by a system rule that is executed when engaging in a specific task or a specific series of tasks. FIG. 17A to FIG. 18B An example of a specific change management task or a specific series of tasks is shown where the system creates a new environment. However, there may be a variety of situations where a controller may create and set up a new environment.

[0304] Therefore, reference Fig. 15B In setting up a new environment, the controller selects environment rules (step 1500.1). Based on the environment rules, using global system rules 210 and templates 230, the controller searches for resources for the environment (step 1500.2). The rules may have a hierarchy of preferred resource selections that is kept going through until the resources needed for the environment are found. At step 1500.3, the controller uses, for example, Figure 3C or Figure 7BThe technology described in the above allocates the resources found at step 1500.2 to the environment. The controller then configures the networked resources of the system with respect to the new environment to ensure compatible and effective connections between the new environment and other system components (step 1500.4). The system status is updated at step 1500.5 to indicate that each resource is enabled and each template is processed. The controller then sets up and implements the integration and interoperability of the resources of the environment and drives any application to deploy the new environment (step 1500.6). The system status is again updated at step 1500.7 to indicate that the environment has become available.

[0305] Fig. 15C It shows that Fig.15A The controller shown in the figure sets up an example process flow for multiple environments. When setting up multiple environments, you can use Fig. 15B However, it should be understood that the environment can be set up in parallel as described in Fig. 15C The description sets the environment in an ordered sequence or in a series. Fig. 15C At step 1500.10, the controller sets up and deploys the first new environment (this may be similar to Fig. 15B 1). For different types of environments and for how different environments interoperate, there may be different environment rules. At step 1500.11, the controller selects environment rules for the new environment. At step 1500.12, the controller searches for resources according to a preference order that can be defined by system rules 210. At step 1500.13, the controller allocates the resources found at step 1500.12 to the next environment. The environments may share or not share resources. At step 1500.14, the controller uses system rules 210 to configure the networked resources of the system between the next environment and the environments with dependencies. At step 1500.15, the system status is updated to each resource being enabled, each template being processed and the networked resources being configured, including the dependencies of the environments. The controller then sets up and implements the integration and interoperability of the resources between the next environment and the environments, and drives any application to deploy the new environment (step 1500.16). At step 1500.17, the system status is updated to the next environment becoming available.

[0306] One-way communication to support monitoring:

[0307] Fig.16AAn example embodiment is shown in which a first controller 1601 operates as a master controller to set up one or more controllers, such as 1601a, 1601b, and / or 1601b. The master controller 1601 can be used to generate multiple clouds, hosts, systems, and / or applications as environments 1602, 1603, 1604 that may or may not be dependent on each other in their operation using the techniques discussed above with respect to controllers, such as controllers 200 / 1401 / 1501. Fig.16A As shown, IT systems, environments, clouds, and / or any one or more combinations thereof may be generated as environments 1602, 1603, 1604. Environment 1602 includes a second controller 1601a, environment 1603 includes a third controller 1601b, and environment 1604 includes a fourth controller 1601c. Environments 1602, 1603, 1604 may each further include one or more resources 1642, 1643, 1644, respectively. Resources may include one or more applications 1642, 1643, 1644 running thereon. These applications may be connected to allocated resources, whether shared or not. These or other applications may run on one or more shared resources in the Internet or pool 1660, which may also include shared applications or application networks. Applications may provide services to users or one or more of the environments or clouds. Environments 1602, 1603, 1604 may share resources or databases, and / or may include or use resources in pool 1660 that are specifically allocated to a particular environment. The various components of the system, including the main controller 1601 and / or one or more environments, may also be able to connect to an application network or external network 1615 such as the Internet.

[0308] Between any resource, environment, or controller and another resource, environment, controller, or external connection, there may be a FIG. 13A to FIG. 13E 1604, a resource, or an application. FIG. 13A to FIG. 13EFor security purposes discussed, disabling environments 1602, 1603, 1604 or disconnecting the main controller 1601 from the environments may allow the main controller 1601 to transform environments 1602, 1603, 1604 into clouds that may then be separated from the main controller 1601 or other clouds or environments. In this sense, the controller 1601 is configured to generate multiple clouds, hosts, or systems.

[0309] Using the disabling or disconnecting of the elements described herein, a user may be allowed limited access to an environment through the main controller 1601 for a specific purpose. For example, a developer may be provided access to a development environment. As another example, an administrator of an application may be restricted to a specific application or application network. As another example, a log may be consulted by the main controller 1601 to collect data without subjecting the main controller itself to damage by the environment or controller it generates.

[0310] After the main controller 1601 sets up the environment 1602, the environment 1602 can be disconnected from the main controller 1601, so that the environment 1602 can operate independently of the main controller 1601 and / or can be selectively monitored and maintained by the main controller 1601 or other applications associated with or run by the environment 1602.

[0311] An environment such as environment 1602 may be coupled to a user interface or console 1640 that allows a purchaser or user to access environment 1602. Environment 1602 may host a user console as an application. Environment 1602 may be accessed remotely by a user. Each environment 1602, 1603, 1604 may be accessed by a common or separate user interface or console.

[0312] Fig. 16B An example system is shown in which environments 1602, 1603, 1604 can be configured to write to another environment 1641 where logs can be viewed, for example using a console (which can be any console that can be directly or indirectly connected to environment 1641). In this way, environment 1641 can be used as a log server to which one or more of environments 1602, 1603, 1604 write events. The main controller 1601 can then access the log server 1641 to monitor events on environments 1602, 1603, 1604 without maintaining a direct connection with such environments 1602, 1603, 1604 as described below. Environment 1641 can also be selectively disconnected from the main controller 1601 and can be configured to only read from other environments 1602, 1603, 1604.

[0313] The main controller 1601 may be configured to Fig. 16CAs shown, a device disconnected from any of its environments 1602 , 1603 , 1604 is also able to monitor some or all of its environments 1602 , 1603 , 1604 . Fig. 16C The in-band management connection 270 between the master controller 1601 and the environments 1602, 1603, 1604 is shown to have been disconnected, which can help protect the master controller 1601 in the event that the environments 1602, 1603, 1604 are compromised. Fig. 16C As shown, even if the in-band connection 270 between the main controller 1601 and the environment 1602 has been disconnected, the out-of-band connection 260 between the main controller 1601 and the environment such as 1602 can still be maintained. In addition, the environment 1641 can have a connection to the main controller 1601 that can be selectively enabled or disabled. The main controller 1601 can set monitoring as a separate application within the environment 1641 that is isolated or isolated from the environments 1602, 1603, 1604. The main controller 1601 can use one-way communication to perform monitoring. For example, logs can be provided from the environments 1602, 1603, 1604 to the environment 1641 through one-way communication. Through this one-way writing and connection between the environment 1641 and the main controller 1601, even though there is no in-band connection 270 between the main controller 1601 and the environments 1602, 1603, 1604, the main controller 1601 can also collect data through the environment 1641 and monitor the environments 1602, 1603, 1604, thereby reducing the risk of the environments 1602, 1603, 1604 damaging the main controller 1601. The access may be filtered or controlled and / or the access may be independent of the Internet. For example, Fig.16D As shown, if the in-band connection 270 between the main controller 1601 and the environment 1602 is connected, the main controller 1601 can control the network switch 1650 to disconnect the environment 1602 from the external network 1615 such as the Internet. When the environment 1602 is connected to the main controller 1601 through the in-band connection 270, the disconnection of the environment 1602 from the external network 1615 can provide enhanced security for the main controller 1601.

[0314] Therefore, it should be understood that FIG. 16B to FIG. 16DThe example embodiment shows how the master controller can securely monitor environments 1602, 1603, 1604 while minimizing exposure to the environments 1602, 1603, 1604. Thus, the master controller 1601 can disconnect itself (or at least disconnect itself from the in-band link) from the environments 1602, 1603, 1604, while still maintaining a mechanism for monitoring the environments via the log server of environment 1641, to which environments 1602, 1603, 1604 may have one-way write privileges. Accordingly, if in the process of reviewing the logs of environment 1641, the master controller 1601 discovers that environment 1602 may have been compromised by malware, the master controller 1601 may use SDN tools to isolate the environment 1602 so that only an out-of-band connection 260 (e.g., see Fig. 16C ). In addition, the controller 1601 may send a notification about the possible problem to the administrator of the environment 1602. The controller may also isolate the damaged environment 1602 by selectively disabling any connection (e.g., in-band management connection 270) between the damaged environment and any of the other environments 1603, 1604. In another example, the master controller 1601 may discover through logs that resources within the environment 1603 are running too hot. This may allow the master controller to intervene in an application or service of the environment 1603 and migrate the application or service to a different environment (whether the different environment is a pre-existing environment or a newly created environment).

[0315] The controller 1601 can also set up one or more similar systems according to the purchaser or user's request. Fig.16E As shown, a purchase application 1650 may be provided, for example, in a console or other location, which allows a purchaser to purchase a cloud, host, system environment, or application or request that the cloud, host, system environment, or application be set up for the purchaser. The purchase application 1650 may instruct the controller 1601 to set up the environment 1602. The environment 1602 may include a controller 1601a, which will deploy or build an IT system, for example, by allocating or assigning resources to the environment 1602.

[0316] Fig.16FUser interfaces 1632, 1633, 1634 are shown that can be used when environments 1602, 1603, 1604 each operate as a cloud and may or may not include a controller. User interfaces 1632, 1633, 1634 (the user interfaces correspond to environments 1602, 1603, 1604, respectively) can each be connected through a main controller 1601, which manages the connection of the user interfaces to the environments. Alternatively or in addition, interface 1640a (the interface can take the form of a console) can be directly coupled to environment 1602, interface 1640b (the interface can take the form of a console) can be directly coupled to environment 1603, and interface 1640c (the interface can take the form of a console) can be directly coupled to environment 1604. Whether or not the connection to the main controller 1601 is separated, disconnected, or disabled, a user can use one or more of the interfaces to use an environment or cloud.

[0317] Clone and back up systems for change management support:

[0318] Some of the environments 1602, 1603, 1604 may be clones of typical setup software used by developers. The environments may also be clones of current working environments as a measure of scale; for example, cloning an environment in another data center in a different location to reduce latency due to location.

[0319] Therefore, it should be understood that the master controller setting up systems and resources in separate environments or subsystems may allow for cloning or backing up portions of the IT system. This may be used in testing and change management as described herein. Such changes may include, but are not limited to: changes to code, configuration rules, security patches, templates, and / or other changes.

[0320] According to an example embodiment, an IT system or controller as described herein may be configured to clone one or more environments. The new or cloned environment may or may not include the same resources as the original environment. For example, it may be desirable or necessary to use a completely different combination of physical and / or virtual resources in a new or newly cloned environment. It may be desirable to clone an environment to a different location or time zone that can manage the optimization of usage conditions. It may be desirable to clone an environment to a virtual environment. In the process of cloning an environment, the global system rules 210 and global templates 230 of a controller or master controller may include information about how to configure and / or run various types of hardware. The configuration rules within the system rules 210 may indicate the arrangement and use of resources so that resources and applications are more optimal given the specific available resources.

[0321] The master controller structure provides its ability to set up systems and resources in separate environments or subsystems, provide structures for cloning environments, provide structures for creating development environments, and / or provide structures for deploying a set of standardized applications and / or resources. Such applications or resources may include, for example, but are not limited to those that can be used for the following: developing and / or running applications, or backing up parts of an IT system or restoring from backups of the IT system and other disaster recovery applications (e.g., a LAMP (apache, mysql, php) stack, a system containing servers running a web front end and react / redux and resources running node.js, and a mongo database and other standardized "stacks"). Sometimes, the master controller may deploy an environment that is a clone of another environment, and the master controller may obtain configuration rules from a subset of the configuration rules used to create the original environment.

[0322] According to an example embodiment, change management of a system or a subset of a system may be accomplished by cloning one or more environments or configuration rules or a subset of configuration rules of such environments. Changes may be required to make changes such as code, configuration rules, security patches, templates, hardware changes, adding / removing components and dependent applications, and other changes.

[0323] According to example embodiments, such changes to the system may be automated to avoid errors caused by directly manually entering changes. Changes may be tested by the user in a development environment before they are automatically implemented on the active system. According to example embodiments, an active production environment may be cloned by using a controller to automatically drive, supply and / or configure an environment configured using the same configuration rules as the production environment. The cloned environment may be run and operated (while a backup environment may preferably be kept as a contingency in case changes need to be undone). This may be done using a controller as described above with reference to Figures 1 to 16F The description is done by creating, configuring, and / or provisioning a new system or environment using system rules 210, templates 230, and / or system states 220. The new environment can be used as a development environment to test changes that will later be implemented in a production environment. The controller can generate the infrastructure of such an environment from a software-defined structure into the development environment.

[0324] A production environment as defined herein means an environment used to operate the system, rather than an environment used only for development and testing, ie, a development environment.

[0325] When the production environment is cloned, the infrastructure or cloned development environment is configured by the controller according to the global system rules 210 and generated as a production environment. Changes to the development environment may be made to code, templates 230 (changing existing templates or changes related to the creation of new templates), security and / or applications or infrastructure configurations. When new changes implemented in the development environment are ready through development and / or testing as needed, the system automatically changes the development environment and then runs or deploys it as a production environment. The new system rules 210 are then uploaded to the controller and / or main controller of the environment, which will apply the system rule changes to the specific environment. The system state 220 is updated in the controller, and additional or revised templates 230 can be implemented. Therefore, the complete system knowledge of the infrastructure can be maintained by the development environment and / or the main controller, and the complete system knowledge of the infrastructure can be recreated. The complete system knowledge as used herein may include, but is not limited to, system knowledge of resource status, resource availability, and system configuration. Complete system knowledge may be gathered by the controller from system rules 210, system status 220, and / or by querying resources using one or more in-band management connections 270, one or more out-of-band management connections 260, and / or one or more SAN connections 280. In particular, resources may be queried to determine resource, network or application utilization, configuration status, or availability.

[0326] The cloned infrastructure or environment may be software defined via system rules 210; however, this need not be the case. The cloned infrastructure or environment may typically include or not include a front end or user interface, and one or more allocated resources, which may or may not include computing resources, networking resources, storage resources, and / or application networking resources. The environment may or may not be arranged as a front end, middleware, and database. A service or development environment may be started using the system rules 210 of the production environment. In particular, for the purpose of cloning, the infrastructure or environment allocated for use by the controller may be software defined. Therefore, the environment may be able to be deployed via system rules 210 and can be cloned by similar means. Before or when a change is needed, the clone or development environment may be automatically set up by a local controller or a master controller using system rules 210.

[0327] Before the development environment is isolated from the production environment, data of the production environment may be written to a read-only data store so that the data will be used by the development environment during development and testing.

[0328] While the production environment is online, the user or client can make changes to the development environment and test the changes. When testing development and changes in the development environment, the data in the data store can be changed. For volatile or writable systems, hot synchronization of the data with the data of the production environment can also be used after the development environment is set up or deployed. The required changes to the system, application and / or environment can be made and tested in the development environment. The required changes are then made to the script of the system rules 210 to create a new version for the environment or for the entire system and the main controller.

[0329] According to another example embodiment, the newly developed environment can then be automatically implemented as a new production environment while maintaining the previous production environment or making it fully functional, so that recovery of the production environment in an earlier state is possible without losing a large amount of data. The development environment is then started using the new configuration rules within the system rules 210, and the database is synchronized with the production database and switched to a writeable database. The original production database can then be switched to a read-only database. The previous production environment is preserved intact as a copy of the previous production environment for a desired period of time in case recovery back to the previous production environment is needed.

[0330] The environment may be configured as a single server or instance that may include or contain physical and / or virtual hosts, networks, and other resources. In another example embodiment, the environment may be a plurality of servers that contain physical and / or virtual hosts, networks, and other resources. For example, there may be a plurality of servers that form a load-balanced Internet-oriented application; and the servers may be connected to a plurality of API / middleware applications (the API / middleware applications may be hosted on one or more servers). The database of the environment may include one or more databases to which the API communicates queries in the environment. The environment may be constructed from system rules 210 in a static or volatile form. An environment or instance may be virtual or physical, or a combination of each.

[0331] Configuration rules for an application or a system within system rules 210 may specify various compute backends (e.g., bare metal, AMD epyc servers, Intel Haswell on qemu / kvm), and may include rules on how to run an application or service on a new compute backend. Thus, if, for example, there is a situation where the availability of resources for testing is reduced, an application may be virtualized.

[0332] Using and following the examples described herein, the test environment can be deployed on virtual resources where the original environment uses physical resources. Figures 1 to 18BThe controller described, and as further described herein, a system or environment may be cloned from a physical environment to an environment that may or may not include virtual resources in whole or in part.

[0333] Fig.17A An example embodiment of the system 100 is shown, which includes a controller 1701 and one or more environments, such as 1702, 1703, 1704. The system 100 can be a static system, i.e., a system in which active user data does not constantly change the state of the system or frequently manipulate data, such as a system that only hosts static web pages. The system can be coupled to a user (or application) interface 110.

[0334] Controller 1701 may be configured in a similar manner to controllers 200 / 1401 / 1501 / 1601 described herein, and may similarly include global system rules 210, controller logic 205, templates 230, and system state elements 220. Controller 1701 may be configured as described herein with reference to FIG. 14A to FIG. 16F The controller 1701 may be coupled to one or more other controllers or environments in the manner described. The global rules 210 of the controller 1701 may include rules that can manage and control other controllers and / or environments. Such global rules 210, controller logic 205, system state 220, and templates 230 may be used to communicate with the controller 1701 in accordance with the present disclosure. Figures 1 to 16F Systems or environments are set up, provisioned, and deployed in a similar manner as described.Each environment may be configured using a subset of the global system rules 210 that define the environment's operations, including defining the environment's operations with respect to other environments.

[0335] The global system rules 210 may also include change management rules 1711. The change management rules 1711 include a set of rules and / or instructions that may be used when changes to the system 100, the global system rules 210, and / or the controller logic 205 may be needed. The change management rules 1711 may be configured to allow a user or developer to develop changes, test the changes in a test environment, and then implement the changes by automatically converting the changes to a new set of configuration rules within the system rules 210. The change management rules 1711 may be a subset of the global system rules 210 (e.g., Fig.17A 1701 ), or the change management rules may be separate from the global system rules 210. The change management rules may use a subset of the global system rules 210. For example, the global system rules 210 may include a subset of environment creation rules configured to create a new environment. The change management rules 1711 may be configured to set up and use the system or environment configured and set up by the controller 1701 to copy and clone some or all aspects of the system 100. The change management rules 1711 may be configured to permit testing of new changes proposed to the system before implementation by testing and implementation using a clone of the system.

[0336] like Fig.17A The clone 1705 shown may include rules, logic, applications, and / or resources of a particular environment or part of the system 100. The clone 1705 may include hardware similar to or different from the system 100, and may or may not use virtual resources. The clone 1705 may be set up as an application. The clone 1705 may be set up and configured using configuration rules within the system rules 210 of the system 100 or the controller 1701. The clone 1705 may or may not include a controller. The clone 1705 may include allocated networking resources, computing resources, application networks, and / or data storage resources as described in more detail above. Such resources may be allocated under the control of the controller 1701 using change management rules 1711. The clone 1705 may be coupled to a user interface that allows a user to make changes to the clone 1705. The user interface may be the same or different from the user interface 110 of the system 100. The clone 1705 may be used for the entire system 100, or a part of the system 100, such as one or more environments and / or controllers. The clone 1705 may or may not be a complete copy of the system 100. The clone 1705 can be coupled to the system 100 via an in-band management connection 270, an out-of-band management connection 260, and / or a SAN connection 280, which can be selectively enabled and / or disabled entirely, and / or converted to a one-way read and / or write connection. Thus, the connection to data in the clone environment 1705 can be changed to make the clone data read-only when the clone environment 1705 is isolated from the production environment during testing, or before the clone environment 1705 is ready to run as a new production environment. For example, if the clone 1705 has a data connection to the environment 1702, this data connection is made read-only for isolation purposes.

[0337] Optional backup 1706 may or may not be used for the entire system, or a portion of the system, such as one or more environments and / or controllers. Backup 1706 may include networking resources, computing resources, application networks, and / or data storage resources as described in more detail above. Backup 1706 may or may not include a controller. Backup 1706 may be a complete copy of system 100. Backup 1706 may be set up as an application or using hardware similar to or different from system 100. Backup 1706 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or completely disabled, and / or converted to a unidirectional read and / or write connection.

[0338] Fig. 17B Shown for use Fig.17A1785, a user or management application initiates a change to the system. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, hardware changes, adding / removing components and / or dependent applications, and other changes. At step 1786, controller 1701 initiates a change to the system. FIG. 14A to FIG. 16F The environment is set up in the manner described to become a clone environment 1705 (where the clone environment may have its own new controller, or the clone environment may use the same controller as the original environment).

[0339] At step 1787, the controller 1701 may use global rules 210 including change management rules 1711 to clone all or part of one or more environments (e.g., a "production environment") of the system to a clone environment 1705 (e.g., where the clone environment 1705 may be used as a "development environment"). Thus, the controller 1701 identifies and allocates resources, uses system rules 210 to set up and allocate clone resources, and copies any of the following from the environment to the clone environment: data, configuration, code, executable files, and other information required to drive the application. At step 1788, the controller 1701 optionally backs up the system by using configuration rules within the system rules 210 to set up another environment to serve as a backup 1706 (with or without a controller), and copies the template 230, controller logic 205, and global rules 210.

[0340] After clone 1705 is generated from the production environment, clone 1705 can be used as a development environment, where changes can be made to the following: cloned code, configuration rules, security patches, templates and other changes. At step 1789, changes to the development environment can be tested before implementation. During testing, clone 1706 can be isolated from the production environment (system 100) or other components of the system. This can be achieved by causing controller 1701 to selectively disable one or more of the connections between system 100 and clone 1706 (e.g., by disabling in-band management connection 270 and / or disabling application network connection). At step 1790, a determination is made as to whether the changed development environment is ready). If step 1709 determines that the development environment is not yet ready (this is a decision usually made by the developer), the process flow returns to step 1789 to make further changes to the clone environment 1705. If step 1790 determines that the development environment is ready, the development environment and the production environment can be switched at step 1791. That is, the controller transitions the development environment 1705 to the new production environment, and the previous production environment may remain unchanged until the transition to the development environment / new production environment is complete and satisfactory.

[0341] Fig.18A Another example embodiment of a system 100 that can be arranged and used in change management of a system is shown. Fig.18A In the example of , the system 100 includes a controller 1801 and one or more environments 1802, 1803, 1804, 1805. The system is shown with a clone environment 1807 and a backup system 1808.

[0342] Controller 1801 is configured in a similar manner to controllers 200 / 1401 / 1501 / 1601 / 1701 described herein and may include global system rules 210, controller logic 205, templates 230, and system state 220 elements. Controller 1801 may be described herein with reference to FIG. 14A to FIG. 16F The controller 1801 may be coupled to one or more other controllers or environments in the manner described. The global rules 210 of the controller 1801 may include rules that can manage and control other controllers and / or environments. Such global rules 210, controller logic 205, system state 220, and templates 230 may be used to communicate with the controller 1801 in accordance with the present disclosure. Figures 1 to 17B Systems or environments are set up, provisioned, and deployed in a similar manner as described.Each environment may be configured using a subset of the global rules 210 that define the environment's operations, including defining the environment's operations with respect to other environments.

[0343] The global rules 210 may also include change management rules 1811. The change management rules 1811 may include a set of rules and / or instructions that may be used when changes to the system, global rules, and / or logic may be needed. The change management rules may be configured to allow users or developers to develop changes, test the changes in a test environment, and then implement the changes by automatically converting the changes to a new set of configuration rules within the system rules 210. The change management rules 1711 may be a subset of the global system rules 210 (e.g., Fig.18A 1801 ), or the change management rules may be separate from the global system rules 210. The change management rules 1711 may use a subset of the global system rules 210. For example, the global system rules 210 may include a subset of environment creation rules configured to create a new environment. The change management rules 1811 may be configured to set up and use the system or environment set up and deployed by the controller 1801 to copy and clone some or all aspects of the system 100. The change management rules 1811 may be configured to permit testing of new changes proposed to the system before implementation by testing and implementing using a clone of the system.

[0344] like Fig.18AThe illustrated cloning environment 1807 may include: a controller 1807a having rules, controller logic, templates, system state data; and allocated resources 1820, which may be allocated to one or more environments and set according to the global system rules 210 and change management rules 1811 of the controller 1801. The backup system 1808 also includes: a controller 1808a having rules, controller logic, templates, system state data; and allocated resources 1821, which may be allocated to one or more environments and set according to the global system rules 210 and change management rules 1811 of the controller 1801. The system may be coupled to the user (or application) interface 110 or another user interface.

[0345] The clone environment 1807 may include rules, logic, templates, system states, applications, and / or resources of a particular environment or part of a system. The clone 1807 may include hardware similar to or different from the system 100, and the clone 1807 may use or not use virtual resources. The clone 1807 may be set as an application. The clone 1807 may be set and configured using configuration rules within the system rules 210 of the controller 1801 of the system 100 or the environment. The clone 1807 may include or not include a controller, and the clone may share a controller with the production environment. The clone 1807 may include allocated networking resources, computing resources, application networks, and / or data storage resources as described in more detail above. Such resources may be allocated under the control of the controller 1801 using change management rules 1811. The clone 1807 may be coupled to a user interface that allows a user to make changes to the clone 1807. The user interface may be the same as or different from the user interface 110 of the system 100.

[0346] The clone 1807 may be used for the entire system, or a portion of the system, such as one or more environments and / or controllers. In an example embodiment, the clone 1807 may include a hot standby data resource 1820a coupled to the data resource 1820 of the environment 1802. The hot standby data resource 1820a may be used when setting up the clone 1807 and in testing changes. For example, as described herein with respect to Fig.18BAs described, the hot standby data resource 1820a can be selectively disconnected or isolated from the storage resource 1820 during change management. The clone 1807 may or may not be a complete copy of the system 100. The clone 1807 may be coupled to the system 100 via the in-band management connection 270, the out-of-band management connection 260, and / or the SAN connection 280, which can be selectively enabled and / or completely disabled, and / or converted to a one-way read and / or write connection. Thus, the connection to the volatile data in the clone environment 1807 can be changed to make the clone data read-only when the clone environment 1807 is isolated from the production environment during testing, or before the clone environment is ready to operate as a new production environment.

[0347] When switching the old production environment to a new production environment, the controller 1801 may instruct the front end, load balancer, or other applications or resources to point to the new production environment. Therefore, when changes occur, users, applications, resources, and / or other connections may be redirected. This may be accomplished, for example, by a variety of methods, including but not limited to: changing ip / ipoib address lists, infinite bandwidth GUIDs, dns servers, infinite bandwidth partitions / opensm configurations; or changing software-defined networking (SDN) configurations, which may be accomplished by sending instructions to networked resources. The front end, load balancer, or other applications and / or resources may point to systems, environments, and / or other applications, including but not limited to: databases, middleware, and / or other back ends. Therefore, a load balancer may be used in change management to switch from an old production environment to a new environment.

[0348] Clone 1807 and backup 1808 can be set up and used in various aspects of managing system changes. Such changes may include, but are not limited to, changes to the following: code, configuration rules, security patches, templates, hardware changes, adding / removing components and / or dependent applications and other changes. Backup 1808 can be used for the entire system, or a part of the system, such as one or more environments and / or controller 1801. Backup 1808 may include networking resources, computing resources, application networks and / or data storage resources as described in more detail above. Backup 1808 may or may not include a controller. Backup 1808 can be a complete copy of system 100. Backup 1808 may include data required to rebuild the system / environment / application from the configuration rules included in the backup, and may include the used application data. Backup 1808 can be set up as an application or using hardware similar or different from system 100. Backup 1808 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or disabled, and / or converted to unidirectional read and / or write connections.

[0349] Fig.18B It is shown that especially Fig.18A Used when the system includes volatile data or when the database is writable Fig.18A 18. The example process flow for performing change management on a system of FIG. 18. Such a database may be part of the storage resources used by the environments in the system. At step 1870, the system (including the production environment) is deployed using global system rules.

[0350] At step 1871, the production environment is cloned to create a read-only environment using the global system rules 210 including the change management rules 1811 and the resource allocation by the master controller 1801 or controller in the clone environment, wherein the clone environment is prohibited from writing to the system. The clone environment can then be used as a development environment.

[0351] At step 1872, the hot spare 1820a is activated and assigned to the clone environment 1807 to store any volatile data that has changed in the system 100. The clone data is updated so that the new version of the development environment can be tested with the updated data. The hot sync data can be turned off at any time. For example, the hot sync data can be turned off when testing a write from an old environment or a production environment to a development environment.

[0352] At step 1873, the user may then use the cloned environment 1807 as the development environment to process the changes. The changes to the development environment are then tested at step 1874. At step 1875, a determination is made as to whether the changed development environment is ready (typically such a determination is made by the developer). If step 1875 determines that the changes are not ready, the process flow may return to step 1873, where the user may back out and make other changes to the development environment. If step 1875 determines that the changes are ready to run, the process flow proceeds to step 1876, where configuration rules are updated in the system or controller for the particular environment and the new updated environment will be deployed using the configuration rules.

[0353] At step 1877, the development environment (or new environment) may then be redeployed with the changes to the desired final configuration with the desired resources and hardware allocations prior to operation. In the next step at 1878, the write capability of the original production environment is disabled, and the original production environment becomes read-only. Although the original production environment is read-only, as part of 1878, any new data from the original production environment (or possibly the new production environment) may be cached and identified as transitional data. As an example, the data may be cached in a database server or other suitable location (e.g., a shared environment). The development environment (or new environment) and the old production environment are then switched at step 1879 so that the development environment (or new environment) becomes the production environment.

[0354] After this switch, the new production environment is made writable at step 1880. If the new production environment is considered working at step 1881, as determined by the developer, any data lost during the switch process (where such data had been cached at step 1878) can be reconciled with the data written to the new environment at step 1884. After this reconciliation, the change is completed (step 1885).

[0355] If step 1881 determines that the new production environment is not working (e.g., a problem is identified that requires the system to be restored to the old system), then the environment is switched back so that the old production environment becomes the production environment again at step 1882. As part of step 1882, the configuration rules for the subject environment on controller 1801 are restored back to the previous version that was used for the now restored production environment.

[0356] At step 1883, changes to the database may be determined, for example, using cached data; and the data may be restored to the old production environment using the old configuration rules. To support step 1883, the database may maintain a log of changes made thereto, thereby permitting step 1883 to determine changes that may need to be undone. A backup database may be used to cache data as described above, wherein the cached data is tracked and clocked; the clock may be restored to determine which changes were made. Snapshots and logs may be used for this purpose.

[0357] After restoring the cached data at 1883, the process may return to step 1871 if restarting is desired.

[0358] The example change management system discussed herein may be used, for example, when upgrading, adding or removing hardware or software, when patching software, when detecting a system failure, when migrating a host during a hardware failure or detection, for dynamic resource migration, for changing configuration rules or templates, and / or for making any other system-related changes. The controller 1801 or system 100 may be configured to detect a failure, and when a failure is detected, change management rules or existing configuration rules may be automatically implemented on other hardware that may be used for the system or the controller. Examples of available fault detection methods include, but are not limited to, checking the host, querying the application, and running various tests or test suites. The change management configuration rules described herein may be implemented when a failure is detected. When a failure is detected, such rules may trigger the automatic generation of a backup environment implemented by the controller, the automatic migration of data or resources. The selection of backup resources may be based on resource parameters. Such resource parameters may include, but are not limited to, usage information, speed, configuration rules, and data capacity and usage status.

[0359] As described herein, any time a change occurs, the controller will create a log of the change and what was actually performed. For security or system update considerations, the controller described herein can be configured to automatically turn on and off according to configuration rules and update the IT system status. The controller may shut down resources to save power. The controller may turn on or migrate resources at different times for different efficiencies. During migration, configuration rules are followed and the environment or system can be backed up or copied. If there is a security vulnerability, the controller can isolate and shut down the attacked area.

[0360] Although the present invention has been described above with respect to exemplary embodiments of the present invention, various modifications can be made to the present invention that still fall within the scope of the invention. Such modifications to the present invention will be recognized after reviewing the teachings herein.

[0361] Appendix A: Example storage connection procedure

[0362] This describes an example process and example rules associated with sharing storage resources between multiple systems. It should be understood that this is merely an example of a storage connection process, and other techniques for connecting computing resources to storage resources may be used. Unless otherwise specified, these rules apply to all systems attempting to initiate a storage connection.

[0363] Definitions for this Appendix A:

[0364] Storage resources: Blocks, files, or file systems that can be shared via storage transports.

[0365] Storage Transport: The method of sharing storage resources locally or remotely. Examples would be iSCSI / iSER, NVMeoF, NFS, Samba file sharing.

[0366] System: Anything that can attempt to connect to a storage resource over a specified storage transport. A system can support any number of storage transports and can make its own decisions about which transports to use.

[0367] Read-only: A read-only storage resource does not allow modification of the data it contains. This constraint is enforced by the storage daemon that handles operations that export storage resources on the storage transport. For additional insurance, some data stores can set the storage resource backend data to read-only (for example, set LVM LVs to read-only).

[0368] Read-Write (or Volatile): A read-write (volatile) storage resource is a storage resource whose contents can be modified by a system connected to the storage resource.

[0369] Rules: There is a set of rules that the controller must follow when determining whether a system can connect to a given storage resource.

[0370] 1. Read-write storage resources should be exported only over a single storage transfer.

[0371] 2. Read-write storage resources should only be connected to / connected by a single system.

[0372] 3. Read-write storage resources should not be connected as read-only.

[0373] 4. Read-only storage resources may be exported over multiple storage transports.

[0374] 5. Read-only storage resources may be connected to / from multiple systems.

[0375] 6. Read-only storage resources should not be connected as read-write.

[0376] process

[0377] If we consider the connection process as a function, then the function will take 2 arguments:

[0378] 1. Storage resource ID

[0379] 2. List of supported storage transports (prioritized in order)

[0380] First, we determine whether the requested storage resource is read-only or read-write.

[0381] If it is read-write, then we need to check to see if the storage resource is already connected, since we are limiting read-write storage resources to a single connection. If the storage resource does already have a connection, then we should ensure that the system requesting the storage resource is the currently connected system (this might happen in the case of a reconnect, for example). Otherwise, we will get an error, since multiple systems cannot connect to the same read-write storage resource. If the requesting system is

[0382] If a system is connected to this storage resource, we should ensure that one of the available storage transports matches the current export of this storage resource. If it does, we pass the connection information to the requesting system. If not, we report an error because we cannot service a read-write storage resource over multiple storage transports.

[0383] For read-only and unconnected read-write storage resources, we iterate over the list of supplied storage transports and attempt to export the storage resource using the transport. If the export fails, we continue to complete the list until we succeed or run out of storage transports. If we run out, we notify the requesting system that we could not connect to the storage resource. On successful export, we store the connection information in the database, along with the new (resource, transport) => (system) relationship. The requesting system then passes the storage transport connection information.

[0384] Systems: Storage connectivity is currently performed by the controller and compute daemon during normal operation. However, future iterations may enable services to connect directly to storage resources and bypass the compute daemon. This may be a requirement for physical deployments of the example services, and it may be useful to use the same process for virtual machine deployments as well.

[0385] Appendix B: Example Connection to OverlayFS

[0386] Services use OverlayFS to reuse common file system objects and reduce service bundle size.

[0387] A service in this instance consists of 3 or more storage resources:

[0388] 1. Platform. This contains the base Linux file system and is read-only accessible.

[0389] 2. Services. This contains all software directly related to the operation of the service (NetThunder service daemon, OpenRC scripts, binaries, etc.). This storage resource accepts read-only access.

[0390] 3. Volatility. These storage resources contain all changes to the system and are managed by LVM from within the service (for physical, container, and virtual machine deployments).

[0391] When running in a virtual machine, the service is directly kernel-booted in Qemu using a custom Linux kernel with initramfs that contains logic to:

[0392] 1. Assemble LVM volume groups (VG) from available read-write disks

[0393] *This VG contains one logical volume (LV) which contains all volatile storage data for the service.

[0394] 2. Mount platforms, services, and LVs

[0395] 3. Use a union file system (OverlayFS in our case) to combine the three file systems.

[0396] The same process can be used for physical deployments. One option is to remotely provision the kernel to a lightweight OS that is booted via PXE boot or IPMIISO boot, and then enter the new real kernel via kexec. Or skip the lightweight OS and enter our kernel directly via PXE boot. Such a system may require additional logic in the kernel initramfs to connect to storage resources.

[0397] An OverlayFS configuration might look like this:

[0398]

[0399] Due to some limitations of OverlayFS, we allow marking a special directory ' / data' as "out of tree". This directory is available to services if the service creates a ' / data' directory when the service package is created. This special directory is mounted via 'mount --rbind' to allow access to a subset of the volatile layer that is not inside OverlayFS. This is required for applications such as NFS (Network File System) that do not support sharing directories that are part of OverlayFS.

[0400] Kernel file system layout:

[0401]

[0402] We create the / new_root directory and use that directory as the target to configure our OverlayFS. Once OverlayFS has been configured, we exec_root into / new_directory and the system boots up normally with all available resources.

Claims

1. A computer system comprising: Controller; Multiple system rules; Multiple templates; System status; and Physical computing resources; wherein the controller is configured to provide automated management of an infrastructure of the computer system based on the system rules, the templates, and the system status; wherein the controller is configured to (1) add a storage resource to the computer system, and (2) update the system state using (i) a location of the storage resource in the computer system and (ii) a transmission type of the storage resource; wherein the controller or the physical computing resource is configured to query the system status to obtain the location of the storage resource and the transmission type of the storage resource; The controller is configured to couple the storage resource with the physical computing resource based on a response to the query. 2 . The system of claim 1 , wherein the controller is configured to instruct the physical computing resource to boot from the storage resource.

3. The system of claim 2, further comprising an out-of-band management connection, and wherein the physical computing resource is configured to boot from the storage resource via the out-of-band management connection.

4. The system of claim 2, further comprising a storage area network (SAN), and wherein the physical computing resource is configured to boot from the storage resource through the SAN.

5. The system of claim 2, further comprising an in-band management connection or a network connection, and wherein the physical computing resource is configured to boot from the storage resource through the in-band management connection or the network connection. 6 . The system of claim 1 , wherein the controller is configured to process one of the templates to obtain an image to start the storage resource, drive the storage resource, or enable the storage resource.

7. The system of claim 1, wherein the controller is configured to add the physical computing resource to the computer system in a manner that sends a mirror to the physical computing resource.

8. The system of claim 1, further comprising a storage area network (SAN), and wherein the controller is configured to couple the physical computing resources and the storage resources to the SAN; and The controller is configured to use the SAN to start the physical computing resource.

9. The system of claim 1, wherein the controller is configured to (1) add the physical computing resource to the computer system, (2) instruct the added physical computing resource to start, and (3) track the added and started physical computing resource in a system state.

10. The system of claim 1, wherein the transport type is small computer system interface (iscsi) or non-volatile memory express over network architecture (nvmeof).