Automatically deployed information technology (IT) system and method with enhanced security
Through controller technology, IT systems are automatically managed based on global system rules and templates, which solves the problems of IT system setting complexity and security risks in existing technologies, achieves flexibility, security and interoperability, simplifies configuration and management processes, and improves system reliability and compliance.
Patent Information
- Application Number
- CN202510886522.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-11
- Filing Date
- 2020-06-10
- Publication Date
- 2025-09-30
AI Technical Summary
Existing IT systems have complexity and security risks during setup, configuration, and deployment. Performance and security issues are difficult to diagnose and resolve, and the change management process can easily lead to system failures and security vulnerabilities. Flexibility and interoperability are difficult to achieve, and insufficient documentation leads to difficulties in compliance and troubleshooting.
It uses controller technology to automatically manage computer systems based on global system rules, templates and IT system status, including self-assembly rules and operation rules, to achieve resource selection, installation, interconnection and update, provide security and interoperability, automatically record configuration and status, support backup and recovery, dynamic resource migration and load balancing.
It improves the flexibility and security of IT systems, reduces human errors, realizes automated configuration and deployment, supports rapid fault recovery and resource optimization, ensures system interoperability and compliance, and reduces configuration and management complexity.
Smart Images

Figure CN120729583A_ABST
Abstract
Description
[0001] This application is a divisional application of patent application 202080056751.9, filed on June 10, 2020, and entitled “Automatically deployed information technology (IT) system and method with enhanced security.”
[0002] Cross-references and priority claims to related patent applications This patent application claims priority to U.S. Provisional Patent Application No. 62 / 860,148, filed on June 11, 2019, and entitled “Automatically Deployed Information Technology (IT) System and Method with Enhanced Security,” the entire disclosure of which is incorporated herein by reference. Background Art
[0003] In recent decades, the demand, use, and need for computing have skyrocketed. The resulting need for greater storage, speed, computing power, applications, and accessibility has led to a rapidly evolving computing landscape, providing tools for businesses of all types and sizes. Consequently, public virtual computing and cloud computing systems have been developed to provide improved computing resources to large numbers of users and diverse types of users. This exponential growth is likely to continue. At the same time, greater risks of failure and security have made infrastructure setup, management, change management, and updates more complex and expensive. Scalability, or the ability to evolve systems over time, has also become a major challenge in the information technology field.
[0004] Problems in most IT systems can be difficult to diagnose and resolve, many of which involve performance and security. Constraints on the time and resources available to set up, configure, and deploy systems can lead to errors and future IT problems. Over time, many different administrators may be involved in changing, patching, or updating IT systems, including users, applications, services, security, software, and hardware. Often, documentation and history of configurations and changes may be insufficient or lost, making it difficult to understand how a particular system has been configured and is functioning at a later time. This can complicate future changes or troubleshooting. When problems or failures occur, it can be difficult to restore and reproduce IT configurations and settings. Furthermore, system administrators can easily make mistakes, such as incorrect commands or other errors, which can paralyze computers, web databases, and services. Furthermore, while the increased risk of security breaches is commonplace, changes, updates, and patches implemented to mitigate them can result in undesirable downtime.
[0005] Once critical infrastructure is in place, working, and active, the costs or risks often appear to outweigh the benefits of changing the system. Problems involved in making changes to active IT systems or environments can cause extensive and sometimes catastrophic problems for the users or entities that rely on those systems. At the very least, the amount of time spent troubleshooting and fixing failures or problems that arise during change management can require significant resources in time, personnel, and money. Technical problems that potentially arise when making changes to an active environment can have cascading effects and may not be resolved simply by undoing the changes that were made. Many of these problems can make it impossible to quickly rebuild the system if failures exist during change management.
[0006] Additionally, bare metal cloud nodes or resources within an IT system may be vulnerable to security issues, compromised, or accessed by rogue users. Hackers, attackers, or rogue users may be able to divert from, access, or invade any other part of the IT system or the network associated with the node. Bare metal cloud nodes or controllers within an IT system may also be vulnerable to attacks due to resources connected to application networks, which may expose the system to security threats or otherwise compromise the system. According to various exemplary embodiments disclosed herein, an IT system may be configured to improve the security of interfacing with the Internet or application networks, regardless of whether the bare metal cloud nodes or resources are connected to external networks. Summary of the Invention
[0007] According to an exemplary embodiment, an IT system includes bare metal cloud nodes or physical resources. When powering on, configuring, managing, or using bare metal cloud nodes or physical resources, if they may be connected to a network where the nodes may be used by other people or customers, in-band management may be omitted, switched from a controller, disconnected from the controller, or filtered from the controller. Additionally, applications or application networks within the system may be disconnected from the controller, disconnectable, switchable, or filtered from the controller via one or more resources coupled to the controller via the application network.
[0008] Physical resources, including virtual machines or hypervisors, can also be vulnerable to security issues, compromised, or accessed by rogue users, where the hypervisor can be used to transfer to another hypervisor as a shared resource. An attacker could potentially break out of the virtual machine and gain network access to the management and / or administration system through the controller. According to various exemplary embodiments disclosed herein, an IT system can be configured to enhance security by disconnecting one or more physical resources, including virtual resources, from the controller via an in-band management connection, enabling disconnection from the controller, filtering from the controller, or not connecting to the controller.
[0009] According to an exemplary embodiment, the physical resources of the IT system may include one or more virtual machines or a hypervisor system, wherein an in-band management connection between a controller and the physical resources may be omitted, disconnected from the resources, capable of being disconnected from the resources, or filtered / capable of being filtered from the resources.
[0010] According to an exemplary embodiment, a system may include a controller that provides and manages related services within the system using the techniques described herein. As an example, cleanup rules may be created and maintained to manage how modifications are unwound when a service that has dependencies with other services is deleted.
[0011] According to an exemplary embodiment, a system may include a controller that provides storage and / or provides resources to computing resources and connects them to a cloud instance using the techniques described herein.
[0012] Furthermore, according to exemplary embodiments, a system may utilize the architecture described herein to support efficient backup operations, including backups involving multiple interdependent services. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic diagram of a system according to an exemplary embodiment.
[0014] Figure 2A is used for Figure 1 Schematic diagram of an example controller for a system.
[0015] Figure 2B An example flow illustrating the operation of an example set of storage expansion rules.
[0016] Figure 2C and Figure 2D Diagram for execution Figure 2B An alternative example of steps 210.1 and 210.2 in FIG.
[0017] Figure 2E An example template is shown.
[0018] Figure 2F An example process flow of controller logic with respect to processing templates is shown.
[0019] Figure 2G and Figure 2H Shown for Figure 2F Example process flow for steps 205.11, 205.12, and 205.13.
[0020] Figure 2IAnother example template is shown.
[0021] Figure 2J Illustrate another example process flow of controller logic with respect to processing templates.
[0022] Figure 2K An example process flow for managing service dependencies is shown.
[0023] Figure 2L is a diagram of an example image derived from a template according to an exemplary embodiment.
[0024] Figure 2M A set of example system rules is shown.
[0025] Figure 2N Graphics are processed by the controller logic Figure 2M Example process flow for system rules.
[0026] Figure 2O Illustrate an example process flow for provisioning storage resources from file system blobs or other file groups.
[0027] Figure 3A yes Figure 2A Schematic diagram of a controller with added computing resources.
[0028] Figure 3B is a diagram of an example image derived from a template according to an exemplary embodiment.
[0029] Figure 3C An example process flow for adding resources, such as computing resources, storage resources, and / or networking resources, to a system is illustrated.
[0030] Figure 4A yes Figure 2A Figure 1 shows a diagram of a controller with storage resources added.
[0031] Figure 4B is a diagram of an example image derived from a template according to an exemplary embodiment.
[0032] Figure 5A yes Figure 2A A diagram of a controller with JBOD and storage resources added.
[0033] Figure 5B Illustrated is an example process flow for adding storage resources and direct attached storage for the storage resources to a system.
[0034] Figure 6A yes Figure 2A Figure 1. Schematic diagram of a controller with networking resources added.
[0035] Figure 6Bis a diagram of an example image derived from a template according to an exemplary embodiment.
[0036] Figure 7A is a schematic diagram of a system in an example physical deployment according to an exemplary embodiment.
[0037] Figure 7B Illustrate an example process for adding resources to an IT system.
[0038] Figure 7C and Figure 7D An example process flow for deploying an application on multiple computing resources, multiple servers, multiple virtual machines, and / or across multiple sites is shown.
[0039] Figure 8A is a schematic diagram of a system in an example deployment according to an exemplary embodiment.
[0040] Figure 8B An example process flow for scaling from a single-node system to a multi-node system is shown.
[0041] Figure 8C Illustrate an example process flow for migrating storage resources to new physical storage resources.
[0042] Figure 8D An example process flow is shown for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for compute and storage.
[0043] Figure 8E Another example process flow for scaling from a single node to multiple nodes in a system is shown.
[0044] Figure 9A is a schematic diagram of a system in an example physical deployment according to an exemplary embodiment.
[0045] Figure 9B is a diagram of an example image derived from a template according to an exemplary embodiment.
[0046] Figure 9C An example of installing an application from an NT package is shown.
[0047] Figure 9D is a schematic diagram of a system in an example deployment according to an exemplary embodiment.
[0048] Figure 9E An example process flow for adding a virtual computing resource host to an IT system is shown.
[0049] Figure 9FAn exemplary system is illustrated with additional connections to resources including instance 310a on the cloud.
[0050] Figures 9G-1 to 9G-4 About Figure 9F The process flow of the system.
[0051] Figure 9H Graphic Figure 9F An exemplary system with additional instances on the cloud is shown in , where the additional instances connect to the cloud API and connect to the controller through a VPN through an in-band management connection.
[0052] Figure 9I About Figure 9H An exemplary process flow for a system.
[0053] Figure 9J Graphic Figure 9F Another exemplary system shown in , which also includes an additional instance on the cloud, wherein the instance is connected to the cloud API and connected to the controller through a VPN through an in-band management connection.
[0054] Figure 9K About Figure 9J An exemplary process flow for a system.
[0055] Figure 9L Illustrate an exemplary process flow for expanding from a system to a local host.
[0056] Figure 10 is a schematic diagram of a system in an example deployment according to an exemplary embodiment.
[0057] Figure 11A An exemplary embodiment system and method is shown.
[0058] Figure 11B An exemplary embodiment system and method is shown.
[0059] Figure 12 An exemplary embodiment system and method is shown.
[0060] Figure 13A is a schematic diagram of a system according to an exemplary embodiment.
[0061] Figure 13B is another schematic diagram of a system according to an exemplary embodiment.
[0062] Figures 13C to 13E An example process flow is illustrated for a system according to one example embodiment.
[0063] Figure 14AAn example system is shown where a master controller has deployed controllers on different systems.
[0064] Figure 14B and Figure 14C An example flow chart showing possible steps for supplying a controller using a master controller is shown.
[0065] Figure 15A An example system in which a host controller generates an environment is shown.
[0066] Figure 15B Shows an example process flow for a controller to set up an environment.
[0067] Figure 15C Figure 1 shows an example process flow for a controller to set up multiple environments.
[0068] Figure 16A The illustrated controller operates as a master controller to configure an exemplary embodiment of one or more controllers.
[0069] 16B to 16D An example system is shown in which an environment can be configured to write to another environment.
[0070] Figure 16E An example system is shown in which a user can purchase new environments generated by a controller.
[0071] Figure 16F A graphical user interface is provided for interfacing with the example system in an environment generated by the controller.
[0072] 17A to 18B Illustrate an example of change management tasks for a new environment.
[0073] Figure 19A -G illustrates examples of systems and the process flows of these systems with respect to providing and managing related services within the systems.
[0074] Figure 20A -D illustrates an exemplary system and related process flows in which one or more computing resources host one or more services that utilize storage in one or more storage resources.
[0075] Figure 21A -L illustrates a system and examples of the associated process flows for backing up systems, services, or other components within the system.
[0076] Figures 22A-22C Illustrate a system and an example of the associated process flow for updating systems, services, or other components within the system. DETAILED DESCRIPTION
[0077] To provide technical solutions to the aforementioned needs in the art, the inventors disclose various inventive embodiments related to systems and methods for information technology that provide automated IT system setup, configuration, maintenance, testing, change management, and / or upgrades. For example, the inventors disclose a controller configured to automatically manage a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. As another example, the inventors disclose a controller configured to automatically manage the physical infrastructure of a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. Examples of automated management that can be performed by a controller may include: remotely or locally accessing and changing settings or other information on a computer that may be running an application or service; building an IT system; changing an IT system; building a separate stack in an IT system; creating a service or application; loading a service or application; configuring a service or application; migrating a service or application; changing a service or application; removing a service or application; cloning a stack to another stack on a different network; creating, adding, removing, setting up, configuring, reconfiguring and / or changing resources or system components; automatically adding, removing and / or restoring resources, services, applications, IT systems and / or IT stacks; configuring interactions between applications, services, stacks and / or other IT systems; and / or monitoring the health of IT system components. In an exemplary embodiment, the controller may be implemented as a physical or virtual computing resource that may be remote or local. Additional examples of controllers that may be employed include, but are not limited to, one or any combination of the following: processes, virtual machines, containers, remote computing resources, applications deployed by other controllers and / or services. The controller may be distributed across multiple nodes and / or resources and may be in other locations or networks.
[0078] IT infrastructure is most often made up of discrete hardware and software components. The hardware components used typically include servers, racks, power supplies, interconnects, display monitors, and other communication equipment. The methods and techniques for first selecting and then interconnecting these discrete components are highly complex, as a vast array of optional configurations vary with varying degrees of efficiency, cost-effectiveness, performance, and security. The hiring and training costs of individual technicians / engineers skilled in connecting these infrastructure components are very expensive. In addition, the vast number of possible iterations of hardware and software creates complexity in maintaining and updating hardware and software. This presents additional challenges when the individuals and / or engineering firms that initially installed the IT infrastructure are not comfortable performing updates. Software components such as operating systems are generally designed to work on a wide range of hardware, or are completely dedicated to specific components. In most cases, complex plans or blueprints are developed and executed. Changes, developments, expansions, and other challenges all require updating complex plans.
[0079] While some IT users purchase cloud computing services from providers that are growing in the industry, this does not solve the problems and challenges of setting up infrastructure, but rather shifts them from the IT users to the cloud providers. In addition, the large cloud providers have addressed the challenges and problems of setting up infrastructure in ways that may sacrifice flexibility, customization, scalability, and rapid adaptation to new hardware and software technologies. Furthermore, cloud computing services do not provide out-of-the-box bare metal setup, configuration, deployment, and updates, or allow for transitions to, from, or between bare metal and virtual IT infrastructure components. These and other limitations of cloud computing services can result in a number of computing, storage, and networking inefficiencies. For example, speed or latency inefficiencies in computing and networking may be introduced by the cloud or in applications or services that utilize the cloud.
[0080] The system and method of an exemplary embodiment provide a novel and unique IT infrastructure deployment, use and management. According to an exemplary embodiment, the complexity of resource selection, installation, interconnection, management and update is rooted in the core controller system and its parameter files, templates, rules and IT system status. The system includes a set of self-assembly rules and operation rules that are configured to enable components to self-assemble rather than requiring technicians to assemble, connect and manage. In addition, the system and method of an exemplary embodiment allow the use of self-assembly rules to achieve greater customization, scalability and flexibility without the current typical external planning documents. The system and method also allow for efficient resource use and reuse.
[0081] A system and method are provided that improve many problems and issues in current IT systems, whether physical or virtual in whole or in part. The system and method of an exemplary embodiment allows for flexibility, reduces variability and human error, and provides a structure that has the potential to improve system security.
[0082] Although there may be some solutions to one or more problems in current IT systems, such solutions do not comprehensively solve a wide range of problems as the exemplary embodiments described herein do. In addition, such existing solutions may solve specific problems while exacerbating other problems.
[0083] Some of the current challenges addressed include, but are not limited to, issues related to setup, configuration, infrastructure deployment, asset tracking, security, application deployment, service deployment, documentation regarding maintenance and compliance, maintenance, scaling, resource allocation, resource management, load balancing, software failures, updating / patching software and security, testing, recovering IT systems, change management, and hardware updates.
[0084] As used herein, IT systems may include, but are not limited to, servers, virtual and physical hosts, databases, and database applications, including, but not limited to, IT services, business computing services, computer applications, customer-facing applications, web applications, mobile applications, back-end, case management, customer tracking, ticketing, business tools, desktop management tools, billing, email, documentation, compliance, data storage, backup, and / or network management.
[0085] One problem users may face before setting up an IT system is predicting infrastructure requirements. Users may not know how much storage, computing power, or other requirements will be needed initially or over time during development or changes. According to one exemplary embodiment, the flexibility allowed for IT systems and infrastructure is that if system needs change, the self-deploying infrastructure (physical and / or virtual) of one exemplary embodiment can be used to automatically add, remove, or reallocate resources from within the infrastructure at a later time. Thus, the challenge of predicting future requirements when setting up a system is addressed by providing the ability to add to the system using the system's global rules, templates, and system state, and tracking changes to such rules, templates, and system state.
[0086] Other challenges may also relate to correct configuration, configuration consistency, interoperability, and / or interdependencies, which can include, for example, future incompatibilities caused by changes to configured system components or their configurations over time. For example, when initially setting up an IT system, there may be missing components or configuration errors in some components. Additionally, when setting up iterations of components or infrastructure components, there may be a lack of consistency between iterations. When changes are made to the system, configuration refinements may be necessary. Future infrastructure changes present a difficult choice between optimal configuration and flexibility. According to one exemplary embodiment, when the system is first deployed, global system rules are used to self-deploy the configuration from a template to the infrastructure components, making the configuration consistent, repeatable, or predictable, thereby allowing for optimal configuration. This initial system deployment can be performed on physical components, while subsequent components can be added or modified, and these subsequent components may or may not be physical. Furthermore, this initial system deployment can be performed on physical components, while subsequent environments can be cloned from the physical structure, and these subsequent environments may or may not be physical. This allows the system configuration to be optimal while allowing for minimally disruptive future changes.
[0087] During the deployment phase, challenges often arise with interoperability of bare metal and / or software-defined infrastructure. There may also be challenges with interoperability of software with other applications, tools, or infrastructure. These challenges may include, but are not limited to, those arising from deployed products sourced from different vendors. The inventors disclose an IT system that provides infrastructure interoperability, whether bare metal, virtualized, or any combination thereof. Thus, interoperability (the ability of components to work together) can be built into the disclosed infrastructure deployment, where the infrastructure is automatically configured and deployed. For example, different applications may depend on each other and reside on separate hosts. To enable such applications to interact with each other, the controller logic, templates, system state, and system rules described herein contain information and configuration instructions for configuring and tracking application interdependencies. Therefore, the infrastructure features discussed herein provide a way to manage how each application or service communicates with each other. For example, ensuring that an email service communicates correctly with an authentication service; and / or ensuring that a groupware service communicates correctly with an email service. Furthermore, this management can be implemented down to the infrastructure level to allow tracking, for example, how compute resources communicate with storage resources. Otherwise, the complexity of IT systems will increase in O(nn) manner.
[0088] As disclosed, automatic resource deployment does not require pre-configuration of operating system software because the controller can be deployed based on global system rules, templates, and IT system state / system self-awareness. According to an exemplary embodiment, users or IT professionals may not need to know whether the addition, allocation, or reallocation of resources will work together to ensure interoperability. According to an exemplary embodiment, additional resources can be automatically added to the network.
[0089] Using an application requires many different resources, typically including computing, storage, and networking resources. It also requires interoperability of resources and system components, including knowledge of what is in place and running, as well as interoperability with other applications. The application may need to connect to other services and obtain configuration files, and ensure that each component works together correctly. Application configuration can therefore be time and resource intensive. If there are interoperability issues with other applications, the application configuration may have a cascading effect on the rest of the infrastructure. This may result in operational outages or vulnerabilities. The inventors disclose automated application deployment for solving these problems. Thus, as disclosed by the inventors, applications can use knowledge of what is going on on the system and intelligent configuration to self-deploy by reading from IT system status, global system rules, and templates. In addition, according to an exemplary embodiment, pre-deployment testing of the configuration can be performed using the change management features described herein.
[0090] Another challenge addressed by the exemplary embodiment relates to issues that may arise with mid-configuration configurations where it is desirable to switch to a different vendor or other tool. According to one aspect of an exemplary embodiment, template conversion is provided between the controller's rules and templates and application templates from a specific vendor. This allows the system to automatically change vendors of software or other tools.
[0091] Many security issues are caused by misconfigurations, patch failures, and the inability to test patches before deployment. Security issues often arise during the configuration phase of a setup. For example, a misconfiguration could expose a sensitive application to the internet or allow forged emails from an electronic server. The inventors disclose a system setup that automatically configures itself to protect against attackers, avoid unnecessary exposure to attackers, and provide security engineers and application security architects with more knowledge of the system. Automation reduces security vulnerabilities caused by human error or misconfiguration. In addition, the disclosed infrastructure provides introspection between services and can allow rule-based access and limit communications between services to only those communications that are actually needed. The inventors disclose a system and method that has the ability to securely test patches before deployment, for example, as discussed with respect to change management.
[0092] Documentation is often a problematic area for IT management. During setup and configuration, the primary goal may often be to get components to work together. Often, this involves troubleshooting and a trial-and-error process where it is sometimes difficult to know what exactly made the system work. While the exact commands executed are often documented, the troubleshooting or trial-and-error process that may have achieved a working system is often not well documented or even not documented at all. Issues or deficiencies in documentation can create problems with audit trails and auditing. Documentation issues that arise can create problems with demonstrating compliance. Often, compliance issues may not be widely known when a system or its components are built. Applicable compliance decisions may only be known after the IT system is set up and configured. Therefore, documentation is crucial for auditing and compliance. The inventors disclose a system that includes a global system rules database, templates, and an IT system status database, which provides for automatically documented setup and configuration. Any configuration that occurs is recorded in the database. According to an exemplary embodiment, the automatically documented configuration provides an audit trail and can be used to demonstrate compliance. Inventory management can use the automatically recorded and tracked information.
[0093] Another challenge arising from IT system setup, configuration, and operation relates to inventory management of hardware and software. For example, it is often important to know how many servers exist, whether the servers are powered on and still running, what the capabilities of the servers are, in which rack each server is, which power supplies are connected to which servers, what network cards and what network ports each server is using, in which IT system the components are operating, and many other important considerations. In addition to inventory information, passwords and other sensitive information used for inventory management should also be managed effectively. Especially in larger IT systems, data centers, or data centers where equipment is frequently changed, the collection and retention of such information is a time-consuming task that is typically managed manually or using various software tools. Compliance protection for secure passwords is a significant risk factor that can be a significant challenge in ensuring a secure computing environment. The inventors disclose an IT system in which the collection and maintenance of inventory and operational status of all servers and other components is automatically updated, stored, and protected as part of the controller's IT system state, global system rules, templates, and controller logic.
[0094] In addition to the issues of setting up and configuring IT systems, the inventors have also disclosed an IT system that can also solve problems and issues that arise in the maintenance of IT systems. Many problems arise when a data center continues to operate in the presence of hardware failures, such as power failures, memory failures, network failures, network card failures and / or CPU failures, among other failures. Other failures arise when hosts are migrated during hardware failures. Therefore, the inventors disclose dynamic resource migration, such as migrating resources from one resource provider to another when a host goes down. In this case, according to an exemplary embodiment, the IT system can be migrated to other servers, nodes or resources, or other IT systems. A controller can report the status of the system. A copy of the data is on another host with a known and automatically set configuration. If a hardware failure is detected, any resources that may have been providing the hardware can be automatically migrated after the failure is automatically detected.
[0095] A significant issue with many IT systems is scalability. Growing businesses or other organizations typically add to or reconfigure their IT systems as they grow and their needs change. Problems arise when an existing IT system requires more resources, such as adding hard drive space, storage space, CPU processing, more network infrastructure; more endpoints, more clients, and / or more security measures. Problems also arise in configuration, setup, and deployment when different services and applications are needed or changes are made to the infrastructure. According to an exemplary embodiment, a data center can be automatically expanded. Nodes or resources can be dynamically and automatically added to or removed from a resource pool. Resources added and removed from a resource pool can be automatically allocated or reallocated. Services can be quickly provided to and moved to new hosts. A controller can dynamically detect and add more resources to a resource pool and know where to allocate / reallocate resources. A system according to an exemplary embodiment can be expanded from a single-node IT system to an expanded system across multiple data centers or IT systems requiring numerous physical and / or virtual nodes or resources.
[0096] The inventors have disclosed a system for implementing flexible resource allocation and management. The system includes computing resources, storage resources, and networking resources that may be in a resource pool and can be automatically allocated. A controller can identify new nodes or hosts on the network and then configure the new nodes or hosts so that they can become part of the resource pool. For example, whenever a new server is inserted, the controller will configure the new server as part of the resource pool, and the new server can be added to the resources and can begin to be used automatically. Nodes or resources can be detected by the controller and added to different pools. Resource requests can be made to the controller, for example, through an API request. The controller can then deploy or allocate the required resources from the pool according to rules. This allows the controller and / or applications to balance the load and dynamically distribute resources based on the needs of the requests through the controller.
[0097] Examples of load balancing include, but are not limited to: deploying new resources when a hardware or software failure occurs; deploying one or more instances of the same application in response to increased user load; and deploying one or more instances of the same application in response to an imbalance in storage, computing, or networking demands.
[0098] Problems involved in making changes to active IT systems or environments can cause significant, and sometimes catastrophic, problems for users or entities that rely on these systems for continued operation. These outages represent not only a potential loss of system usage, but also financial losses due to data loss and the significant time, personnel, and financial resources required to repair the problems. These problems can be exacerbated by the difficulty of reconfiguring systems due to configuration documentation errors or a lack of system understanding. Due to this issue, many IT system users are reluctant to patch IT resources to eliminate known security risks, making these resources more vulnerable to security breaches.
[0099] Many issues that arise in the maintenance of IT systems are related to software failures due to change management or controls that may require configuration. Situations where such failures may occur include, but are not limited to: upgrading to a new software version; migrating to a different piece of software; changing password or authentication management; switching between services or between different service offerings.
[0100] Manually configured and maintained infrastructure is often difficult to recreate. Recreating infrastructure can be important for a number of reasons, including but not limited to: undoing problematic changes, recovering from power outages or other disasters. Problems in manually configured systems are difficult to diagnose. Manually configured and maintained infrastructure is difficult to recreate. Furthermore, system administrators can easily make mistakes, such as issuing incorrect commands, which have been known to disable computer systems.
[0101] Making changes to active IT systems or environments can cause significant, and sometimes catastrophic, problems for users or entities that rely on these systems for continued operation. These outages not only represent a potential loss of system usage, but they can also result in data loss and financial losses due to the significant time, personnel, and financial resources required to repair the problems. These problems can be exacerbated by the difficulty of reconfiguring systems due to errors in configuration documentation or a lack of system understanding. Furthermore, in many cases, restoring a system to its previous state after a significant or major change has occurred is difficult.
[0102] Furthermore, potential technical issues that arise when changes are made to a live environment can have cascading effects. These cascading effects can make reverting to a pre-change state challenging and sometimes impossible. Consequently, even if a problem arises with an implemented change and the change needs to be rolled back, the state of the system has already been altered. Undoing infrastructure and system administration errors, as well as erroneous changes to a production environment, has recently been identified as an unresolved problem. Furthermore, testing system changes before deployment to a live environment is known to be problematic.
[0103] Thus, the inventors have disclosed several exemplary embodiments of systems and methods configured to restore changes to an active system back to a pre-change state. Furthermore, the inventors have disclosed (provided) a system and method configured to achieve substantial restoration of the state of a system or environment that is subject to real-time changes, which may prevent or ameliorate one or more of the problems described above.
[0104] According to a variation of an exemplary embodiment, an IT system has complete system knowledge of global system rules, templates, and IT system state. Infrastructure can be cloned using this complete system knowledge. Systems or system environments can be cloned into software-defined infrastructure or environments. A system environment, including an active volatile database (referred to as a production environment), can be written to a non-volatile, read-only database to serve as a development environment during development and testing. Required changes can be made and tested in the development environment. Users or controller logic can make changes to global rules to create new versions. Rule versions can be tracked. According to another aspect of an exemplary embodiment, the newly developed environment can then be automatically implemented. The previous production environment can also be maintained or kept in a fully functional state, allowing modifications to the previous state of the production environment without data loss. The development environment can then be started using new specifications, rules, and templates, and the database or system can be synchronized with the production database, switching to a writeable database. The original production database can then be switched to a read-only database, to which the system can be restored if a recovery is required.
[0105] With respect to upgrading or patching software, if it is detected that a service needs upgrading or patching, a new host can be deployed. In the event of a failure due to upgrading or patching, a new service can be deployed when change recovery is possible as described above.
[0106] Hardware upgrades are crucial in many situations, especially when the latest hardware is essential. An example of this type of situation occurs in the high-frequency trading industry, where IT systems with millisecond speed advantages can enable users to achieve superior trading results and profits. In particular, ensuring interoperability with existing infrastructure presents challenges, ensuring that the new hardware understands how to communicate with the protocols and work with the existing infrastructure. In addition to ensuring interoperability, these components also need to be integrated with the existing setup.
[0107] refer to Figure 1 , shows an IT system 100 according to an exemplary embodiment. The system 100 can be one or more types of IT systems, including but not limited to those described herein.
[0108] A user interface (UI) 110 is shown coupled to a controller 200 via an application programming interface (API) application 120, which may or may not reside on a separate physical or virtual server. The controller 200 may be deployed on one or more processors and one or more memories to implement any of the control operations discussed herein. Instructions for execution by the one or more processors to implement such control operations may reside on a non-transitory computer-readable storage medium, such as a processor memory. The API 120 may include one or more API applications, which may be redundant and / or operate in parallel. The API application 120 receives requests to configure system resources, parses the requests, and passes them to the controller 200. The API application 120 receives one or more responses from the controller, parses the one or more responses, and passes them to the UI (or application) 110. Alternatively or in addition, applications or services may communicate with the API application 120. The controller 200 is coupled to one or more computing resources 300, one or more storage resources 400, and one or more networking resources 500. Resources 300, 400, 500 may or may not reside on a single node. One or more of the resources 300, 400, 500 may be virtual. The resources 300, 400, 500 may or may not reside on multiple nodes, or may reside on multiple nodes in various combinations. A physical device may include one or more, or each, of the resource types including, but not limited to, computing resources 300, storage resources 400, and networking resources 500. Resources 300, 400, 500 may also include resource pools, whether physically located or not, and whether virtual or not. Bare metal computing resources may also be used to enable the use of virtual or container computing resources.
[0109] In addition to the known definition of a node, as used herein, a node can be any system, device, or resource connected to one or more networks, or other functional unit that performs a function on a standalone device or a network-connected device. A node can also include, but is not limited to, for example, a server, a service / application / multiple services on a physical or virtual host, a virtual server, and / or multiple or single services running on a multi-tenant server or within a container.
[0110] The controller 200 may include one or more physical or virtual controller servers, which may also be redundant and / or operate in parallel. The controller may run on a physical or virtual host that serves as a computing host. As an example, the controller may include a controller running on a host that is used for other purposes, such as because it has access to sensitive resources. The controller receives requests from the API application 120, parses the request, and makes appropriate task assignments to other resources and directs them to those other resources; monitors resources and receives information from them; maintains the system's status and change history; and may communicate with other controllers in the IT system. The controller may also include the API application 120.
[0111] As defined herein, computing resources may include a single real or virtual computing node or a resource pool having one or more computing nodes. A computing resource or computing node may include one or more physical or virtual machines or container hosts that can host one or more services or run one or more applications. A computing resource may also be on hardware designed for multiple purposes, including but not limited to computing, storage, caching, networking, specialized computing, including but not limited to GPUs, ASICs, coprocessors, CPUs, FPGAs, and other specialized computing methods. PCI Express switches or similar devices may be added to such devices, and such devices may be dynamically added in this manner. A computing resource or computing node may include or run multiple different virtual machines that contain running services or applications, or may be one or more hypervisors or container hosts that are virtual computing resources. While the focus of a computing resource may be on providing computing functionality, it may also include data storage and / or networking capabilities.
[0112] Storage resources as defined herein may include storage nodes or storage resource pools. Storage resources may include any data storage medium, such as fast, slow, hybrid, cache storage medium and / or RAM. Storage resources may include one or more types of networks, machines, devices, nodes, or any combination thereof that may or may not be directly attached to other storage resources. According to aspects of an exemplary embodiment, storage resources may be bare metal or virtual resources, or a combination thereof. Although the focus of storage resources may be on providing storage functionality, it may also include computing and / or networking capabilities.
[0113] One or more networking resources 500 may include a single networking resource, multiple networking resources, or a pool of networking resources. The one or more networking resources may include one or more physical or virtual devices, one or more appliances, switches, routers, or other interconnects between system resources, or applications used to manage networking. Such system resources may be physical or virtual and may include computing resources, storage resources, or other networking resources. Networking resources provide connectivity between external networks and application networks and may host core network services, including but not limited to Domain Name System (DNS), Dynamic Host Configuration Protocol (DHCP), subnet management, Layer 3 routing, Network Address Translation (NAT), and other services. Some of these services may be deployed on computing resources, storage resources, or networking resources in physical or virtual machines. Networking resources may utilize one or more architectures or protocols, including but not limited to InfiniBand, Ethernet, Remote Direct Memory Access (DMA) over Converged Ethernet (RoCE), Fibre Channel, and / or Omnipath, and may include interconnects between multiple architectures. Networking resources may or may not have Software Defined Networking (SDN) capabilities. The controller 200 may be able to directly modify the networking resources 300 to configure the topology of the IT system using SDN virtual local area networks (VLANs), etc. Although the focus of the networking resources may be on providing networking functions, it may also include computing and / or storage capabilities.
[0114] As used herein, an application network refers to a networked resource used to connect or couple applications, resources, services, and / or other networks, or to couple users and / or clients to applications, resources, and / or services, or any combination thereof. An application network may include a network used by a server to communicate with other application servers (physical or virtual) and with clients. An application network may communicate with machines or networks external to system 100. For example, an application network may connect a web front end to a database. A user may connect to a web application via the Internet or another network that may or may not be managed by a controller.
[0115] According to an exemplary embodiment, computing resources 300, storage resources 400, and networking resources 500 may each be automatically added, removed, provisioned, allocated, reallocated, configured, reconfigured, and / or deployed by controller 200. According to an exemplary embodiment, additional resources may be added to the resource pool.
[0116] Although a user interface 110 is shown, such as a Web UI or other user interface that a user 105 can utilize to access and interact with the system, alternatively or in addition, applications can also communicate or interact with the controller 200 through one or more API applications 120 or in other ways. For example, a user 105 or an application can send requests including, but not limited to: building an IT system; building a separate stack within an IT system; creating a service or application; migrating a service or application; changing a service or application; removing a service or application; cloning a stack to another stack on a different network; creating, adding, removing, setting up or configuring, or reconfiguring resources or system components.
[0117] Figure 1 The system 100 may include a server having connectors or other communication interfaces to various elements, components or resources that may be physical or virtual, or any combination thereof. According to one variation, Figure 1 The illustrated system 100 may include bare metal servers with connectors.
[0118] As described in more detail herein, the controller 200 can be configured to power on resources or components, automatically set up, configure, and / or control the powering on of resources, add resources, allocate resources, manage resources, and update available resources. The actuation process can begin by actuating the controller so that the order in which devices are actuated can be consistent, rather than depending on the user actuating the device. The process can also involve detecting that a resource has been actuated.
[0119] refer to Figures 2A to 10 , showing a controller 200 , controller logic 205 , a global system rules database 210 , an IT system state 220 , and a template 230 .
[0120] The system 100 includes global system rules 210. The global system rules 210 may, among other things, declare rules for setting up, configuring, starting, allocating, and managing resources that may include computing, storage, and networking. The global system rules 210 include the minimum requirements for the system 100 to be in a correct or desired state. The requirements may include the IT tasks that are expected to be completed, as well as an updateable list of expected hardware that is required to predictably build the required system. The updateable list of expected hardware allows the controller to verify that the required resources are available (from the start of the rule or use of the template, for example, before starting the rule or use of the template). The global rules may include a list of operations required for various tasks and corresponding instructions related to the sequencing of the operations and tasks. For example, the rules may specify the order of: driving components; starting resources, applications, and services; dependencies; and the order in which different tasks are initiated, such as loading, configuring, starting, reloading applications, or updating hardware. The rules 210 may also include one or more of the following: a list of resource allocations required by applications and services; a list of templates that can be used; a list of applications to be loaded and how they are configured; a list of services to be loaded and how they are configured; a list of application networks and which applications run on which networks; a list of configuration variables and user-specific application variables specific to different applications; an expected state that allows the controller to review the system state to verify that the state is as expected and that the results of each instruction are as expected; and / or a version list that includes a list of rule changes (e.g., snapshots) that allows tracking changes to the rules and the ability to test or revert to different rules in different situations. The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on physical resources. The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on virtual resources. The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on a combination of physical and virtual resources.
[0121] Figure 2M A set of example system rules 210 that may take the form of global system rules is shown. Figure 2M The illustrated set of exemplary system rules 210 may be loaded into the controller 200 or obtained by querying the system status (see 210.1). Figure 2MIn the example of , system rules 210 include a set of instructions, which may take the form of configuration routines 210.2, and also include data 210.3 for creating and / or recreating an IT system or environment. Configuration rules within system rules 210 may specify how to locate templates 230 (which may reside in a file system, disk, storage resource, or within the system rules) via a required template list 210.7. Controller logic 205 may also locate templates 230 before processing them and ensure that they exist before invoking system rules 210. System rules 210 may include system rule subsets 210.15, and these subsets 210.15 may be executed as part of configuration routines 210.2.
[0122] Furthermore, subsystem rules 210.15 can be used, for example, as a tool to build a system with integrated IT applications (which can then be processed using system rule execution routine 210.16 and subsequently updated to reflect the addition of 210.15 to the system state and current configuration rules). Subsystem rules 210.15 can also be located elsewhere and loaded into system state 220 through user interaction. For example, you can also treat subsystem rules 210.15 as a playbook, which can be retrieved and run (which then causes global system rules 210 to be updated, so you can replay the playbook if you want to clone a system).
[0123] A configuration routine 210.2 may be a set of instructions for building a system. The configuration routine 210.2 may also include subsystem rules 210.15 or system state pointers 210.8, if desired by the practitioner. When running the configuration routine 210.2, the controller logic 205 may process a series of templates in a specific order (210.9), optionally allowing for parallel deployment, but maintaining proper dependency handling (210.12). The configuration routine 210.2 may optionally invoke API calls 210.10, which may set configuration parameters 210.5 on the application that may be configured by processing the templates according to 210.9. Additionally, required services 210.11 are services that need to be started and running if the system is to make the API calls 210.10.
[0124] Routine 210.2 may also include a process, program, or method for performing a data load (210.13) with respect to volatile data 210.6, including, but not limited to, copying data, transferring a database to a computing resource, pairing a computing resource with a storage resource, and / or updating system state 220 with the location of volatile data 210.6. Volatile data pointers (see 210.4) may be maintained using data 210.3 to locate volatile data that may be stored elsewhere. Data load routine 210.13 may also be used to load configuration parameters 210.5 if they are located in a non-standard data store (e.g., contained in a database).
[0125] The system rules 210 may also include a resource list 210.18 that may indicate which components are assigned to which resources and will allow the controller logic 205 to determine whether appropriate resources and / or hardware are available. The system rules 210 may also include an alternative hardware and / or resource list 210.19 for alternative deployments (e.g., for a development environment where a software engineer may want to perform real-time testing but does not want to allocate an entire data center). The system rules may also include data backup and / or standby routines 210.17 that provide instructions on how to back up the system and use spare parts to achieve redundancy. Examples of data backup systems and / or backup routines that implement backup rules include, but are not limited to, the data backup routines referenced herein. Figure 21A -Those mentioned by J.
[0126] After each action is taken, the system state 220 may be updated and the query (which may include a write) saved as a system state query 210.14.
[0127] Figure 2N The diagram is processed by the controller logic 205 Figure 2M 20. At step 210.20, the controller logic 205 checks to ensure that appropriate resources are available (see Figure 2M Otherwise, alternative configurations may be checked at step 210.21. A third option may include prompting the user to select a configuration that may be affected by Figure 2M List of supported alternative configurations for template 230 referenced in 210.7.
[0128] At step 210.22, the controller logic can then ensure that the computing resource (or any appropriate resource) gains access to the volatile data. This may involve connecting to a storage resource or adding the storage resource to the system state 220. At step 210.23, the configuration routines are processed, and as each routine is processed, the system state 220 is updated (step 210.24). The system state 220 can also be queried to see if certain steps have been completed before continuing (step 210.25).
[0129] The configuration routine processing steps shown in Figure 210.23 may include any of the procedures (or combinations thereof) of 210.26. The configuration routine processing steps may also include other procedures. For example, the processing at 210.26 may include template processing (210.27), loading configuration data (210.28), loading static data (210.29), loading dynamic volatile data (210.30), and / or coupling services, applications, subsystems, and / or environments (210.31). Such procedures within 210.26 may be repeated in a loop or run in parallel, as some system components may be independent while other system components may be interdependent. Controller logic, service dependencies, and / or system rules may indicate which services may be dependent on each other, and the services may be coupled to further construct the IT system from the system rules.
[0130] Global system rules 210 may also include storage expansion rules. Storage expansion rules provide a set of rules for automatically adding storage resources to existing storage resources within the system, for example. Furthermore, storage expansion rules may provide triggers by which applications running on one or more computing resources will know when to request storage expansion (or by which the controller 200 may know when to expand the storage of a computing resource or application). Controller 200 may allocate and manage new storage resources and may merge or integrate these storage resources with existing storage resources for specific operational resources. Such specific operational resources may include, but are not limited to, computing resources within the system, applications running computing resources within the system, virtual machines, containers, or physical or virtual computing hosts, or any combination thereof. An operational resource may signal to the controller 200, for example, through a storage space query, that the operational resource is running out of storage space. In-band management connection 270, SAN connection 280, or any other networking or coupling to controller 200 may be used for such queries. Out-of-band management connection 260 may also be used. These storage expansion rules (or a subset of these storage expansion rules) may also be applied to resources that are not operational.
[0131] Storage expansion rules specify how to locate, connect, and configure new storage resources within the system. The controller registers the new storage resource in the system state 220 and informs the operational resources where and how to connect to it. The operational resources use this registration information to connect to the storage resource. The controller 200 may merge the new storage resource with existing storage resources, or it may add the new storage resource to a volume group.
[0132] Figure 2B An example flow illustrating the operation of a set of example storage expansion rules is shown. At step 210.41, the operational resource determines that its storage is low based on a trigger point or other factors. At step 210.42, the operational resource connects to the controller 200 via an in-band management connection 270, a SAN connection 280, or another type of connection visible to the operating system. Through this connection, the operational resource can notify the controller 200 that its storage is low. At step 210.43, the controller configures storage resources to expand storage capacity for the operational resource. At step 210.44, the controller provides the operational resource with information regarding the location of the newly configured storage resource. At step 210.45, the operational resource connects to the newly configured storage resource. At step 210.46, the controller adds a mapping of the new storage resource location to the system state 220. The controller can then add the new storage resource to the volume group assigned to the operational resource (step 210.47), or the controller can add the assignment of the new storage resource to the operational resource to the system state 220 (step 210.48).
[0133] Figure 2C Diagram for execution Figure 2B 41 and 210.42 in . At step 210.50, the controller sends a critical command via the out-of-band management connection 260 to view a storage status update on a monitor or console regarding the running resource. For example, the monitor may be an ipmi console that can be accessed via the out-of-band connection 260. As an example, the out-of-band connection 260 may be plugged into a USB as a keyboard / mouse and into a VGA monitoring port. At step 210.51, the running resource displays the information on the screen. At step 210.52, the controller then reads the information presented on the monitor or console via the out-of-band management connection 260 and screen scraping or similar operation; wherein this read information may indicate a low storage condition based on a trigger point. The process flow may then continue Figure 2B Step 210.43.
[0134] Figure 2D Diagram for execution Figure 2BAnother alternative example of steps 210.41 and 210.42. At step 210.55, the operating resource automatically displays information on a monitor or console for the controller to read. At step 210.56, the controller automatically, periodically, or continuously reads the monitor or console to view the operating resource. In response to this reading, the controller learns that the operating resource's memory is low (step 210.57). The process flow can then continue Figure 2B Step 210.43.
[0135] Controller 200 also includes a library of templates 230, which can include bare metal and / or service templates. These templates may include, but are not limited to, email, file storage, IP telephony, software billing, software XMPP, wikis, version control, account authentication management, and third-party applications that may be configurable via a user interface. Templates 230 can be associated with resources, applications, or services and can serve as recipes defining how to integrate such resources, applications, or services into the system.
[0136] Likewise, a template may include a set of information for creating, configuring, and / or deploying resources, or for establishing an application or service loaded on a resource. Such information may include, but is not limited to, a kernel, an initrd file, a file system or file system image, files, configuration files, configuration file templates, information for determining appropriate settings for different hardware and / or computing backends, and / or other options that may be used to configure resources to drive an application and to allow and / or facilitate the creation, startup, or running of an operating system image for an application.
[0137] The template may contain information that can be used to deploy the application on a variety of supported hardware types and / or computing backends, including but not limited to: multiple physical server types or components, multiple hypervisors running on multiple hardware types, and container hosts that can be hosted on multiple hardware types.
[0138] Templates can be used to create boot images for applications or services running on computing resources. Templates and images derived from them can be used to create applications, deploy applications or services, and / or arrange resources for various system functions, which enables and / or facilitates the creation of applications. Templates can contain variable parameters within files, file systems, and / or operating system images, which can be overridden by configuration options from default settings or settings specified by a controller. Templates can contain configuration scripts for configuring applications or other resources, and they can utilize configuration variables, configuration rules, and / or default rules or variables. These scripts, variables, and / or rules can include specific rules, scripts, or variables for specific hardware or other resource-specific parameters, such as hypervisor (in the case of virtualization) or available memory. Templates can contain files in the form of binary resources, compilable source code that specifies binary resources or hardware or other resource-specific parameters, multiple sets of specific binary resources, or source code with compilation instructions for specific hardware or other resource-specific parameters, such as hypervisor (in the case of virtualization) or available memory. Templates can include a set of information independent of what is currently running on a resource.
[0139] A template may include a base image. The base image may include a base operating system file system. The base operating system may be read-only. The base image may also include basic operating system tools that are independent of the running system. The base image may include basic directories and operating system tools. The template may include a kernel. The kernel or kernels may include an initrd, or multiple kernels may be configured for different hardware types and resource types. The image may be derived from a template ad loaded into one or more resources or deployments. The loaded image may also include startup files, such as the kernel or initrd corresponding to the template.
[0140] An image may include template file system information that can be loaded into a resource based on a template. A template file system may configure an application or service. A template file system may include a shared file system that is shared by all resources, or similar resources, for example to save storage space for storing a file system or to facilitate the use of read-only files. A template file system or image may include a set of files that are shared by deployed services. A template file system may be preloaded onto a controller or downloaded. A template file system may be updated. A template file system may allow for relatively faster deployment because it does not require reconstruction. Sharing a file system with other resources or applications may allow for reduced storage because files are not copied unnecessarily. This may also allow for easier recovery from failures because only files other than the template file system need to be restored.
[0141] The template boot file may include a kernel and / or initrd or similar file system used to assist in the boot process. The boot file may start the operating system and set up the template file system. The initrd may include a small temporary file system that explains how to set up the template so that the template can be started.
[0142] The template may also include template BIOS settings. Template BIOS settings can be used to set optional settings for applications running on the physical host. If used, as described in this article about Figures 1 to 12 The described out-of-band management 260 can be used to start resources or applications. The physical host can use the out-of-band management network 260 or CDROM to start resources or applications. The controller 200 can set application-specific BIOS settings defined in such a template. The controller 200 can use the out-of-band management system to make direct BIOS changes through APIs specific to the specific resource. The settings can be verified through the console and image recognition. Therefore, the controller 200 can use the console feature and utilize a virtual keyboard and mouse to make BIOS changes. The controller can also use a UEFI shell and can directly input into the console, and can use image recognition to verify successful results, correctly enter commands, and ensure successful setting changes. If there is a bootable operating system available for BIOS changes or updates to a specific BIOS version, the controller 200 can remotely load a disk image or ISO to start the application running the operating system, update the BIOS, and allow configuration changes to be made in a reliable manner.
[0143] A template may also include a list of template-specific supporting resources, or a list of resources required to run a specific application or service.
[0144] The template image or portions of the image or template may be stored on the controller 200 , or the controller 200 may move or copy them to the storage resource 410 .
[0145] Figure 2EAn example template 230 is shown. The template contains all the information needed to create an application or service. Template 230 may also include information for different hardware types, alternative data, files, and binaries that provide similar or identical functionality. For example, there may be a file system blob 232 for / usr / bin and / or bin and binaries 234 compiled for different architectures. Template 230 may also include a daemon 233 or script 231. Daemon 233 is a binary or script that may run at boot time when the host is powered on and ready. In some cases, daemon 233 may drive an API that may be accessible to the controller and may allow the controller to change host settings (and the controller may subsequently update active system rules). Daemons can also be shut down and restarted through out-of-band management 260 or in-band management 270, discussed above and below. These daemons may also drive common APIs to provide dependency services for new services (for example, a common web server API that communicates with the API that controls Nginx or Apache). The script 231 may be an installation script that may be run when or after the image is booted, or after a daemon is started or a service is enabled.
[0146] The template 230 may also include a kernel 235 and a pre-boot file system 236. The template 230 may also include multiple kernels 235 and one or more pre-boot file systems (such as an initrd or initramfs for Linux, or a read-only ramdisk for BSD) for different hardware and different configurations. The initrd may also be used to mount a file system blob 232 presented as an overlay, and to mount the root file system on remote storage by booting into the initramfs 236, which may optionally be connected to storage resources via a SAN connection 280 as described below.
[0147] File system blob 232 is a file system image that can be divided into separate blobs. Blobs may be interchangeable based on configuration options, hardware type, and other setup differences. A host booted from template 230 can boot from a union file system containing multiple blobs (such as overlayfs), or from an image created from one or more file system blobs.
[0148] Template 230 may also include or be linked to additional information 237, such as volatile data 238 and / or configuration parameters 239. For example, volatile data 238 may be included within template 230, or the volatile data may be included externally. The volatile data may be in the form of a file system blob 232 or other data store, including but not limited to a database, a flat file, a file stored in a directory, a compressed file archive, or a Git or other version control repository. Furthermore, configuration parameters 239 may be included externally or internally within template 230 and may optionally be included in system rules and applied to template 230.
[0149] System 100 also includes an IT system state 220 that tracks, maintains, changes, and updates the status of system 100, including but not limited to resources. System state 220 can track available resources, which informs controller logic whether resources are available to implement rules and templates, as well as what resources are available to implement them. System state can also track used resources, allowing controller logic 205 to check efficiency and utilization, and to determine whether resources need to be switched for upgrades or other reasons, such as to improve efficiency or achieve priority. System state can also track what applications are running. Based on the system state, controller logic 205 can compare expected application performance with actual application performance and determine whether any corrections are needed. System state 220 can also track where applications are running. Controller logic 205 can use this information for efficiency assessment, change management, updates, troubleshooting, or audit trails. System state can also track networking information, such as which networks are operating or currently running, or configuration values and history. System state 220 can also track change history. System state 220 can also track which templates are used in which deployments based on global system rules that dictate which templates are used. The history can be used for auditing, alarming, change management, building reports, tracking versions and configurations or configuration variables related to hardware and applications. The system state 220 can maintain configuration history for the purpose of auditing, compliance testing, or troubleshooting.
[0150] The controller includes logic 205 for managing all information contained in the system state, templates, and global system rules. Controller logic 205, global system rules database 210, IT system state 220, and templates 230 are managed by controller 200 and may or may not reside on controller 200. Controller logic or application 205, global system rules database 210, IT system state 220, and templates 230 may be physical or virtual and may or may not be distributed services, distributed databases, and / or files. API application 120 may be included in controller logic / controller application 205.
[0151] The controller 200 may run as a standalone machine and / or may include one or more controllers. The controller 200 may include a controller service or application and may run within another machine. The controller machine may first start the controller service to ensure an orderly and / or consistent startup of the entire stack or set of stacks.
[0152] Controller 200 may control the computing resources, storage resources, and networking resources of one or more stacks. Each stack may or may not be controlled by a different subset of the rules within global system rules 210. For example, there may be a pre-production stack, a production stack, a development stack, a test stack, a parallel stack, a backup stack, and / or other stacks with different functions within the system.
[0153] The controller logic 205 can be configured to read and interpret global system rules to achieve a desired IT system state. The controller logic 205 can be configured to use templates to build system components, such as applications or services, according to the global rules, and allocate, add, or remove resources to achieve the desired IT system state. The controller logic 205 can read the global system rules, generate a list of tasks to achieve the correct state, and issue instructions for fulfilling the rules based on available operations. The controller logic 205 can include logic for: performing operations, such as starting the system, adding, removing, reconfiguring resources; and identifying what can be done. The controller logic can check the system state at startup time and at regular time intervals to see if the hardware is available, and if so, the hardware can perform the task. If the necessary hardware is not available, the controller logic 205 uses the global system rules 210, templates 220, and the hardware available according to the system state 230 to present alternative options and modify the global rules and / or system state 220 accordingly.
[0154] Controller logic 205 may know what variables are required, what user input is needed to proceed, or what the user requires to run in the system. Controller logic can use a list of templates from global system rules and compare it to the templates required in the system state to ensure the required templates are available. Controller logic 205 can determine from the system state database whether resources on the template's specific supported resource list are available. Controller logic 205 can allocate resources, update the state, and proceed to the next set of tasks to implement the global rules. Controller logic 205 can launch / run applications on the allocated resources as specified in the global rules. The rules may specify how to build applications from templates. Controller logic 205 can retrieve one or more templates and configure the application based on the variables. The templates may tell controller logic 205 which kernel, startup files, file systems, and supported hardware resources are required. Controller logic 205 can then add information about the application deployment to the system state database. After each instruction, controller logic 205 can check the system state database against the expected state according to the global rules to verify that the expected operation was completed correctly.
[0155] The controller logic 205 may use versions according to version rules.The system state 220 may have a database of which rule versions have been used in different deployments.
[0156] The controller logic 205 may include validation logic for rule optimization and efficient sequencing. The controller logic 205 may be configured to optimize resources. Information in the system state, rules, and templates related to applications currently running or expected to run may be used by the controller logic to implement efficiency or priority with respect to resources. The controller logic 205 may use information in the "used resources" section of the system state 220 to determine efficiency or the need to switch resources for upgrades, reuse, or other reasons.
[0157] The controller can view application execution based on the system state 220 and compare the application execution with the expected application execution according to the global rules. If the application is not running, the controller can start the application. If the application should not run, the controller can stop the application and reallocate resources as appropriate. The controller logic 205 may include a database of resource (computing, storage, networking resources) specifications. The controller logic may include logic for identifying the types of resources available for the system that can be used. This can be performed using the out-of-band management network 260. The controller logic 205 can be configured to use out-of-band management 260 to identify new hardware. The controller logic 205 can also obtain information about change history, rules used, and versions from the system state 220 for auditing, building reports, and change management purposes.
[0158] Figure 2FAn exemplary process flow is shown for the controller logic 205 regarding processing a template 230 and obtaining an image to power on, enable, and / or enable a resource, which for the purposes of this example may be referred to as a host. This process may also include configuring storage resources and coupling storage and compute hosts and / or resources. The controller logic 205 is aware of the hardware resources available in the system 100, and the system rules 210 may indicate which hardware resources can be utilized. The controller logic 205 parses the template 230 at step 205.1, which may include an instruction file that may be executed to cause the controller logic to collect information generated by Figure 2E The instruction file can be in JSON format. At step 205.2, the controller logic collects a list of required file buckets. Furthermore, at step 205.3, the controller logic 205 collects the required hardware-specific files into buckets. These hardware-specific files are referenced by the hardware and, optionally, by a hypervisor (or container host system, or multi-tenant system). If the hardware is to run on a virtual machine, a hypervisor (or container host system, or multi-tenant system) may also be required to reference them.
[0159] If hardware specific files exist, the controller logic will collect the hardware specific files at step 205.4. In some cases, the file system image may contain a kernel and initramfs as well as a directory containing kernel modules (or the kernel modules will eventually be placed into a directory). The controller logic 205 then selects a compatible, appropriate base image at step 205.5. A base image contains operating system files that may not be specific to the application or image derived from the template 230. Compatibility in this context means that the base image contains the files required to turn the template into a working application. The base image can be managed outside of the template as a mechanism for saving space (and typically, the base image may be the same for several applications or services). Additionally, at step 205.6, the controller logic 205 selects one or more buckets with executable files, source code, and hardware specific configuration files. The template 230 may reference other files including, but not limited to, configuration files, configuration file templates (which are configuration files containing placeholders or variables filled with variables from the system rules 210 that may become known in the template 230, such that the controller 200 may convert the configuration template into a configuration file and optionally change the configuration file through an API endpoint), binaries, and source code (which may be compiled when the image is booted). At step 205.7, the hardware specific instructions corresponding to the components selected at steps 205.4., 205.5, and 205.6 may be loaded as part of the image being booted. The controller logic 205 obtains the image from the selected components. For example, there may be different pre-installed scripts for a physical host compared to a virtual machine, or there may be differences for a powerpc compared to an x86.
[0160] At step 205.8, the controller logic 205 mounts overlayfs and repackages the main files into a single file system blob. When using multiple file system blobs, the image can be created by extracting the compressed archive and / or fetching git from the multiple blobs. If step 205.8 is not performed, the file system blobs can remain separate, and the image can be created as a set of file system blobs and mounted using a file system that can mount multiple smaller file systems together, such as overlayfs. The controller logic 205 can then locate a compatible kernel (or a kernel specified in the system rules 210) at step 205.9 and an applicable initrd at step 205.10. A compatible kernel can be one that satisfies the dependencies of the template or the resources used to implement the template. A compatible initrd can be one that loads the template onto the required computing resources. Typically, an initird is available for physical resources, allowing it to mount storage resources before a full boot (because the root file system may be remote). The kernel and initrd can be packaged into a filesystem blob for direct kernel booting, or for use on a physical host using kexec to change the kernel on the active system after booting the initial operating system.
[0161] The controller then configures one or more storage resources to allow one or more computing resources to drive one or more applications and / or one or more images using any of the techniques described in 205.11, 205.12, and / or 205.13. At 205.11, an overlayfs file may be provided as a storage resource. At 205.12, a file system is presented. For example, a storage resource may present a composite file system, or a computing resource may simultaneously mount multiple file system blobs using a file system similar to overlayfs. At 205.13, the blobs are sent to the storage resource before presenting the file system.
[0162] Figure 2G and Figure 2H Shown for Figure 2F An example process flow for steps 205.11 and 205.12 is provided below. Further, the system may employ processes and rules for connecting computer resources to storage resources, which may be referred to as a storage connection process. Appendix A of the accompanying documentation provides a description of such a storage connection process, except for the steps described in the example process flow. Figure 2G and Figure 2H Instances other than those shown. Figure 2GAn example process flow for connecting storage resources is shown. Some storage resources may be read-only, and other storage resources may be writable. The storage resource may manage its own write locks so that there are no simultaneous writes that would cause a race condition, or the system state 220 may track (see, e.g., step 205.20) which connections can write to the storage resource and / or prevent multiple read-write connections from connecting to the resource (step 205.21). The controller logic or resource itself may query the controller system status 220 for the location and transport type of the storage resource (e.g., Internet Small Computer System Interface (ISCSI, iSCSI, or iscsi), ISCSI for Remote Direct Storage Access (RDMA, or rdma) (ISER, iSER, or iser), Non-Volatile Memory Host Controller Interface Specification over Fibre Channel (NVMEOF, or nvmeof), Fibre Channel (FC, or fc), Fibre Channel over Ethernet (FCOE, or FCoE), Network File System (NFS, or nfs), nfs over rdma, Distributed File System (AFS, or afs), Common Internet File System (CIFS, or cifs), Windows Share) (step 205.22). If the computing resource is virtualized, the hypervisor (e.g., via a hypervisor daemon) may handle the connection to the storage resource (step 205.23). This may have desirable security advantages, as the virtual machine may not be aware of the SAN 280.
[0163] Referring to step 205.24, the process for connecting computing resources and storage resources may be specified in system rules 210. The controller logic then queries system state 220 to ensure that the resources are available and writable (if necessary) (step 205.22). System state 220 can be queried via any of a variety of techniques, such as SQL queries (or other types of database queries), JSON parsing, etc. The query will return the information necessary for the computing resource to connect to the storage resource. The controller 200, system state 220, or system rules 210 may provide authentication credentials for the computing resource to connect to the system state (step 205.25). The computing resource will then update system state 220 directly or via the controller (step 205.26).
[0164] Figure 2HAn example boot process is shown in which a physical, virtual, or other type of computing resource, application, service, or host drives and connects to a storage resource. The storage resource may optionally utilize a fused file system and / or expandable volumes. In the event that a controller or other system enables a physical host, the physical host may be preloaded with an operating system for configuring the system. Thus, at step 205.31, the controller may preload a boot disk with initramfs. Additionally, the controller 200 may use the out-of-band management connection 260 to network boot a preliminary operating system (step 205.30) and then optionally preload the preliminary operating system onto the host (step 205.31). The initramfs is then loaded at step 205.32 and the boot disk may be preloaded using the initramfs at step 205.33. Figure 2G The storage resources may be connected using the method shown. Then, if an expandable volume exists, the coupled subvolumes or devices are optionally assembled into a volume group at step 205.34 if Logical Volume Management (LVM) is in use. Alternatively, other disk grouping methods may be used to couple the volumes at step 205.34.
[0165] If the fusion file system is in use, the files are combined at step 205.36, and the boot process continues (step 205.46). If using overlayfs in Linux to work around known issues, the following subprocess can be run. A / data directory can be formed in each mounted file system blob, which may be volatile (step 205.37). The new_root directory can then be created at step 205.38, and the overlayfs mounted to the directory at step 205.39. The initramfs then runs exec_root on / new_root (step 205.40).
[0166] If the host is a virtual machine (VM), additional tools such as direct kernel booting may be available. In this case, the hypervisor can connect to the storage resources before booting the VM (step 205.41), or the hypervisor can do so at boot time. A direct kernel boot of the VM can then be performed, along with loading the initramfs (step 205.42). The initramfs is then loaded at step 205.43, and the hypervisor can then connect to the storage resources, which may be remote (step 205.44). To accomplish this, the hypervisor host may require an incoming interface (for example, if InfiniBand needs to connect to an iSER target, it may use pci-passthrough to incoming SR-IOV-based virtual functions, or in some cases, a paravirtualized network interface). These connections are available to the initramfs. If the virtual machine is not already ready, the virtual machine can then connect to the storage resources at step 205.45. The virtual machine may also receive its storage resources through the hypervisor (optionally via paravirtualized storage). The process can be similar for virtual machines that optionally mount fusion file systems and LVM-type disks.
[0167] Figure 2O An example process flow for configuring storage resources from file system blobs or other file groups, as at 205.13, is illustrated. The blobs are collected at step 205.75 and can be copied directly to the storage resource host at 205.73 (if the storage resource host is different from the device storing the file system blobs 232). Once the storage resources are in place, the system state is updated at 205.74 using the storage resource's location and available transport (e.g., iSER, nvmeof, iSCSI, FCoE, Fibre Channel, nfs, nfs over rdma). Some of these blobs may be read-only, in which case the system state remains unchanged and new computing resources or hosts can connect to the read-only storage resources (e.g., when connecting to a base image). In some cases, it may be desirable to place files into a single file system image, as shown at 205.70, to avoid any fused file system overhead. This can be done by mounting the blobs as a fused file system (step 205.71), then copying the blobs into the new file system or repacking them into a single file system (step 205.72), and then optionally copying the new file system image to the appropriate location where the new file system image will appear as a storage resource. Some fused file systems may allow the merge to be done without first mounting the fused file system at step 205.71 and merging them in a single step.
[0168] Figure 2I As shown in the figure Figure 2E Another example template 230 is shown. In this example, the controller can be configured to use Figure 2I A template 230 with a broker configuration tool is shown. According to one exemplary embodiment, the broker configuration tool may include a common API for coupling a new application or service with dependent applications or services. Therefore, template 230 may also include a list of dependencies 244 that may be required to configure the template's services. Template 230 may also include connection rules 245, which may include calls to common APIs for dependencies. Template 230 may also include one or more common APIs 243 and a list of common APIs and versions 242. Common APIs 243 may have methods, functions, scripts, or instructions that may or may not be callable from an application or controller, allowing the controller to configure dependent applications or services so that they can then be coupled to the new application built using template 230. Controllers may communicate with common APIs 243 and / or make API calls to configure the coupling of the new service or application with dependent services or services. Alternatively, instructions may allow the application or service to directly communicate with and / or send calls to common APIs 243 on dependent applications or services. The template 230 connects rules 245 , which is a set of rules and / or instructions that may include API calls for connecting a new service or application with dependent services or applications.
[0169] The system state 220 may also include a list of running services 246. The running service list 246 can be queried by the controller logic 205 to try to satisfy the dependencies 244 from the templates 230. The controller may also include a list 247 of different common APIs that can be used for a particular service / application or a certain type of service / application, and may also include templates containing common APIs. The list may reside in the controller logic 205, the system rules 210, the system state 220, or in a template store accessible to the controller. The controller also maintains a common API index 248 compiled from all existing or loaded templates.
[0170] Figure 2J The controller logic 205 is shown as follows: Figure 2F The processing template 230 is shown but there is an example process flow for step 255 of managing service dependencies by the controller. Figure 2K Shown for Figure 2J255 . At step 255.1, the controller collects dependency list 244 from the template. The controller also collects common API list 243 from the template. (A) At step 255.2, the controller narrows the list of possible dependent applications or services by comparing common API list 243 from the template with common API index 248 and based on the type of application or service seeking to satisfy the dependency. At step 255.3, the controller determines whether system rules 210 specify a method for satisfying the dependency.
[0171] If the determination at step 255.3 is yes, the controller determines whether the dependent service or application is running by querying a list of running templates (step 255.4). If the determination at step 255.4 is no, the service application is run (and / or configured first, then run), which may include the controller logic processing the templates for the dependent service / application (step 255.5). If the dependent service or application is found to be running at step 255.4, the process flow proceeds to step 255.6. At step 255.6, the controller uses the template to couple the new service or application being built to the dependent service or application. In the process of coupling the new service or application with the dependent application / service, the controller completes the template it is processing and runs connection rules 245. Based on the connection rules 245, the controller sends commands to the common API 243 regarding how to satisfy dependencies 244 and / or couple applications / services. The generic API 243 translates the instructions from the controller to connect the new service or application with the dependent application or service, which may include but is not limited to: calling the service's API function, changing the configuration, running scripts, calling other programs. After step 255.6, the process flow proceeds to Figure 2J Step 205.2.
[0172] If step 255.3 determines that system rules 210 do not specify how to satisfy the dependency, the controller will query system status 220 at step 255.7 to see if the appropriate dependency application or service is running. At step 255.8, the controller makes its own determination based on the query as to whether the appropriate dependency application or service is running. If the determination at step 255.8 is negative, the controller may notify the administrator or user to take action (step 255.9). If the determination at step 255.8 is positive, the process flow proceeds to step 255.6, which can be operated as described above. The user may optionally be queried as to whether the new application should connect to a running dependency application. In this case, the controller may couple the new application or service to the dependency application or service at step 255.6 as follows: the controller will complete the template 230 it is processing and run connection rules 245. The controller will then send commands to the common API 243 based on the connection rules 245 regarding how to satisfy the dependency relationship 244. The common API 243 translates instructions from the controller to connect the new service or application with the dependent applications or services.
[0173] A user communicates with the controller 200 through an external user interface or Web UI, or an application through the API application 120 , which may also be incorporated into the controller application or logic 205 .
[0174] Controller 200 communicates with the stack or resources via one or more of a plurality of networks, interconnects, or other connections that the controller can utilize to direct the operation of computing resources, storage resources, and networked resources. Such connections may include: out-of-band management connection 260; in-band management connection 270; SAN connection 280; and optional networking in-band management connection 290.
[0175] Out-of-band management can be used by the controller 200 to detect, configure, and manage components of the system 100 through the controller 200. The out-of-band management connection 260 enables the controller 200 to detect resources that are plugged in and available, but not turned on. The resources can be added to the IT system state 220 when they are plugged in. Out-of-band management can be configured to load a boot image, configure, and monitor resources that are subordinate to the system 100. Out-of-band management can also start a temporary image for diagnosing the operating system. Out-of-band management can be used to change BIOS settings, and console tools can also be used to run commands on the running operating system. The settings can also be changed by the controller using the console, keyboard, and mirror recognition of video signals from physical or virtual monitoring ports on the hardware resource, such as VGA, DVI, or HDMI ports, and / or using APIs provided by out-of-band management, such as Redfish.
[0176] As used herein, out-of-band management may include, but is not limited to, a management system capable of connecting to a resource or node independent of the operating system and motherboard. Out-of-band management connection 260 may include a network, or various types of direct or indirect connections or interconnects. Examples of out-of-band management connection types include, but are not limited to, IPMI, Redfish, SSH, telnet, other management tools, keyboard, video, and mouse (KVM) or KVM over IP, serial console, or USB. Out-of-band management is a tool that can be used over a network to power on and off nodes or resources, monitor temperature and other system data, make BIOS changes and other low-level changes that may be outside the control of the operating system, connect to a console and send commands, and control inputs including, but not limited to, a keyboard, mouse, and monitor. Out-of-band management can be coupled to out-of-band management circuitry in a physical resource. Out-of-band management can connect to a disk image as a disk that can be used to boot the installation media.
[0177] A management network or in-band management connection 270 may allow the controller to gather information about computing resources, storage resources, networking resources, or other resources for direct communication to the operating system on which the resource is running. The storage resources, computing resources, or networking resources may include a management interface that interfaces with connections 260 and / or 270, whereby the resource may communicate with the controller 200 and inform the controller of what is running and what is available for the resource, and receive commands from the controller. As used herein, an in-band management network includes a management network that is capable of communicating with a resource directly to the operating system of the resource. Examples of in-band management connections may include, but are not limited to, SSH, telnet, other management tools, serial console, or USB.
[0178] While out-of-band management is described herein as a physically or virtually separate network from the in-band management network, they can be combined or can operate in conjunction with each other for efficiency purposes as described in more detail herein. Furthermore, and accordingly, out-of-band management and in-band management, or aspects thereof, can communicate through the same port of the controller or be coupled using a combined interconnect. Optionally, one or more of connections 260, 270, 280, 290 can be separate or combined with other such networks and may or may not include the same architecture.
[0179] Furthermore, computing resources, storage resources, and controllers may or may not be coupled to a storage network (SAN) 280 in a manner that enables controller 200 to use the storage network to boot each resource. Controller 200 can send boot images or other templates to individual storage resources or other resources, enabling them to boot from them. The controller can indicate where to boot in such situations. The controller can power on the resource, instructing it where to power on and how to configure itself. Controller 200 instructs the resource on how to boot, what image to use, and where the image is located if it is on another resource. Resources can pre-configure the BIOS. The controller can also or alternatively configure the BIOS through out-of-band management so that it boots from a storage area network. Controller 200 can also be configured to boot an operating system from an ISO and enable resources to copy data to a local disk. The local disk can then be used for booting. The controller can configure other resources, including other controllers, in a manner that enables them to boot. Some resources may include applications that provide computing, storage, or networking functionality. Furthermore, it is possible for the controller to boot a storage resource and then have it be responsible for provisioning the boot image for subsequent resources or services. Storage can also be managed over a different network that is used for another purpose.
[0180] Optionally, one or more of the resources may be coupled to a networking in-band management connection 290. Connection 290 may include one or more types of in-band management as described with respect to in-band management connection 270. Connection 290 may connect the controller to an application network to utilize the network, or to manage the network through the in-band management network.
[0181] Figure 2LThe image 250 is shown. This image 250 can be loaded directly or indirectly (via another resource or database) from a template 230 onto a resource to start the resource or to load an application or service on the resource. Image 250 may include boot files 240 specific to the resource type and hardware. Boot files 240 may include a kernel 241 corresponding to the resource, application, or service to be deployed. Boot files 240 may also include an initrd or similar file system to assist in the boot process. Boot system 240 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 250 may include a file system 251. File system 251 may include a base image 252 and a corresponding file system, a service image 253 and a corresponding file system, and a volatile image 254 and a corresponding file system. The loaded file system and data may vary depending on the resource type and the application or service to be run. Base image 252 may include a base operating system file system. The base operating system may be read-only. Base image 252 may also include basic operating system tools, regardless of the running application. Base image 252 may include basic directories and operating system tools. The service file system 253 may include configuration files and specifications for resources, applications, or services. The volatile file system 254 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file system can be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.
[0182] As described above, the controller 200 may be used to add resources such as computing resources, storage resources, and / or networking resources to the system. Figure 11AAn example method for adding a physical resource, such as a bare metal node, to system 100 is illustrated. A resource, such as a computing resource, storage resource, or networking resource, is inserted into a controller via a network connection 1110. The network connection may include an out-of-band management connection. The controller recognizes that the resource is inserted via the out-of-band management connection 1111. The controller identifies information related to the resource, which may include, but is not limited to, the resource's type, capabilities, and / or attributes 1112. The controller adds the resource and / or information related to the resource to its system state 1113. An image derived from a template is loaded onto a physical component of the system, which may include, but is not limited to, another resource, such as a storage resource, or a resource on the controller 1114. The image includes one or more file systems, which may include configuration files. Such configurations may include BIOS and boot parameters. The controller instructs the physical resource to boot using the image's file system 1115. Additional resources or multiple different types of bare metal or physical resources can be added in this manner using an image of a template, or at least a portion thereof.
[0183] Figure 11B An example method for automatically allocating resources using global system rules and templates, according to an exemplary embodiment, is illustrated. A request is made to a system requiring resource allocation to satisfy the request 1120. A controller understands its resource pool based on its system state database 1121. The controller uses templates to determine the required resources 1122. The controller dispatches the resources and stores the information in the system state 1123. The controller deploys the resources using the template 1124.
[0184] refer to Figure 12 , using the system 100 described herein, illustrates an example method for automatically deploying an application or service. A user or application issues a request for a service 1210. The request is translated to an API application 1220. The API application transmits the request to a controller 1230. The controller interprets the request 1240. The controller takes into account the state of the system and its resources 1250. The controller uses its rules and templates to deploy the service 1260. The controller 1270 sends the request to the resource 1270, deploys the image derived from the template 1280, and updates the IT system state.
[0185] Other more detailed examples of operations such as adding resources, allocating resources, and deploying applications or services are discussed in more detail below.
[0186] Adding computing resources to the system refer to Figure 3A, shows the addition of a computing resource 310 to system 100. When computing resource 310 is added, it is coupled to controller 200 and can be shut down. Note that if computing resource 310 is pre-loaded with an image, alternative steps can be followed, where any network connection can be used to communicate with the resource, start the resource, and add information to the system state. If the computing resource and controller are on the same node, the service running the computing resource is shut down.
[0187] like Figure 3A As shown, computing resource 310 is coupled to the controller via the following networks: out-of-band management connection 260, in-band management connection 270, and optionally, SAN 280. Computing resource 310 is also coupled to one or more application networks 390, where services, application users, and / or clients can communicate with each other. Out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or circuitry within computing resource 310 that is activated when computing resource 310 is plugged in. Device 315 may enable features including, but not limited to, powering on and off the device, attaching to a console and entering commands, monitoring temperature and other computer health-related elements, and setting BIOS settings and other features outside the scope of the operating system. Controller 200 can view computing resource 310 via out-of-band management network 260. The controller can also identify the type of computing resource and determine whether to configure the computing resource using in-band or out-of-band management. Controller logic 205 is configured to review added hardware in out-of-band management 260 or in-band management 270. If a computing resource 310 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource will be configured automatically or through interaction with the user. If the resource is added automatically, the settings will follow the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the computing resource. The controller 200 can query the API application or otherwise request the user or any program in the control stack to confirm that the new resource has been authorized. The authorization process can also be done automatically and securely using cryptography to confirm the legitimacy of the new resource. The controller logic 205 adds the computing resource 310 to the IT system state 220, which includes the switch or network into which the computing resource 310 is plugged.
[0188] If the computing resource is physical, the controller 200 can power on the computing resource via the out-of-band management network 260, and the computing resource 310 can be powered on from an image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, via the SAN 280. The image can be loaded via other network connections or indirectly from another resource. Once powered on, information related to the computing resource 310 received via the in-band management connection 270 can also be collected and added to the IT system state 220. The computing resource 310 can then be added to the storage resource pool, and the computing resource will become a resource managed by the controller 200 and tracked in the IT system state 220.
[0189] If the computing resource is virtualized, the controller 200 can power on the computing resource via the in-band management network 270 or via out-of-band management 260. The computing resource 310 can be booted from an image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, via the SAN 280. The image can be loaded via other network connections or indirectly from another resource. Once booted, information related to the computing resource 310 received via the in-band management connection 270 can also be collected and added to the IT system state 220. The computing resource 310 can then be added to the storage resource pool and become a resource managed by the controller 200 and tracked in the IT system state 220.
[0190] The controller 200 may be able to automatically turn resources on and off according to global system rules and update the IT system status for reasons determined by the IT system user, such as turning off resources to save power, or turning on resources to improve application performance, or any other reason the IT system user may have.
[0191] Figure 3BImage 350 is loaded directly or indirectly (via another resource or database) from template 230 into computing resource 310 to boot the computing resource and / or load an application. Image 350 may include boot files 340 for the resource type and hardware. Boot files 340 may include a kernel 341 corresponding to the resource, application, or service to be deployed. Boot files 340 may also include an initrd or similar file system to assist in the boot process. Boot system 340 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 350 may include a file system 351. File system 351 may include a base image 352 and a corresponding file system, a service image 353 and a corresponding file system, and a volatile image 354 and a corresponding file system. The loaded file system and data may vary depending on the resource type and the application or service to be run. Base image 352 may include a base operating system file system. The base operating system may be read-only. Base image 352 may also include basic operating system tools, regardless of the running application. Base image 352 may include basic directories and operating system tools. The service file system 353 may include configuration files and specifications for resources, applications, or services. The volatile file system 354 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file system can be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.
[0192] Figure 3C 1 illustrates an example process flow for adding a resource, such as computing resource 310, to system 100. Although in this example, the subject resource will be described as computing resource 310, it should be understood that Figure 3C The subject resources of the process flow may also be storage resources 410 and / or networking resources 510. Figure 3C In the example of , the added resource 310 is not on the same node as the controller 200. At step 300.1, the resource 310 is coupled to the controller 200 in a disconnected state. Figure 3C In the example shown, out-of-band management connection 260 is used to connect to resource 310. However, it should be understood that other network connections can be used if desired by the practitioner. At steps 300.2 and 300.3, controller logic 205 examines the system's out-of-band management connections and uses out-of-band management connection 260 to identify and discern the type of resource 310 being added and its configuration. For example, the controller logic may consult the resource's BIOS or other information (such as serial number information) as a reference to obtain type and configuration information.
[0193] At step 300.4, the controller uses global system rules to determine whether a particular resource 310 should be automatically added. If not, the controller waits until its use is authorized (step 300.5). For example, at step 300.4, a user may respond to the query stating that they do not want to use a particular resource 310, or that the particular resource may be automatically put on hold until it is ready for use. If step 300.4 determines that a resource 310 should be automatically added, the controller uses its rules to perform the automatic setup (step 300.6) and proceeds to step 300.7.
[0194] At step 300.7, the controller selects and uses the template 230 associated with the resource to add the resource to the system state 220. In some cases, a template 230 may be specific to a particular resource. However, some templates 230 may cover multiple resource types. For example, some templates 230 may be hardware-spanning. At step 300.8, the controller activates the resource 310 via its out-of-band management connection 260 according to the global system rules 210. At step 300.9, using the global system rules 210, the controller searches for and loads the boot image for the resource from one or more selected templates. The resource 310 is then booted from the image derived from the master template 230 (step 300.10). After booting the resource 310, additional information about the resource 310 may be received from the resource 310 via the in-band management connection 270 (step 300.11). This information may include, for example, firmware version, network card, and any other device to which the resource may be connected. The new information may be added to the system state 220 at step 300.12. The resource 310 may then be considered to have been added to the resource pool and is ready for allocation (step 300.13).
[0195] about Figure 3C If the resource and the controller are on the same node, it should be understood that the service running the resource may be remote from the node. In this case, the controller can use inter-process communication techniques for the resource, such as Unix sockets, a loopback adapter, or other inter-process communication techniques to communicate with the resource. Based on system rules, the controller can install a virtual host, or hypervisor, or container host to run the application using a template known from the controller. The resource application information can then be added to the system state 220, and the resource will be ready for allocation.
[0196] Add storage resources to the system: Figure 4A FIGURE 4 illustrates the addition of storage resources 410 to system 100. In one exemplary embodiment, the following Figure 3CAn example process flow is provided for adding a storage resource 410 to the system 100, where the added storage resource 410 is not on the same node as the controller 200. Additionally, it should be noted that if the storage resource 410 is pre-loaded with an image, alternative steps may be followed, where any network connection may be used to communicate with the storage resource 410, start the storage resource 410, and add the information to the system state 220.
[0197] When a storage resource 410 is added, it is coupled to the controller 200 and can be disconnected. Storage resource 410 is coupled to the controller via the following networks: out-of-band management network 260, in-band management connection 270, SAN 280, and optionally, connection 290. Storage resource 410 may or may not be coupled to one or more application networks 390, where services, application users, and / or clients can communicate with each other. Applications or clients can access the resource's storage directly or indirectly through the application, without accessing the resource's storage via the SAN. Application networks may have built-in storage or be accessible and distinguishable as storage resources in the IT system state. Out-of-band management connection 260 may be coupled to a separate out-of-band management device 415 or circuitry within storage resource 410 that is activated when storage resource 410 is inserted. Device 415 may enable features including, but not limited to, powering on and off the device, attaching to a console and entering commands, monitoring temperature and other computer health-related elements, and setting BIOS settings and other features outside the operating system. The controller 200 can view storage resources 410 via the out-of-band management network 260. The controller can also identify the type of storage resource and determine whether to configure the storage resource using in-band management or out-of-band management. The controller logic 205 is configured to carefully check the added hardware in the out-of-band management 260 or in-band management 270. If a storage resource 410 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource 410 will be configured automatically or through user interaction. If the resource is added automatically, the settings will follow the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the storage resource. The controller 200 can query one or more API applications or otherwise request the user or any program in the control stack to confirm that the new resource is authorized. The authorization process can also be completed automatically and securely using cryptography to confirm the legitimacy of the new resource. The controller logic 205 adds the storage resource 410 to the IT system state 220 , which includes the switch or network into which the storage resource 410 is plugged.
[0198] Controller 200 can power on storage resource 410 via out-of-band management network 260, and storage resource 410 can be powered on from image 450 loaded from template 230, for example, via SAN 280, using global system rules 210 and controller logic 205. The image can also be loaded via other network connections or indirectly from another resource. Once powered on, information related to storage resource 410 received via in-band management connection 270 can also be collected and added to IT system state 220. Computing resource 410 is now added to the storage resource pool, and the storage resource will become a resource managed by controller 200 and tracked in IT system state 220.
[0199] The storage resources may include a single pool of storage resources, or multiple pools of storage resources that can be used or accessed by the IT system independently or simultaneously. When storage resources are added, the storage resources may provide one storage pool, multiple storage pools, a portion of a storage pool, and / or multiple portions of multiple storage pools to the IT system state. The controller and / or storage resource may manage the various storage resources of the pool, or groupings of such resources within the pool. The storage pool may include multiple storage pools running on multiple storage resources. For example, a flash disk or array, a cache disk or array, or a storage pool on a dedicated compute node coupled with a pool on a dedicated storage node is used to optimize both bandwidth and latency.
[0200] Figure 4BThe diagram illustrates an image 450 that is loaded directly or indirectly (via another resource or database) from template 230 into storage resource 410 to boot the storage resource and / or load an application. Image 450 may include boot files 440 specific to the resource type and hardware. Boot files 440 may include a kernel 441 corresponding to the resource, application, or service to be deployed. Boot files 440 may also include an initrd or similar file system to assist in the boot process. Boot system 440 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 450 may include a file system 451. File system 451 may include a base image 452 and a corresponding file system, a service image 453 and a corresponding file system, and a volatile image 454 and a corresponding file system. The loaded file system and data may vary depending on the resource type and the application or service to be run. Base image 452 may include a base operating system file system. The base operating system may be read-only. Base image 452 may also include basic operating system tools, regardless of the running operating system. Base image 452 may include basic directories and operating system tools. The service file system 453 may include configuration files and specifications for resources, applications, or services. The volatile file system 454 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file system can be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.
[0201] Figure 5A An example is shown in which another storage resource, directly attached storage 510 (which may take the form of a node with JBOD or other type of directly attached storage) is coupled to storage resource 410 as an additional storage resource for the system. JBOD is an external disk array that is typically connected to the node providing the storage resource, and the JBOD will be used as Figure 5A , but it should be understood that other types of direct attached storage may also be employed as 510.
[0202] For example, as regards Figure 5AAs described, controller 200 can add storage resources 410 and JBODs 510 to its system. JBODs 510 are coupled to controller 200 via out-of-band management connections 260. Storage resources 410 are coupled to the following networks: out-of-band management connection 260, in-band management connection 270, SAN 280, and optionally connection 290. Storage nodes 410 communicate with the storage of JBODs 510 via SAS or other disk drive architecture 520. JBODs 510 may also include out-of-band management devices 515 that communicate with the controller via out-of-band management connections 260. Through out-of-band management 260, controller 200 can detect JBODs 510 and storage resources 410. Controller 200 can also detect other parameters not controlled by the operating system, such as those described herein with respect to various out-of-band management circuits. Global system rules 210 of controller 200 provide configuration startup rules for starting or activating JBODs and storage nodes that have not yet been added. The order in which storage resources are powered on can be controlled by the controller logic 205 using global rules 220. Based on one set of global system rules 220, the controller may first power on the JBOD 510, and the controller 200 may then use the loaded image 450 to power on the storage resource 410 in a manner similar to that described with respect to FIG. 4. In another set of global system rules, the controller 200 may first power on the storage resource 410, and then power on the JBOD 510. In other global system rules, timing or delays between activating various devices may be specified. The readiness or operational status of various resources may be determined and / or used by the controller 200 for device allocation management via the controller logic 205, global system rules 210, and / or templates 230. The IT system state 220 may be updated by communicating with the storage resource 410. The storage node 410 learns the storage parameters and configuration of the JBOD 510 by accessing the JBOD through the disk fabric 520. The storage resource 410 provides information to the controller 200, which then updates the IT system state 220 with information about the amount of available storage and other attributes. The controller updates the IT system state 220 when the storage resource 410 is activated and the storage resource 410 is identified as part of the pool of storage resources 400 for the system 100. The storage node processes the logic for controlling the JBOD storage resource using the configuration set by the controller 200. For example, the controller can instruct the storage node to configure the JBOD to create a pool from a RAID 10 or other configuration.
[0203] Figure 5BAn example process flow is illustrated for adding a storage resource 410 and a direct-attached storage 510 for the storage resource 410 to the system 100. At step 500.1, the direct-attached storage 510 is coupled to the controller 200 in a disconnected state via the out-of-band management connection 260. At step 500.2, the storage resource 410 is coupled to the controller 200 in a disconnected state via the out-of-band management connection 260 and the in-band management connection 270, while the storage resource 410 is coupled to the direct-attached storage 510 via, for example, a SAS 520, such as a disk drive fabric.
[0204] The controller logic 205 may then examine the out-of-band management connection 260 to detect the storage resource 410 and the directly attached storage 510 (step 500.3). While any network connection may be used, in this example, out-of-band management is used by the controller logic to identify and discern the type of resource being added (in this case, the storage resource 410 and the directly attached storage 510) and its configuration (step 500.4).
[0205] At step 500.5, the controller 200 selects and uses a template 230 for the specific storage type of each type of storage device to add the resources 410 and 510 to the system state 220. At step 500.6, the controller powers on the direct storage and storage nodes in this order through the out-of-band management connection 260 according to the global system rules 210 (which may specify a power-on sequence, a power-on sequence) (500.6). Using the global system rules 210, the controller finds and loads the boot image of the storage resource 410 from the template 230 selected for the storage resource 410, and then boots the storage resource from the image (step 500.7). The storage resource 410 learns the storage parameters and configuration of the direct attached storage 510 by accessing the direct attached storage 510 through the disk architecture 520. The storage resource 410 can then be connected to the storage resource 510 by connecting to the disk architecture 520. The in-band management connection 270 provides additional information about the storage resource 410 and / or the direct-attached storage 510 to the controller (step 500.8). At step 500.9, the controller updates the system state 220 with the information obtained at step 500.8. At step 500.10, the controller processes the direct-attached storage 510 for the storage resource 410 and sets a configuration for how the direct-attached storage should be configured. At step 500.11, a new resource comprising the combination of the storage resource 410 and the direct-attached storage 510 can then be added to the resource pool and made available for allocation within the system.
[0206] According to another aspect of an exemplary embodiment, the controller can use out-of-band management to identify other devices in the stack that may not be participating in the computation or service. For example, such devices may include, but are not limited to, cooling towers / air conditioners, lighting, temperature devices, sound devices, alarm devices, power systems, or any other devices associated with the system.
[0207] Add networking resources to the system: Figure 6A FIGURE 6 illustrates the addition of a networked resource 610 to the system 100. In one exemplary embodiment, the Figure 3C , wherein the added networking resource 610 is not located on the same node as the controller 200. Additionally, it should be noted that if the networking resource 610 is preloaded with an image, alternative steps may be followed, wherein any network connection may be used to communicate with the network resource 610, start the network resource 610, and add information to the system state 220.
[0208] When a networked resource 610 is added, it is coupled to the controller 200 and can be disconnected. The networked resource 610 can be coupled to the controller 200 via the following connections: an out-of-band management connection 260 and / or an in-band management connection 270. The networked resource can optionally be plugged into a SAN 280 and / or a connection 290. The networked resource 610 may also or may not be coupled to one or more application networks 390, where services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 can be coupled to a separate out-of-band management device 615 or circuitry within the networked resource 610 that is activated when the networked resource 610 is added. The device 615 can enable features including, but not limited to, powering on and off the device, attaching to a console and entering commands, monitoring temperature and other computer health-related elements, and setting BIOS settings and other features outside the operating system. The controller 200 can view the networked resource 610 via the out-of-band management connection 260. The controller can also identify the type of networked resource and / or network architecture and determine whether to configure it using in-band or out-of-band management. The controller logic 205 is configured to review the added hardware in out-of-band management 260 or in-band management 270. If a networked resource 610 is detected, the controller logic 205 can use global system rules 220 to determine whether the networked resource 610 will be configured automatically or through user interaction. If the resource is added automatically, the configuration will follow the global system rules 210 within the controller 200. If added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wishes to handle the resource. The controller 200 can query one or more API applications or otherwise request the user or any program in the control stack to confirm that the new resource is authorized. The authorization process can also be performed automatically and securely using cryptography to confirm the legitimacy of the new resource. The controller logic 205 can then add the networked resource 610 to the IT system state 220. For switches that cannot identify themselves to the controller, the user can manually add the switch to the system state.
[0209] If the networking resource is physical, the controller 200 can power on the networking resource 610 via the out-of-band management connection 260, and the networking resource 610 can be powered on from an image 605 loaded from the template 230, for example, via the SAN 280, using the global system rules 210 and the controller logic 205. The image can also be loaded via other network connections or indirectly from other resources. Once powered on, information related to the networking resource 610 received via the in-band management connection 270 can also be collected and added to the IT system state 220. The networking resource 610 can then be added to the storage resource pool, and the networking resource will become a resource managed by the controller 200 and tracked in the IT system state 220. Optionally, some networking resource switches can be controlled via a console port connected to the out-of-band management 260 and can be configured at the driver level, or can have a switch operating system installed via a boot loader, such as through ONIE.
[0210] If the networking resource is virtual, the controller 200 can power on the networking resource through the in-band management network 270 or through out-of-band management 260. The networking resource 610 can be booted from the image 650 loaded from the template 230 through the SAN 280 using the global system rules 210 and the controller logic 205. Once booted, information related to the networking resource 610 received through the in-band management connection 270 can also be collected and added to the IT system state 220. The networking resource 610 can then be added to the storage resource pool, and the networking resource will become a resource managed by the controller 200 and tracked in the IT system state 220.
[0211] The controller 200 can instruct networking resources (whether physical or virtual) to assign, reassign, or move ports to connect to different physical or virtual resources, i.e., connectivity, storage, or compute resources as defined herein. This can be accomplished using technologies including, but not limited to, SDN, InfiniBand zoning, VLANs, and vXLAN. The controller 200 can instruct virtual switches to move or assign virtual interfaces to networks or interconnects that communicate with a particular virtual switch or resources hosting a virtual switch. Some physical or virtual switches can be controlled by an API coupled to the controller.
[0212] The controller 200 may also instruct the computing resources, storage resources, or networking resources to change fabric types if such a change is possible. A port may be configured to switch to a different fabric, such as a hybrid InfiniBand / Ethernet interface.
[0213] The controller 200 can give instructions to networking resources, which may include switches or other networking resources that switch multiple application networks. The switches or network devices may include different fabrics, or, for example, they may be plugged into InfiniBand switches, ROCE switches, and / or other switches that preferably have SDN capabilities and multiple fabrics.
[0214] Figure 6B The diagram illustrates an image 650 loaded directly or indirectly (e.g., via another resource or database) from template 230 into networked resources 610 to boot the networked resources and / or load an application. Image 650 may include boot files 640 specific to the resource type and hardware. Boot files 640 may include a kernel 641 corresponding to the resource, application, or service to be deployed. Boot files 640 may also include an initrd or similar file system to assist in the boot process. Boot system 640 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 650 may include a file system 651. File system 651 may include a base image 652 and a corresponding file system, a service image 653 and a corresponding file system, and a volatile image 654 and a corresponding file system. The loaded file system and data may vary depending on the resource type and the application or service to be run. Base image 652 may include a base operating system file system. The base operating system may be read-only. Base image 652 may also include basic operating system tools, regardless of the running operating system. Base image 652 may include basic directories and operating system tools. The service file system 653 may include configuration files and specifications for resources, applications, or services. The volatile file system 654 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file system can be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.
[0215] Deploy applications or services on resources: Figure 7A The illustrated system 100 includes: a controller 200; physical and virtual computing resources including a first computing node 311, a second computing node 312, and a third computing node 313; storage resources 410; and network resources 610. The resources are shown as described herein with respect to Figures 1 to 6B The described approach sets and adds to the IT system state 220 .
[0216] Although multiple computing nodes are shown in this figure, a single computing node may also be used according to an exemplary embodiment. A computing node may host physical or virtual computing resources, and applications may run on a physical or virtual computing node. Similarly, although a single network provider node and storage node are shown, it is contemplated that multiple resource nodes of these types may or may not be used in a system according to an exemplary embodiment.
[0217] A service or application may be deployed in any system according to an exemplary embodiment. An instance of deploying a service on a computing node may be about Figure 7A , but may similarly be used with different arrangements of the system 100. For example, Figure 7A The controller 200 in FIG. 2 can automatically configure computing resources 310 in the form of computing nodes 311, 312, and 313 according to the global system rules 210. The computing resources can then be added to the IT system state 220. The controller 200 can thus identify the computing resources 311, 312, and 313 (which may or may not be disconnected) and any physical or virtual applications that may be running on the computing resources or nodes. The controller 200 can also automatically configure one or more storage resources 410 and one or more networking resources 610 according to the global system rules 210 and the template 230, and add them to the IT system state 220. The controller 200 can identify storage resources 410 and networking resources 610 that may or may not be disconnected initially.
[0218] Figure 7B An example process for adding a resource to the IT system 100 is shown. At step 700.1, a new physical resource is coupled to the system. At step 700.2, the controller becomes aware of the new resource. The resource may be connected to remote storage (step 700.4). At step 700.3, the controller configures a method for starting the new resource. All connections made to the resource may be recorded in the system state 220 (step 700.5). Figure 3C Provides information such as Figure 7B More details of an exemplary embodiment of the process flow are shown.
[0219] Figure 7C and Figure 7DAn example process flow for deploying an application on multiple computing resources, multiple servers, multiple virtual machines, and / or across multiple sites is shown. This example process differs from a standard template deployment in that the IT system 100 will require components to couple redundant and dependent applications and / or services. The controller logic may process a meta-template at step 700.11, where the meta-template may contain multiple templates 230, a file system blob 232, and other components required to configure a multi-homed service (which may be in the form of other templates 230).
[0220] At step 700.12, the controller logic 205 checks the system state 220 for available resources; however, if insufficient resources exist, the controller logic may reduce the number of redundant services that may be deployed (see 700.16, where the number of redundant services is identified). At step 700.13, the controller logic 205 configures the networking resources and interconnects needed to connect the services together. If a service or application is deployed across multiple sites, the meta-template may include (or the controller logic 205 may configure) services that are optionally configured from a template to enable data synchronization and interoperability across sites (see 700.15).
[0221] At step 700.16, the controller logic 205 may determine the meta-template data, resource availability, and the number of redundant services (if redundant services will exist on multiple hosts) from system rules. At 700.17, coupling to other redundant services and coupling to the motherboard are established. If multiple redundant hosts exist, the controller logic 205 or logic within the template (which may include a binary 234, a daemon 232, or a file system blob containing settings in the boot operating system) may prevent network address and host name conflicts. Optionally, the controller logic will provide network addresses (see 700.18) and register each redundant service in DNS (700.19) and system state 220 (700.18). System state 220 will track redundant services and, if the controller logic 205 notices that a service with conflicting parameters, such as host name (e.g., Software Defined Access (SDA) name), DNS name, network address, etc., is already in system state 220, the controller logic will not allow duplicate registrations.
[0222] Depend on Figure 7DThe configuration routine shown processes one or more templates from the meta-template. The configuration routine processes all redundant services, deploys multi-host or clustered services to multiple hosts, and deploys services to couple the hosts. Any process that can deploy an IT system from a system rule can run the configuration routine. In the case of a multi-host service, the example routine might process the service template as at 700.32, provision storage resources as at 700.33, power the host as at 700.35, and couple the host / compute resources with the storage resources as at 700.36 (and register in the system state 220). (Then repeat for the number of redundant services (700.38), each time registering in the system state 220 (see 700.20) and using controller logic to record information to track the individual services and prevent conflicts (see 700.31).
[0223] Some of the service templates may contain services and tools that can couple multiple host services. Some of these services may be considered dependencies (700.39) and then a coupling routine at 700.40 may be used to couple the services and register the coupling in the system state 220. Additionally, one of the service templates may be a master template and then the dependent service template at 700.39 would be a slave or secondary service; and the coupling routine at 700.40 would connect the services. The routines may be defined in a meta template; for example, for a redundant dns configuration, the coupling routine at 700.40 may include a connection from the dns to the primary dns and configuration for zone transfers along with dnssec. Some services may use physical storage (see 700.34) to improve performance and the physical storage may be loaded with Figure 5B Tools for coupling services can be included in the template itself, and configuration between services can be accomplished using APIs accessible by the controller and / or other hosts in a multi-node application / service.
[0224] The controller 200 can allow a user or controller to determine the appropriate computing backend for an application. The controller 200 can allow a user or controller to optimally place an application on the appropriate physical or virtual computing resources by determining resource usage. When a hypervisor or other computing backend is deployed to a computing node, the hypervisor or other computing backend can report resource utilization statistics back to the controller via the in-band management connection 270. When the controller decides to create an application on a virtual computing resource based on its own logic and global system rules, or based on user input, the controller can automatically select the hypervisor on the optimal host and power on the virtual computing resources on that host.
[0225] For example, the controller 200 deploys an application or service to one or more computing resources using one or more templates 230. Such an application or service may be, for example, a virtual machine running the application or service. In one example, Figure 7A 1. The diagram illustrates the deployment of multiple virtual machines (VMs) on multiple compute nodes, as shown in which the controller 200 can identify that there are multiple compute resources 310 in the form of compute nodes 311, 312, and 313 in its compute resource pool. The compute nodes can, for example, utilize a hypervisor or alternatively be deployed on bare metal, where the use of virtual machines may be undesirable for speed reasons. In this example, the compute resource 310 is loaded with a hypervisor application and has VM (1) 321 and VM (2) 322 configured and deployed on the compute node 311. If, for example, the compute node 311 does not have resources for attaching VMs, or if other resources are preferred for a particular service, the controller 200 can identify that there are no available resources on the compute node 311, or that it is preferred to set up a new VM in a different resource, based on the stack state 220. It may also be identified that the hypervisor is loaded on, for example, the compute resource 312, but not on the resource 313, which may be a bare metal compute node used for other purposes. Thus, based on the requirements of the service or application template being installed, and the status of the system state 220 , the controller may, in this instance, select a compute node 313 for deploying the next required resource, VM ( 3 ) 323 .
[0226] The computing resources of the system may be configured to share storage on the storage resources of the storage nodes.
[0227] The user can request to set up services for the system 100 through the user interface 110 or the application. The services may include but are not limited to: email services; web services; user management services; network providers; LDAP; Dev tools; VOIP; authentication tools; and billing.
[0228] The API application 120 translates the user or application request and sends a message to the controller 200. The controller 200's service template or image 230 is used to identify which resources the service requires. The resources to be used are then identified based on their availability according to the IT system state 220. The controller 200 issues a request to one or more of the compute nodes 311, 312, or 313 for the required computing service, requests storage resources 410 for the required storage resources, and requests network resources 610 for the required networking resources. The IT system state 220 is then updated to identify the resources to be allocated. The service is then installed on the allocated resources using the global system rules 210 based on the service or application's template 230.
[0229] According to an exemplary embodiment, multiple computing nodes may be used for either the same service or different services, and, for example, a storage service and / or network provider pool may be shared among the computing nodes.
[0230] refer to Figure 8A , shows a system 100 in which a controller 200, as well as computing resources 300, storage resources 400, and networking resources 600 are on the same or shared physical hardware, such as a single node. Figures 1 to 10 The various features described and shown in the above can be incorporated into a single node. When the node is driven, the controller image is loaded on the node. The computing resources 300, storage resources 400, and networking resources 600 are configured using the template 230 and the global system rules 210. The controller 200 can be configured to load computing backends 318, 319 as computing resources, which may or may not be added to the node or to one or more different nodes. Such backends 318, 319 may include, but are not limited to: virtualization technology, containers, and multi-tenant processes for creating virtual computing resources, networking resources, and storage resources.
[0231] Applications or services 725, such as web, email, core network services (DHCP, DNS, etc.), and collaboration tools, can be installed on virtual resources on nodes / devices that are shared with the controller 200. These applications or services can be moved to physical or virtual resources independently of the controller 200. Applications can run on virtual machines on a single node.
[0232] Figure 8B A method for scaling from a single-node system to a multi-node system (such as one with Figure 8A Nodes 318 and / or 319 are shown as example process flows. Figure 8A and Figure 8B , we consider an IT system with a controller 200 running on a single server; wherein it is desired to scale out the IT system to a multi-node IT system. Thus, before the scaling out, the IT system is in a single-node state. Figure 8A As shown, the controller 200 runs on a multi-tenant single-node system to drive various IT system management applications and / or resources, which may include but are not limited to: storage resources, computing resources, hypervisor systems and / or container hosts.
[0233] At step 800.2, a new physical resource is coupled to the single-node system by connecting it via out-of-band management connection 260, in-band management connection 270, SAN 280, and / or network 290. For the purposes of this example, this new physical resource may also be referred to as hardware or a host. The controller 200 may detect the new resource on the management network and then query the device. Alternatively, the new device may broadcast a message announcing itself to the controller 200. For example, the new device may be identified by its MAC address, out-of-band management, and / or booting into a preliminary OS and using in-band management to identify the hardware type. In either event, at step 800.3, the new device provides information about its node type and its currently available hardware and software resources to the controller. The controller 200 then learns about the new device and its capabilities.
[0234] At step 800.4, tasks assigned to the system running controller 200 may be assigned to the new host. For example, if the host is pre-loaded with an operating system (such as a storage host operating system or hypervisor), controller 200 may allocate new hardware resources and / or capabilities. The controller may then provide the image and provision the new hardware, or the new hardware may request the image from the controller and configure itself using the methods disclosed above and below. If the new host is unable to host storage resources or virtual computing resources, the new resources may be made available to controller 200. Controller 200 may then move and / or assign existing applications to the new resources, or use the new resources for newly created or subsequently created applications.
[0235] At step 800.5, the IT system can keep its current applications running on the controller or migrate them to the new hardware. If migrating virtual computing resources, VM migration techniques (such as qemu+kvm migration tools) can be used to update the system state and new system rules. The change management techniques discussed below can be used to make these changes reliably and securely. As more applications may be added to the system, the controller can use any of a variety of techniques to determine how to allocate the system's resources, including but not limited to: polling, weighted polling, least utilization, weighted least utilization, prediction techniques with training assistance based on utilization, scheduling techniques, expected capacity techniques, and capacity capping techniques.
[0236] Figure 8CAn example process flow for migrating a storage resource to a new physical storage resource is illustrated. The storage resource may then be mirrored, migrated, or a combination thereof (e.g., storage may be mirrored and then the original storage resource disconnected). At step 820, the storage resource is coupled to the system by having the new storage resource contact the controller or have the controller discover the new storage resource. This can be accomplished using an out-of-band management connection 260, an in-band management connection 270, a SAN network 280, or a flat network that may be used by the application network, or a combination thereof. Under in-band management, the operating system can be pre-booted and the new resource can be connected to the controller.
[0237] At step 822, a new storage target is created on the new storage resource, and this may be recorded in a database at step 824. In one example, the storage target may be created by copying a file. In another example, the storage target may be created by creating a block device and copying data (which may be in the form of one or more file system blobs). In another example, the storage target may be created by mirroring two or more storage resources between block devices (e.g., creating a RAID) and optionally connecting via one or more remote storage transports, including but not limited to ISCSI, ISER, NVMEof, NFS, NFS over RDMA, FC, FCoe, SRP, etc. The database entry at step 824 may include information about computing resources (or other types of resources and / or hosts) connected to the new storage resource, either remotely or locally (if the storage resource is on the same device as the other resources or hosts).
[0238] At step 826, the storage resources are synchronized. For example, the storage may be mirrored. As another example, the storage may be taken offline and synchronized. At step 826, a technology such as RAID 1 (or other types of RAID, but typically RAID 1 or RAID 0, and RAID 110 (mirrored RAID 10) if desired (mdadm, zfs, btrfs, hardware RAID)) may be employed.
[0239] Then, after the database record at step 828, the data from the old storage resource is optionally connected (if the operation occurs later, the database may contain information about the circumstances under which the data was copied, if such data must be recorded). If the storage target is being migrated away from the original host (e.g., as previously described according to Figure 8A and Figure 8BIf the system is moving from a single-node system to a multi-node system and / or a distributed IT system (as described above), the new storage resource may be designated as the primary storage resource by the controller, system state, computing resources, or a combination thereof at step 830. This may be done as part of the removal of the old storage resource. In some cases, the physical or virtual hosts connected to the resource may then need to be updated, and in some cases, the physical or virtual hosts may be shut down (and subsequently restarted) during the transition at step 832 (this may utilize the techniques disclosed herein to drive the physical or virtual hosts).
[0240] Figure 8D An example process flow is shown for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for compute and storage. At step 850, the controller 200 creates a virtual machine, container, and / or process that may be on the new node (e.g., see Figure 8A 8 and 9). At step 852, the old application host may then be shut down. Then, at step 854, the data is copied or synchronized. By first shutting down the host at step 852 and then copying / synchronizing at step 854, the migration will be safer in the case where it involves migrating a VM from a single node. Shutting down will also be beneficial for moving from VMs to physical resources. Step 854 may also be completed via a data pre-synchronization step 862 before shutting down, which can help minimize the associated downtime. In addition, the host may not be shut down as at step 852, in which case the old host remains online until the new host is ready (or the new storage resources are ready). Techniques for avoiding the shutting down step 852 will be discussed in more detail below. At step 854, data may optionally be synchronized unless a hot standby is used to mirror or synchronize the storage resources.
[0241] The new storage resource is now operational and can be recorded in the database at step 856, enabling the controller 200 to connect the new host to the new storage resource at step 858. When migrating from a single node with multiple virtual hosts, this process may need to be repeated for multiple hosts (step 860). The boot order can be determined by the controller logic using the dependencies of the applications (if they are tracked).
[0242] Figure 8EAnother example process flow for scaling from a single node to multiple nodes in a system is shown. At step 870, a new resource is coupled to the single-node system. The controller may have a set of system rules and / or scaling rules for the system (or the controller may derive scaling rules based on the services running, their templates, and the dependencies that services have on each other). At step 872, the controller checks for such rules to facilitate scaling.
[0243] If the new physical resources include storage resources, the storage resources may be moved from a single node or other form of simpler IT system (or the storage resources may be mirrored) at step 874. If the storage resources are moved, the computing resources or runtime resources may be reloaded or restarted at step 876 after the storage resources are moved. In another example, the computing resources may be connected to the mirrored storage resources at step 876 and the computing resources may be kept running, while the old storage resources on the single node system or the hardware resources of the previous system may be disconnected or disabled. For example, a runtime service may be coupled to two mirrored block devices - one on a single node server (e.g., using mdadm raid 1) and the other on the storage resource; and once the data is synchronized, the drive on the single node server may be disconnected. The previous hardware may still include multiple parts of the IT system and the system may be run in a mixed mode with the controller on the same node (step 878). The system may continue to iterate through this migration process until the original node drives only the controller, whereupon the system is distributed (step 880). Additionally, in Figure 8E At each step of the process flow, the controller may update the system state 220 and record any changes to the system in the database (step 882 ).
[0244] refer to Figure 9A , application 910 is installed on resource 900. Resource 900 can be as described herein Figures 1 to 10 The computing resources 310, storage resources 410, or networking resources 610 described herein. Resource 900 may be a physical resource. A physical resource may include a physical machine or a physical IT system component. Resource 900 may be, for example, a physical computing resource, a storage resource, or a networking resource. Resource 900 may be similar to the one described herein. Figures 2A to 10 The other computing resources, networking resources, or storage resources described are coupled together to the controller 200 in the system 100 .
[0245] The resource 900 may be initially powered off. The resource 900 may be coupled to the controller via the following networks: an out-of-band management connection 260, an in-band management connection 270, a SAN 280, and / or a network 290. The resource 900 may also be coupled to one or more application networks 390 where services, application users, and / or clients may communicate with one another. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 915 or circuitry of the resource 900 that is turned on when the resource 900 is plugged in. The device may enable features including, but not limited to: powering on / off the device, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings 195 and other features outside the scope of the operating system.
[0246] The controller 200 can detect resource 900 via the out-of-band management network 260. The controller can also identify the resource type and determine whether to configure the resource using in-band or out-of-band management. The controller logic 205 can be configured to carefully check for attached hardware in either out-of-band management 260 or in-band management 270. If a resource 900 is detected, the controller logic 205 can use global system rules 220 to determine whether the resource 900 will be configured automatically or through user interaction. If the resource is added automatically, the configuration will follow the global system rules 210 within the controller 200. If the resource is added by a user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wishes to handle the computing resource. The controller 200 can query an API application or otherwise request the user or any program in the control stack to confirm that the new resource is authorized. The authorization process can also be performed automatically and securely using cryptography to confirm the legitimacy of the new resource. The resource 900 is then added to the IT system state 220, which includes the switch or network into which the resource 900 is plugged.
[0247] The controller 200 can power on the resource via the out-of-band management network 260. The controller 200 can use the out-of-band management connection 260 to power on the physical resource and configure the BIOS 195. The controller 200 can automatically use the console 190 and select the desired BIOS options. This can be accomplished by the controller 200 reading the console image using image recognition and controlling the console 190 via out-of-band management. The boot status can be determined by using image recognition through the console of the resource 900, or by out-of-band management using a virtual keyboard to query the services listening on the resource, or by querying the services of the application 910. Some applications may have processes that allow the controller 200 to monitor settings in the application 910 or, in some cases, change those settings using in-band management 270.
[0248] On physical resource 900 (or as described herein Figures 1 to 10Applications 910 of the described resources (300, 310, 311, 312, 313, 400, 410, 411, 412, 600, 610) can be booted via SAN 280 or another network using BIOS boot options or configuring remote boot, such as enabling PXE boot or other methods like Flex Boot. Additionally or alternatively, the controller 200 can use out-of-band management 260 and / or in-band management connection 270 to instruct the physical resource 900 to boot the application image in image 950. The controller can configure boot options for the resource or use an existing enabled remote boot method such as PXE boot or Flex Boot. The controller 200 can optionally or alternatively use out-of-band management 260 to boot from an ISO image, configure local disks, and then instruct the resource to boot from one or more local disks 920. The one or more local disks can load boot files. This can be accomplished using out-of-band management 260, image identification, and a virtual keyboard. The resource may also have boot files and / or a boot loader installed. The resources 900 and applications can be started from an image 950 loaded from a template 230 using global system rules 210 and controller logic 205, for example, via SAN 280. The global system rules 220 can specify a startup order. For example, the global system rules 220 may require that the resources 900 be started first, followed by the applications 910. Once the resources 900 are started using the image 950, information related to the resources 900 received via the in-band management connection 270 can also be collected and added to the IT system state 220. The resources 900 can be added to a storage resource pool and become managed by the controller 200 and tracked in the IT system state 220. The applications 910 can also be started using the images 950 or application images 956 loaded on the resources 900 in the order specified by the global system rules 220.
[0249] The controller 200 can configure the networked resources 610 using an out-of-band management connection 260 or another connection to connect the application 910 to the application network 390. The physical resource 900 can be connected to remote storage, such as block storage resources, including, but not limited to, ISER (iSCSI over RDMA), NVMEOF FCoE, FC, or iSCSI, or another storage backend such as SWIFT, GFUSTER, or CEPHFS. The IT system state 220 can be updated using the out-of-band management connection 260 and / or the in-band management connection 270 when a service or application is up and running. The controller 200 can use the out-of-band management connection 260 or the in-band management connection 270 to determine the power state of the physical resource 900, i.e., whether it is powered on or powered off. The controller 200 can use the out-of-band management connection 260 or the in-band management connection 270 to determine whether the service or application is running or in the booted state. The controller can take other actions based on the information it receives and the global system rules 210.
[0250] Figure 9B An image 950 is shown loaded from template 230 directly or indirectly (eg, through another resource or database) to a compute node to launch application 910. Image 950 may include a custom kernel 941 for application 910.
[0251] Image 950 may include boot files 940 for resource types and hardware. Boot files 940 may include a kernel 941 corresponding to the resource, application, or service to be deployed. Boot files 940 may also include an initrd or similar file system to assist in the boot process. Boot system 940 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 950 may include a file system 951. File system 951 may include a base image 952 and a corresponding file system, a service image 953 and a corresponding file system, and a volatile image 954 and a corresponding file system. The loaded file system and data may vary depending on the resource type and the application or service to be run. Base image 952 may include a base operating system file system. The base operating system may be read-only. Base image 952 may also include basic operating system tools that are independent of the running application. Base image 952 may include basic directories and operating system tools. Service file system 953 may include configuration files and specifications for resources, applications, or services. The volatile file system 594 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file system may be mounted as a separate file system using a technology such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.
[0252] Figure 9C This example illustrates installing an application from an NT package, which can be a type of template 230. At step 900.1, the controller determines that a package blob needs to be installed. At step 900.2, the controller creates a storage resource for the blob type (block, file, file system) on the default data store. At step 900.3, the controller connects to the storage resource via a storage transport appropriate for the storage resource type. At step 900.4, the controller copies the package blob to the connected storage resource. The controller then disconnects from the storage resource (step 900.5) and sets the storage resource to read-only (step 900.6). The package blob is then successfully installed (step 900.7).
[0253] In another example, the accompanying Appendix B describes example details on how the system connects computing resources to overlayfs. Such techniques can be used to facilitate the following operations: Figure 9A Install the application on the resource, or Figure 2F Step 205.11 activates computing resources from storage resources.
[0254] Figure 9DThe diagram shows an application 910 deployed on a resource 900. The resource 900 may include a computing node, which may include a virtual computing resource, for example, a hypervisor 920, one or more virtual machines 921, 922 and / or containers. The resource 900 may be related to the Figures 1 to 10A similar approach is described for configuration using an image 950 loaded onto a resource 900. In this example, resource 920 is shown as virtual machines 921 and 922 managed by a hypervisor. Controller 200 can use in-band management 270 to communicate with resource 900 hosting hypervisor 920 to create and configure the resource and allocate appropriate hardware resources, including but not limited to CPU, RAM, GPU, remote GPU (which can use RDMA to remotely connect to another host), network connections, network fabric connections, and / or virtual and physical connections to partitioned and / or segmented networks. Controller 200 can use a virtual console 190 (e.g., including but not limited to SPICE or VNC) and image identification to control resource 900 and hypervisor 920. Additionally or alternatively, controller 200 can use out-of-band management 260 or in-band management connection 270 to instruct hypervisor 920 to launch application image 950 from template 230 using global system rules 210. Image 950 can be stored on controller 200, or controller 200 can move or copy the image to storage resource 410. The boot image for VMs 921, 922 can be stored locally as a file, for example, on image 950, a block device, or on a remote host, and shared using image types such as qcow2 or raw images via file sharing, such as NFS over RDMA / NFS, or the boot image can use remote block devices using iSCSI, ISER, NVMEOF, FC, FCoE. Portions of image 950 can be stored on storage resource 410 or compute node 310. Controller 200 can use global rules and / or templates to appropriately configure networking resources 610 to support the application via out-of-band management connection 260 or another connection. Application 910 on resource 900 can be booted using an image 950 loaded via SAN 280 or another network, using a BIOS boot option, or by allowing hypervisor 920 on resource 900 to connect to block storage resources, such as, but not limited to, ISER (iSCSI over RDMA), NVMEOF FCoE, FC or iSCSI, or another storage backend such as SWIFT, GFUSTER, or CEPHFS. Storage resources can be copied from a template target for a particular storage resource. IT system state 220 can be updated by querying hypervisor 920 for information. In-band management connection 270 can communicate with hypervisor 920 and can be used to determine the power state of the resource, i.e., whether it is powered on or off, or to determine the boot state. Hypervisor 920 can also use a virtual in-band connection 923 to the virtualized application 910 and use hypervisor 920 to implement functions similar to out-of-band management.Depending on whether the service or application is driven or started, this information may indicate whether the service or application is started and running.
[0255] The startup status can be determined through mirror recognition via the resource 900's console 190, or by out-of-band management 260 using a virtual keyboard to query the listening services on the resource, or by querying the services of the application 910 itself. Some applications may have processes that allow the controller 200 to monitor settings in the application 910 or, in some cases, change these settings using in-band management 270. Some applications may reside on virtual resources, and the controller 200 can monitor them by communicating with the hypervisor 920 using in-band management 270 (or out-of-band management 260). The application 910 may not have such a process for monitoring (or such a process may be disabled to save resources) and / or for adding input; in such cases, the controller 200 can use the out-of-band management connection 260 and the mirrored process and / or virtual keyboard to log in to the system to make changes and / or open a management process. Similar to virtual computing resources, a virtual machine console 190 can be used.
[0256] Figure 9E An example process flow for adding a virtual computing resource host to the IT system 100 is shown. At step 900.11, a host capable of being a virtual computing resource is added to the system. The controller may Figure 15B The process flow configures a bare metal server (step 900.12); alternatively, an operating system may be preloaded and / or the host may be preconfigured (step 900.13). The resource is then added to the system state 220 as a pool of virtual computing resources (step 900.14), and the resource becomes accessible to the controller 200 via an API (step 900.15). The API is typically accessed via the in-band management connection 270; however, the in-band management connection 270 can be selectively enabled and / or disabled using a virtual keyboard; and the controller can communicate over the out-of-band management connection 260 using an out-of-band management connection 260 and a virtual keyboard and monitor (step 900.16). At step 900.17, the controller can now utilize the new resource as a virtual computing resource.
[0257] Example multi-controller system: refer to Figure 10 , shows a system 100 having: as herein described Figures 1 to 10The computing resources 300, 310 described herein include multiple physical computing nodes 311, 312, 313; the storage resources 400, 410 described herein are in the form of multiple storage nodes 411, 412 and JBOD 413; multiple controllers 200a, 200b, the multiple controllers 200a, 200b include components 205, 210, 220, 230 ( Figures 1 to 9C ) and configured as the controller 200 described herein; networking resources 600, 610 as described herein, said networking resources 600, 610 comprising a plurality of frameworks 611, 612, 613; and an application network 390.
[0258] Figure 10 The illustration of one possible arrangement of components of the system 100 is for purposes of illustration and does not limit the possible arrangements of components of the system 100 .
[0259] The user interface or application 110 communicates with the API application 120, which communicates with either or both of the controllers 200a or 200b. The controllers 200a, 200b may be coupled to an out-of-band management connection 260, an in-band management connection 270, a SAN 280, or a networked in-band management connection 290. Figures 1 to 9C As depicted, controllers 200a, 200b are coupled to compute nodes 311, 312, 313, storage 411, 412 (including JBOD 413), and networking resources 610 via connections 260, 270, 280, and optionally 290. Application network 390 is coupled to compute nodes 311, 312, 313, storage resources 411, 412, 413, and networking resources 610.
[0260] Controllers 200a, 200b may operate in parallel. Either controller 200a or 200b may initially be as described herein with respect to Figures 1 to 9CControllers 200a and 200b are described as operating as the primary controller 200. Controllers 200a and 200b can be arranged to configure the entire system 100 from a disconnected state. One of the controllers 200a and 200b can also populate the system state 220 from an existing configuration by probing the other controllers via out-of-band connections 260 and in-band connections 270. Either controller 200a or 200b can access or receive resource status and related information from resources or other controllers via one or more connections 260 and 270. The controller or other resource can update the other controller. Thus, when additional controllers are added to the system, the additional controllers can be configured to restore the system 100 to the system state 220. In the event of a failure of one of the controllers or the primary controller, the other controller can be designated as the primary controller. The IT system state 220 may also be capable of being reconstructed from status information available or stored on the resource. For example, an application can be deployed on a computing resource, where the application is configured to create a virtual computing resource at which the system state is stored or replicated. Global system rules 210, system state 220 and templates 230 may also be saved or copied on a resource or resource combination. Thus, if all controllers are forced offline and a new controller is added, the system may be configured to allow the new controller to restore the system state 220.
[0261] Networking resources 610 may include multiple network architectures. Figure 10 As shown, multiple network fabrics may include one or more of the following: an SDN Ethernet switch 611, a ROCE switch 612, an InfiniBand switch 613, or other switches or fabrics 614. A hypervisor system including virtual machines on a compute node can utilize one or more of these fabrics to connect to a physical or virtual switch. This networking arrangement can allow for limiting the physical network, for example, by segmenting the network, for security or other resource optimization purposes.
[0262] System 100 may be implemented as described herein. Figures 1 to 10The controller 200 described in the preceding text automatically configures services or applications. Users can request services for the system 100 through the user interface 110 or an application. These services may include, but are not limited to, email services; web services; user management services; network providers; LDAP; developer tools; VOIP; authentication tools; and billing software. The API application 120 translates the user or application request and sends a message to the controller 200. The controller 200's service template or image 230 is used to identify which resources are required for the service. The required resources are identified based on availability according to the system state 220. The controller 200 requests computing resources 310 or computing nodes 311, 312, or 313 for the required computing service, storage resources 410 for the required storage resources, and network resources 610 for the required networking resources. The system state 220 is then updated to identify the resources to be allocated. The service is then installed on the allocated resources using global system rules 210 according to the service template.
[0263] Enhanced system security: refer to Figure 13A , shows an IT system 100, wherein the system 100 includes a resource 1310, wherein the resource 1310 can be a bare metal or physical resource. Figure 13A Only a single resource 1310 connected to system 100 is shown, but it should be understood that system 100 may include multiple resources 1310. One or more resources 1310 may be or include bare metal cloud nodes. Bare metal cloud nodes may include, but are not limited to, resources connected to an external network 1380 that allow remote access to physical hosts or virtual machines, allow creation of virtual machines, and allow external users to execute code on one or more resources. One or more resources 1310 may be directly or indirectly connected to external network 1380 or application network 390. External network 1380 may be the Internet or one or more other resources not managed by controller 200 or the multiple controllers of IT system 100. External network 1380 may include, but is not limited to, the Internet, one or more Internet connections, one or more resources not managed by a controller, other wide area networks (e.g., Stratcom, peer-to-peer mesh networks, or other external networks that may or may not be publicly accessible), or other networks.
[0264] When a physical resource 1310 is added to the IT system 100a, the physical resource is coupled to the controller 200 and can be disconnected. The resource 1310 is coupled to the controller 200a via one or more of the following networks: an out-of-band management (OOBM) connection 260, an optional in-band management (IBM) connection 270, and an optional SAN connection 280. As used herein, a SAN 280 may or may not include a configuration SAN. A configuration SAN may include a SAN used to drive or configure physical resources. The configuration SAN may be part of the SAN 280 or may be separate from the SAN 280. In-band management may also include a configuration SAN, which may or may not be a SAN 280 as described herein. The configuration SAN may also be disabled, disconnected, or unavailable when the resource is in use. While the OOBM connection 260 is not visible to the OS of the system 100, the IBM connection 270 and / or the configuration SAN may be visible to the OS of the system 100. Figure 13A The controller 200 can be used with reference to Figures 1 to 12 B. The resource 1310 may include internal storage. In some configurations, the controller 200 may populate the storage and may temporarily configure the resource to connect to a SAN to retrieve data and / or information. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or circuitry of the resource 1310 that turns on when the resource 1310 is plugged in. The device 315 may allow features including, but not limited to: powering on / off the device, attaching to a console and entering commands, monitoring temperature and other computer health related elements, and setting BIOS settings and other features outside the scope of the operating system. The controller 200 may view the resource 1310 through the out-of-band management network 260. The controller may also identify the type of resource and identify the configuration of the resource using in-band management or out-of-band management. Figures 13C to 13E Various process flows are illustrated for adding physical resources 1310 to IT system 100a and / or starting or managing system 100 in a manner that enhances system security.
[0265] As used herein with reference to a network, networking resource, network device, and / or networking interface, the term "disable" refers to the action by which such network, networking resource, network device, and / or networking interface is disconnected (manually or automatically), physically disconnected, and / or virtually or in some other manner (e.g., filtered) from a network, i.e., a virtual network (including, but not limited to, VLAN, VXLAN, InfiniBand partitions). The term "disable" also encompasses a one-way or unilateral restriction of operability, such as preventing a resource from sending or writing data to a destination (while still being able to receive or read data from the resource) or preventing a resource from receiving or reading data from a source (while still being able to send or write data to the destination). Such a network, networking resource, network device, and / or networking interface can be disconnected from an additional network, virtual network, or from the resource's coupling, while remaining connected to a previously connected network, virtual network, or resource. Furthermore, such a networking resource or device can switch its coupling from one network, virtual network, or resource to another.
[0266] As used herein with reference to a network, networking resource, network device, and / or networking interface, the term "enable" refers to the action by which such network, networking resource, network device, and / or networking interface is actuated (manually or automatically), physically connected, and / or virtually or in some other manner connected to a network, i.e., a virtual network (including, but not limited to, VLANs, VXLANs, InfiniBand partitions). Such a network, networking resource, network device, and / or networking interface can be connected to an additional network, virtual network, or resource while already connected to another system component. Furthermore, such a networking resource or device can switch its connection from one network, virtual network, or resource to another. The term "enable" also encompasses unidirectional or one-sided permission of operability, such as allowing a resource to send, write, or receive data to a destination (while still having the ability to restrict data from a certain source), or allowing a resource to send, receive, or read data from a certain source (while still having the ability to restrict data from the destination).
[0267] The controller logic 205 is configured to carefully check the added hardware in the out-of-band management connection 260 or the in-band management connection 270 and / or the configuration SAN 280. If a resource 1310 is detected, the controller logic 205 can use the global system rules 220 to determine whether the resource will be configured automatically or through interaction with the user. If the resource is added automatically, the settings will follow the global system rules 210 within the controller 200. If the resource is added by the user, the global system rules 210 within the controller 200 may require the user to confirm the addition of the resource and how the user wants to handle the resource 1310. The controller 200 can query the API application or otherwise request the user or any program in the control stack to confirm that the new resource has been authorized. The authorization process can also be completed automatically and securely using cryptography to confirm the legitimacy of the new resource. The controller logic 205 then adds the resource 1310 to the IT system state 220, which includes the switch or network into which the resource 1310 is inserted.
[0268] If the resource is physical, the controller 200 can power on the resource via the out-of-band management network 260, and the resource 1310 can be powered on using the global system rules 210 and the controller logic 205, for example, via the SAN 280, from an image 350 loaded from the template 230. The image can be loaded via another network connection or indirectly from another resource. Once powered on, information related to the resource 1310 can also be collected and added to the IT system state 220. This can be accomplished through in-band management and / or configuring a SAN or out-of-band management connection. The resource 1310 can be powered on using the global system rules 210 and the controller logic 205, for example, via the SAN 280, from an image 350 loaded from the template 230. The image can be loaded via another network connection or indirectly from another resource. Once powered on, information related to the computing resource 310 received via the in-band management connection 270 can also be collected and added to the IT system state 220. The resource 1310 can then be added to the storage resource pool, and the resource will become a resource managed by the controller 200 and tracked in the IT system state 220 .
[0269] In-band management and / or configuration SANs can be used by controllers 200 to set up, manage, use, or communicate with resources 1310 and to execute any commands or tasks. However, the in-band management connection 270 can optionally be configured by controllers 200 to be turned off or disabled at any time or during the setup, management, use, or operation of the system 100 or controllers 200. In-band management can also be configured to be turned on or enabled at any time or during the setup, management, use, or operation of the system 100 or controllers 200. Optionally, controllers 200 can controllably or switchably disconnect resources 1310 from in-band management connections 270 connected to one or more controllers 200. This disconnection or disconnectability can be physical, such as using an automated physical switch or some other type of switch to cut off the resource's in-band management connection and / or configuration SAN from the network. Disconnection can be accomplished, for example, by a network switch shutting off power to the port connected to the in-band management 270 and / or configuration SAN 280 of the resource 1310. This disconnection or partial disconnection can also be accomplished using software-defined networking, or can be physically filtered out of the controller using software-defined networking. This disconnection can be accomplished by the controller through in-band management or out-of-band management. According to an exemplary embodiment, at any point before, during, or after a resource 1310 is added to the IT system, the resource 1310 can be disconnected from the in-band management connection 270 in response to a selective control instruction from the controller 200.
[0270] Using software-defined networking, the in-band management connection 270 and / or configuration SAN 280 may or may not retain certain functionality. The in-band management 270 and / or configuration SAN 280 can function as a limited connection for communications to or from the controller 200 or to other resources. Connection 270 can be restricted to prevent attackers from pivoting to the controller 200, other networks, or other resources. The system can be configured to prevent devices such as the controller 200 and resources 1310 from communicating openly to prevent them from compromising the resources 1310. For example, the in-band management 270 and / or configuration SAN 280 might only allow the in-band management and / or configuration SAN to transmit data but not receive anything, either through software-defined networking or hardware modifications (such as electronic restrictions). The in-band management and / or configuration SAN can be configured as a one-way write component or one-way write connection from the controller 200 to the resources 1310, either physically or using software-defined networking that only allows writes from the controller to the resources. The one-way write nature of the connection can also be controlled and enabled or disabled based on security expectations and different phases or times of system operation. The system may also be configured such that writing or communication from the resource to the controller is limited to, for example, conveying logs or alarms. Interfaces may also be moved to other networks or added and removed from networks using techniques including, but not limited to, software defined networking, VLANS, VXLANS and / or InfiniBand partitioning. For example, an interface may be connected to a setup network, removed from that network and moved to a network used for runtime. Communication from the controller to the resource may be cut off or restricted such that the controller may not be physically able to respond to any data sent from the resource 1310. According to one example, once the resource 1310 is added and started, the in-band management 270 may be turned off or filtered out, either physically or using software defined networking. In-band management may be configured such that it is able to send data to another resource dedicated to log management.
[0271] In-band management can be turned on and off using out-of-band management or software-defined networking. In the event that in-band management is disconnected, the daemon may not need to be running and keyboard functionality can be used to re-enable in-band management.
[0272] Additionally, optionally, resource 1310 may not have an in-band management connection, and the resource may be managed via out-of-band management.
[0273] Alternatively or in addition, out-of-band management can be used to manipulate various aspects of the system, including, but not limited to, using the keyboard, a virtual keyboard, a disk mount console, attaching a virtual disk, changing BIOS settings, changing boot parameters and other aspects of the system, running existing scripts that may exist on a bootable image or installation CD, or other features of out-of-band management that allow the controller 200 and resources 1310 to communicate with or without exposing the operating system running on the resource 1310. For example, the controller 200 can use such tools to send commands via out-of-band management 260. The controller 200 can also use image identification to assist in controlling the resource 1310. Thus, using an out-of-band management connection, the system can prevent or avoid unintended manipulation of the resources connected to the system via the out-of-band management connection. The out-of-band management connection can also be configured as a one-way communication system during operation of the system or at selected times during operation of the system.
[0274] Additionally, if desired by a practitioner, the out-of-band management connection 260 may also be selectively controlled by the controller 200 in the same manner as the in-band management connection.
[0275] The controller 200 may be able to automatically turn resources on and off according to global system rules and update the IT system status for reasons determined by the IT system user, such as turning off resources to save power, turning on resources to improve application performance, or any other reasons the IT system user may have. The controller may also be able to turn the configuration SAN, in-band management connections, and out-of-band management connections on and off, or designate such connections as one-way write connections at any time during system operation and for various security purposes (e.g., disabling the in-band management connection 270 or the configuration SAN 280 when the resource 1310 is connected to the external network 1380 or the internal network 390). One-way in-band management may also be used, for example, to monitor the health of the system, monitoring logs and information that may be visible to the operating system.
[0276] The resources 1310 may also be coupled to one or more internal networks 390, such as application networks, where services, application users, and / or clients may communicate with each other. Such application networks 390 may also be connected to or capable of connecting to external networks 1380. Figures 2A to 12 In an exemplary embodiment of B, in-band management can be disconnected, can be disconnected from the resource or application network 390, or can provide one-way writes from the controller to provide additional security in the event that the resource or application network is connected to an external network, or in the event that the resource is connected to an application network that is not connected to an external network.
[0277] Figure 13AThe IT system 100 can be configured similarly as Figure 3B In the illustrated IT system 100, image 350 can be loaded directly or indirectly (via another resource or database) from template 230 into resource 1310 to boot the computing resource and / or load an application. Image 350 may include boot files 340 specific to the resource type and hardware. Boot files 340 may include a kernel 341 corresponding to the resource, application, or service to be deployed. Boot files 340 may also include an initrd or similar file system to assist in the boot process. Boot system 340 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 350 may include a file system 351. File system 351 may include a base image 352 and its corresponding file system, a service image 353 and its corresponding file system, and a volatile image 354 and its corresponding file system. The loaded file system and data may vary depending on the resource type and the application or service to be run. Base image 352 may include a base operating system file system. The base operating system may be read-only. Base image 352 may also include basic operating system tools, regardless of the running operating system. Base image 352 may include basic directories and operating system tools. The service file system 353 may include configuration files and specifications for resources, applications, or services. The volatile file system 354 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file system can be mounted as a separate file system using technologies such as overlayFS to allow some read-only file systems and some read-write file systems, thereby reducing the amount of duplicated data for applications.
[0278] Figure 13B A plurality of resources 1310 are shown, each of which includes one or more hypervisor systems 1311 hosting or including one or more virtual machines. The controller 200a is coupled to the resources 1310, each of which includes bare metal resources. Figure 13B As shown and described, resources 1310 are each coupled to controller 200a. According to exemplary embodiments herein, in-band management connection 270, configuration SAN 280, and / or out-of-band management connection 260 may be similar to those described with respect to FIG. Figure 13AAs described above, one or more virtual machines or hypervisors may be compromised or damaged. In conventional systems, other virtual machines on other hypervisors may be compromised. This may occur, for example, due to a hypervisor vulnerability operating within a virtual machine. For example, a transfer may occur from a compromised hypervisor to controller 200a, and from a compromised controller 200a to another hypervisor coupled to controller 200a. For example, the transfer may occur between the compromised hypervisor and the target hypervisor using a network connected to both. Figure 13B The illustrated arrangement of in-band management 270, configuration SAN 280, or out-of-band management 260 of the controller 200a and the resource 1310, wherein either or both of the controller and the resource can be selectively controlled to disable the in-band connection (or configuration SAN) and / or the out-of-band connection in a given link between the controller 200a and the resource 1310, can prevent a compromised virtual machine being used from escaping one hypervisor system and being transferred to other resources.
[0279] The above about Figures 1 to 12 The described in-band management connection 270 and out-of-band management connection 260 may also be used with respect to Figure 13A and Figure 13B Configure in a similar manner as described.
[0280] Figure 13C 1. The diagram illustrates an example process flow for adding physical resources, such as bare metal nodes, to the system 100, or managing the physical resources. Figure 13A and Figure 13B As shown or as about Figures 1 to 12 The resources 1310 are shown connected to the controller of the system 100 .
[0281] After connecting to the instance of the resource, the external network and / or application network is disabled at step 1370. As described above, any of a variety of techniques can be used for such disabling. For example, before using an in-band management connection or configuring a SAN to set up the system, add the resource, test the system, update the system, or perform other tasks or commands, such as with respect to Figure 13A and Figure 13B As described, components of the system 100 (or only those components that are vulnerable) are disabled, disconnected from any external network or application network, or filtered out of any external network or application network.
[0282] Following step 1370, the in-band management connection and / or configuration SAN is then enabled at step 1371. The combination of steps 1370 and 1371 thus isolates the resource from the external network and / or application network while the in-band management and / or SAN connection is active. Commands may then be run on the resource via the in-band management connection under the control of the controller 200 (see step 1372). For example, setup and configuration steps (such as, but not limited to, those described herein with respect to the in-band management and / or configuration SAN) may then be performed at step 1372 using the in-band management and / or configuration SAN. Figures 1 to 13B Alternatively or in addition, in-band management and / or configuration of the SAN may be used at step 1372 to perform other tasks including, but not limited to: operating, updating, or managing the system (which may include, but is not limited to, any change management or system updates), testing, updating, transferring data, collecting information about performance and health (including, but not limited to, errors, cpu usage, network usage, file system information, and storage usage), and collecting logs and other data that may be used as described herein. Figures 1 to 13B Other commands for managing system 100 are described in .
[0283] After adding the resources, setting up the system and / or executing such tasks or commands, the system may be started at step 1373 as described herein. Figure 13A and Figure 13B The description describes disabling the in-band management connection 270 and / or configuration SAN 280 between the resource and the controller or other components of the system in one or more directions. This disabling may be achieved by disconnecting, filtering, etc. as described above. After step 1373, the connection to the external network and / or application network may be restored at step 1374. For example, the controller may inform the networked resource that the resource 1310 is allowed to connect to the application network or the Internet. The same steps may be followed when testing or updating the system, i.e., the in-band management connection or the external network and / or application network may be disconnected or filtered before enabling the in-band management connection or connecting the in-band management connection (unidirectionally or bidirectionally) to the resource. Thus, steps 1373 and 1374 operate together to prevent the resource from connecting to the controller via the in-band management connection and / or configuration SAN when the resource is connected to the external network and / or application network.
[0284] Out-of-band management can be used to manage systems or resources, set up systems or resources, configure, start, or add systems or resources. Out-of-band management, in any embodiment used herein, can use a virtual keyboard to send commands to the machine to change settings before booting, and can also send commands to the operating system by typing into the virtual keyboard; if the machine is not logged in, out-of-band management can use the virtual keyboard to enter a username and password and can use image recognition to verify the login, verify the commands it entered, and check to see if the commands were executed. If the physical resource only has a graphical console, a virtual mouse can also be used and image recognition will allow out-of-band management to make changes.
[0285] Figure 13D Another example process flow for adding physical resources such as bare metal nodes to the system 100 or managing the physical resources is shown in FIG. 1380. Figure 13A and Figure 13B As shown or Figures 1 to 12The resource shown is connected to a system or resource. A disk can be virtually connected by providing access to a disk image (e.g., an ISO image) via out-of-band management with the help of a controller (see step 1381). The resource or system can then be booted from the disk image (step 1382), and files can be copied from the disk image to a bootable disk (see step 1383). This can also be used to boot a system where resources have been configured in this manner using out-of-band management. This can also be used to configure and / or boot multiple resources that may be coupled together (including but not limited to coupling using networked resources), regardless of whether these multiple resources also include a controller or constitute a system. Thus, a virtual disk can be used to allow a controller to connect the disk image to the resource as if the virtual disk were attached to the resource. Out-of-band management can also be used to send files to the resource. At step 1383, _Data can be copied from the virtual disk to a local disk. The disk image may contain files that the resource can copy and use in its operations. Files can be copied or used by a scheduled program or instructions from out-of-band management. The controller can log in to the resource using a virtual keyboard via out-of-band management and enter commands to copy files from the virtual disk to its own disk or other storage accessible to the resource. At step 1384, the system or resource is configured to boot by setting BIOS, EFI, or boot order settings, thereby booting the system or resource from a bootable disk. Boot configuration can use an EFI manager, such as efibootmgr, in the operating system. The EFI manager can be run directly via out-of-band management or by including it in an installer script (for example, when the resource boots, the resource automatically runs a script using efibootmgr). Additionally, boot options and any other BIOS changes can be set using out-of-band management tools such as Supermicro Boot Manager using boot order commands or by uploading a BIOS configuration (such as XML BIOS configuration supported by Supermicro Update Manager). The BIOS can also be configured from the console using keyboard and image recognition to set appropriate BIOS settings, including the boot order. The installer can be run on the loaded pre-configured image. The configuration can be tested by viewing the screen and using image recognition. After configuration, the resource may then be enabled (eg, driven, started, connected to an application network, or a combination thereof) (step 1385 ). Figure 13E Another example process flow for adding physical resources, such as bare metal nodes, to the system 100, or managing the physical resources, in this case using PXE, Flex Boot, or similar network boot, is shown. At step 1390, the management system 100 may be connected to the system 100 via (1) an in-band management connection 270 and / or SAN and (2) an out-of-band management connection 260. Figure 13A and Figure 13B As shown or as about Figures 1 to 12The resources 1310 shown are connected to the controller of the system 100. The external network and / or application network connection may then be disabled (e.g., physically, using SDN, or virtually, filtered out or disconnected) in whole or in part at step 1391 (similar to what was discussed above with respect to step 1370). For example, before using an in-band management connection or SAN to set up the system, add the resources, test the system, update the system, or perform other tasks or commands, as described with respect to Figure 13A and Figure 13B As described, components of the system 100 (or only those components that are vulnerable) are disabled, disconnected from any external network or application network, or filtered out of any external network or application network.
[0286] At step 1392, the type of resource is determined. For example, an out-of-band management tool may be used, or information about the resource may be gathered from the mac address by connecting a disk image (e.g., an ISO image) to the resource as if the disk were attached to the resource, to temporarily boot an operating system that has tools that can be used to discern resource information. Then, at step 1393, the resource is configured or discerned as pre-configured for PXE or flex boot, etc. Thereafter, at step 1394, the resource is driven to perform a PXE, Flex boot, or similar boot (or in a situation where it is temporarily booted and then driven again). Then, at step 1395, the resource is booted from an in-band management connection or SAN, or the resource is booted from the in-band management connection or SAN. At step 1396, the resource is booted from the reference Figure 13D The data is copied to the disk accessible by the resource in a similar manner as described in step 1383 of the embodiment. Then, in step 1397, the data is copied to the disk accessible by the resource in a similar manner as described in the embodiment of ... Figure 13D The resource is configured to boot from one or more disks in a manner similar to that described in step 1384. If the resource is identified as preconfigured for PXE, flexboot, etc., files may be copied at any of steps 1393 through 1396. If in-band management is enabled, it may be disabled at step 1398, and the application network or external network may be reconnected or enabled at step 1399.
[0287] Furthermore, it should be understood that techniques other than OOBM can also be used to remotely enable (such as power on) resources and verify that the resources have been powered on. For example, the system can prompt the user to press the power button and manually inform the controller that the system has been powered on (or using a keyboard / console connection to the controller). In addition, once the system has been powered on, the system can check with IBM for the controller, and the controller can log in and tell the system to restart (for example, through methods such as ssh, telnet, or another method implemented over the network). For example, the controller can be introduced via ssh and send a restart command. If PXE is being used and OOBM is not present, then in any case, the system should have a way to remotely instruct the resource to power on, or tell the user to manually drive the resource.
[0288] Deploy the controller and / or environment: In an exemplary embodiment, controllers may be deployed within a system from an originating controller 200 (where such an originating controller 200 may be referred to as a "master controller"). Thus, the master controller may set up a system or environment that may be an isolated or isolatable IT system or environment.
[0289] As described herein, an environment refers to a collection of resources within a computer system that are capable of interoperating with each other. A computer system may include multiple environments within it; however, this need not be the case. One or more resources of an environment may include one or more instances, applications, or sub-applications running on the environment. Further, an environment may include one or more environments or sub-environments. An environment may or may not include a controller, and the environment may operate one or more applications. Such resources of an environment may include, for example, networking resources, computing resources, storage resources, and / or application networks for running a particular environment, including applications in the environment. Therefore, it should be understood that an environment may provide the functionality of one or more applications. In some examples, the environment described herein may or can be physically or virtually separated from other environments. Additionally, in other examples, the environment may have network connections to other environments, where such connections may be disabled or enabled as needed.
[0290] In addition, the main controller can set up, deploy, and / or manage one or more additional controllers in various environments or in separate systems. Such additional controllers can be or become independent of the main controller. Even if independent or quasi-independent of the main controller, such additional controllers can obtain instructions from the main controller (or a separate monitor or environment using a monitoring application) or send information to the main controller at various times during operation. The environments can be configured for security purposes (for example, by enabling the environments to be isolated from each other and / or from the main controller) and / or for various management purposes. One environment can be connected to an external network, while another related environment can be connected or not connected, or can be connected or not connected to the external network.
[0291] A master controller can manage environments or applications, regardless of whether they are separate systems and regardless of whether they include controllers or sub-controllers. The master controller can also manage shared storage for global configuration files or other data. The master controller can also resolve global system rules (e.g., system rules 210) or a subset thereof to different controllers based on their capabilities. Each new controller (referred to as a "sub-controller") can receive new configuration rules, which may be a subset of the master controller's configuration rules. The subset of global configuration rules deployed to a controller may depend on or correspond to the type of IT system being set up. The master controller can set up or deploy new controllers or separate IT systems, which may later be permanently separated from the master controller, for example, for shipping, distribution, or other reasons. Global configuration rules (or a subset thereof) can define the framework for setting up applications or sub-applications in various environments and how they can interact with each other. Such applications or environments can run on sub-controllers that include a subset of the global configuration rules deployed by the master controller. In some examples, such applications or environments can be managed by the master controller. However, in other examples, such applications or environments are not managed by the primary controller. If a new controller is being generated from the primary controller to manage the application or environment, the application dependencies can be viewed across multiple applications to facilitate control by the new controller.
[0292] Thus, in one exemplary embodiment, a system may include a primary controller configured to deploy another controller, or an IT system including such another controller. Such an implemented system may be configured to be completely disconnected from the primary controller. Once independent, such a system may be configured to operate as a standalone system; or the system may be controlled or monitored by another controller (or an environment with an application), such as the primary controller, at various discrete or continuous times during operation.
[0293] Figure 14AAn example system is shown in which a master controller 1401 has deployed controllers 1401a and 1401b on different systems 1400a and 1400b, respectively (where 1400a and 1400b may be referred to as subsystems; however, it should be understood that subsystems 1400a and 1400b may also be used as environments). Master controller 1401 may be configured in a manner similar to controller 200 discussed above. Thus, the master controller may include controller logic 205, global system rules 210, system state 220, and templates 230.
[0294] Systems 1400a and 1400b include controllers 1401a, 1401b, respectively, coupled to resources 1420a, 1420b. The master controller 1401 may be coupled to one or more other controllers, such as controller 1401a of subsystem 1400a and controller 1401b of subsystem 1400b. The global rules 210 of the master controller 1400 may include rules that may manage and control the other controllers. The master controller 1401 may use such global rules 210, along with controller logic 205, system state 220, and templates 230, to control the controllers 1401a, 1401b in a manner consistent with the present disclosure. Figures 1 to 13E Subsystems 1400a, 1400b are set up, provisioned, and deployed in a similar manner as described.
[0295] For example, the main controller 1401 may load global rules 210 (or a subset thereof) as rules 1410a and 1410b onto subsystems 1400a and 1400b, respectively, in the following manner: Global rules 210 (or a subset thereof) dictate the operation of controllers 1401a and 1401b and their subsystems 1400a and 1400b, respectively. Each controller 1401a and 1401b may have rules 1410a and 1410b that may be the same or different subsets of global rules 210. For example, which subset of global rules 210 is supplied to a given subsystem may depend on the type of subsystem being deployed. The controller 1401 may also load or direct data to be loaded to system resources 1420a and 1420b or controllers 1401a and 1401b.
[0296] The master controller 1401 may be connected to the other controllers 1401a, 1401b via one or more in-band management connections 270 and / or one or more out-of-band management connections 260 or SAN connections 280, which may be connected at various stages of deployment or management in the manner described herein, for example, with reference to FIG. 13A to 13EUsing the selective enabling and disabling of the in-band management connection 270 or the out-of-band management connection 260, the subsystems 1400a, 1400b can be deployed in a manner in which the subsystems 1400a, 1400b may have no knowledge (or have limited, controlled, or restricted knowledge) of the main system 100 or the controller 1401, or about each other at various times.
[0297] In an exemplary embodiment, the main controller 1401 can operate a centralized IT system having local controllers 1401a, 1401b deployed and configured by the main controller 1401, so that the main controller 1401 can deploy and / or run multiple IT systems. Such IT systems may be independent or not independent of each other. The main controller 1401 can set monitoring as a separate application that is isolated or isolated from the IT system it has created. A separate console for monitoring can be provided to be connected between the main controller and one or more local controllers and / or to be connected between environments that can be selectively enabled or disabled. The controller 1401 can deploy, for example, an isolation system for various purposes, including but not limited to: a commercial system, a manufacturing system with data storage, a data center, and other different functional nodes, each of which has a different controller to prevent operational interruption or damage. This isolation can be complete or permanent, or it can be quasi-isolation, for example, temporary, time or task related, communication direction related, or other parameter related. For example, the main controller 1401 can be configured to provide instructions to the system, which may or may not be limited to certain predefined scenarios, while a subsystem may have limited or no ability to communicate with the main controller. Consequently, such a subsystem may be unable to compromise the main controller 1401. The main controller 1401 and subcontrollers 1401a, 1401b can be isolated from each other, for example, by disabling in-band management 270, by one-way writes, and / or by restricting communication to out-of-band management 260, as described herein (with specific examples discussed below). For example, if a vulnerability occurs, one or more controllers can disable in-band management connections 270 with one or more other controllers to prevent the vulnerability or access from spreading. System segments can be shut down or isolated.
[0298] The subsystems 1400a, 1400b may also share resources with or connect to another environment or system through in-band management 270 or out-of-band management 260.
[0299] Figure 14B and Figure 14C is an example flow illustrating possible steps for supplying a sub-controller from a main controller.
[0300] exist Figure 14BIn step 1460, the master controller supplies or sets up a resource, such as resource 1420a or 1420b. At step 1461, the master controller supplies or sets up a sub-controller. The master controller may perform steps 1460 and 1461 using the techniques discussed above for setting up resources within the system. Figure 14B Step 1460 is shown as being performed before step 1461, but it should be understood that this need not be the case. Using its system rules 210, the master controller 1401 can determine which resources are needed and locate the resources on the system or network. The master controller can set up or deploy the sub-controller by loading the system rules 210 onto the system (or by providing instructions to the sub-controller on how to set up and obtain its own system rules) at step 1461. These instructions may include, but are not limited to: configuring resources, configuring applications, global system rules for creating IT systems run by the sub-controllers, instructions for reconnecting to the master controller to collect new or changed rules, and instructions for disconnecting from the application network to make room for a new production environment. After deploying the resources, at step 1463, the master controller can then dispatch the resources to the sub-controllers via updates to the system rules 210 and / or system status 220.
[0301] Figure 14C An alternative process flow for deployment is shown. Figure 14C In the example of , the master controller deploys the sub-controller at step 1470 (which may be done as described with respect to step 1461). Then, at step 1475, the sub-controller uses a Figure 3C and Figure 7B The resources are deployed using the techniques shown.
[0302] Figure 15A An example system is shown in which a main controller 1501 for system 100 generates environments 1502, 1503, and 1504. Environment 1502 includes resources 1522, environment 1503 includes resources 1523, and environment 1504 includes resources 1524. In addition, environments 1502, 1503, and 1504 can share access to a shared resource pool 1525. Such shared resources can include, but are not limited to, shared data sets, APIs, or applications running that need to communicate with each other.
[0303] exist Figure 15AIn the example of , each environment 1502, 1503, 1504 shares a master controller 1501. The global system rules 210 of the master controller 1501 may include rules for deploying and managing the environments. Resources 1522, 1523, and / or 1524 may be required by their respective environments 1501, 1502, 1503 to manage one or more applications. Configuration rules for such applications may be implemented by the master controller (or by local controllers in the environment, if any) to define how each such environment operates and interacts with other applications and environments. The master controller 1401 may use the global rules 210 in conjunction with the controller logic 205, system state 220, and templates 230 to configure the environment in a manner consistent with the present invention. Figures 1 to 14C The described resources and system deployment are similarly configured, provisioned, and deployed in an environment. If the environment includes a local controller, the master controller 1501 may load the global rules 210 (or a subset thereof) onto the local controller or associated storage in a manner that defines the operation of the environment.
[0304] The controller 1501 may deploy and configure the resources 1522, 1523, 1524 and / or shared resources 1525 of the environments 1502, 1503, 1504 using configuration rules and system rules 210. The controller 1501 may also monitor the environments, or configure the resources 1522, 1523, 1524 (or shared resources 1525) to allow monitoring of the respective environments 1502, 1503, 1504. Such monitoring may be performed using a connection to a separate monitoring console that may be enabled or disabled, or may be performed through the master controller. The master controller 1501 may be connected to one or more of the environments 1502, 1503, 1504 via one or more in-band management connections 270 and / or one or more out-of-band management connections 260 or SAN connections 280, which may be used at various stages of deployment or management as described herein. 13A to 13E and Figure 14A Using the enabling and disabling of the in-band management connection 270 or the out-of-band management connection 260 or the SAN connection 280, the environments 1502, 1503, 1504 can be deployed in a manner that may have no knowledge, or limited or controlled knowledge, or no connection, or limited or controlled connection, with respect to each other or to the host system 100 or the controller 1501 at various times.
[0305] The environment may include one or more resources that are coupled or interact with other resources, or coupled to an external network 1580 that is connected to an external, external environment. The environment can be physical or non-physical. Non-physical in this context means that the environments share one or more of the same physical hosts, but are virtually separate from each other. The environment and system can be deployed on the same, similar but different, or different hardware. In some examples, environments 1502, 1503, 1504 can be valid copies of each other; however, in other examples, environments 1502, 1503, 1504 can provide different functionality from each other. As an example, the resources of an environment can be servers.
[0306] Placing systems and resources in separate environments or subsystems according to the techniques described herein can allow for the isolation of applications for security and / or performance reasons. Separate environments can also mitigate the impact of compromised resources. For example, one environment might contain sensitive data and be configured for less internet exposure, while another environment might host internet-facing applications.
[0307] Figure 15B As shown in the figure Figure 15A The controller shown sets up an example process flow for an environment. In this example, the system can be tasked with creating and setting up a new environment. This can be triggered by a user request or by system rules that execute when a specific task or series of tasks is performed. 17A to 18B The example of a specific change management task or a specific series of tasks is illustrated where the system creates a new environment. However, there may be a variety of situations where a controller may create and set up a new environment.
[0308] Therefore, reference Figure 15B In setting up a new environment, the controller selects the environment rules (step 1500.1). Based on the environment rules, using the global system rules 210 and templates 230, the controller searches for resources for the environment (step 1500.2). The rules may have a hierarchy of preferred resource selections that is continually reviewed until the resources required for the environment are found. At step 1500.3, the controller uses, for example, Figure 3C or Figure 7BThe techniques described in [1500.2] allocate the resources found in step 1500.2 to the environment. The controller then configures the system's networked resources for the new environment to ensure compatibility and effective connectivity between the new environment and other system components (step 1500.4). At step 1500.5, the system status is updated to indicate that each resource is enabled and each template is processed. The controller then configures and implements the integration and interoperability of the environment's resources and drives any applications to deploy the new environment (step 1500.6). At step 1500.7, the system status is again updated to indicate that the environment has become available.
[0309] Figure 15C As shown in the figure Figure 15A The controller shown in the figure sets up an example process flow for multiple environments. When setting up multiple environments, you can use Figure 15B However, it will be appreciated that the environment can be set up in parallel using the techniques described in Figure 15C The description sets the environment in an ordered sequence or in a continuous manner. Figure 15C At step 1500.10, the controller sets up and deploys the first new environment (this can be similar to Figure 15B The controller 210 is executed as described in step 1500.1 of the preceding text. Different environment rules may exist for different types of environments and how they interact with each other. At step 1500.11, the controller selects environment rules for the new environment. At step 1500.12, the controller searches for resources based on a preference order defined by system rules 210. At step 1500.13, the controller allocates the resources found in step 1500.12 to the next environment. These environments may or may not share resources. At step 1500.14, the controller uses system rules 210 to configure the system's networked resources for the next environment and between environments with dependencies. At step 1500.15, the system status is updated to indicate that each resource is enabled, each template is processed, and networked resources are configured, including environment dependencies. The controller then configures and implements the integration and interoperability of resources between the next environment and the environments, and drives any applications to deploy the new environment (step 1500.16). At step 1500.17, the system status is updated to indicate that the next environment has become available.
[0310] One-way communication to support monitoring: Figure 16AThe diagram shows an exemplary embodiment in which a first controller 1601 operates as a master controller to set up one or more controllers, such as 1601a, 1601b, and / or 1601b. The master controller 1601 can be used to generate multiple clouds, hosts, systems, and / or applications as environments 1602, 1603, 1604 that may or may not be dependent on each other in their operation using the techniques discussed above with respect to controllers, such as controllers 200 / 1401 / 1501. Figure 16A As shown, IT systems, environments, clouds, and / or any one or more combinations thereof may be generated as environments 1602, 1603, and 1604. Environment 1602 includes a second controller 1601a, environment 1603 includes a third controller 1601b, and environment 1604 includes a fourth controller 1601c. Environments 1602, 1603, and 1604 may each further include one or more resources 1642, 1643, and 1644, respectively. Resources may include one or more applications 1642, 1643, and 1644 running thereon. These applications may be connected to allocated resources, whether shared or not. These or other applications may run on one or more shared resources in the Internet or in a pool 1660, which may also include shared applications or application networks. Applications may provide services to users or to one or more of the environments or clouds. Environments 1602, 1603, and 1604 may share resources or databases and / or may include or use resources from pool 1660 that are specifically allocated to a particular environment. Various components of the system, including the main controller 1601 and / or one or more environments, may also be able to connect to an application network or external network 1615 such as the Internet.
[0311] Between any resource, environment, or controller and another resource, environment, controller, or external connection, there may be a connection that can be configured as described in this document about 13A to 13E 1604. For example, any resource, controller, environment, or external connection may be disabled or disconnected from controller 1601, environment 1602, environment 1603, and / or environment 1604, resource, or application via in-band management connection 270, out-of-band management connection 270, or SAN connection 280, or by being physically disconnected. As an example, in order to protect controller 1601, an in-band management connection 270 between controller 1601 and any of environments 1602, 1603, 1604 may be disabled. As another example, such one or more in-band management connections 270 may be selectively disabled or enabled during operation of environments 1602, 1603, 1604. In addition to the information herein regarding 13A to 13EFor the security purposes discussed, disabling environments 1602, 1603, 1604 or disconnecting the master controller 1601 from the environments can allow the master controller 1601 to transform the environments 1602, 1603, 1604 into clouds that can then be separated from the master controller 1601 or other clouds or environments. In this sense, the controller 1601 is configured to generate multiple clouds, hosts, or systems.
[0312] By disabling or disconnecting the elements described herein, a user can be allowed limited access to an environment through the main controller 1601 for specific purposes. For example, a developer can be provided access to a development environment. As another example, an application administrator can be restricted to a specific application or application network. As another example, a log can be reviewed by the main controller 1601 to collect data without exposing the main controller itself to damage from the environment or controller it generates.
[0313] After the main controller 1601 sets up the environment 1602, the environment 1602 can be disconnected from the main controller 1601, so that the environment 1602 can operate independently of the main controller 1601 and / or can be selectively monitored and maintained by the main controller 1601 or other applications associated with the environment 1602 or run by the environment.
[0314] An environment such as environment 1602 may be coupled to a user interface or console 1640 that allows a purchaser or user to access environment 1602. Environment 1602 may host a user console as an application. Environment 1602 may be accessed remotely by a user. Each environment 1602, 1603, 1604 may be accessed by a common or separate user interface or console.
[0315] Figure 16B An example system is shown in which environments 1602, 1603, and 1604 can be configured to write to another environment 1641, where logs can be viewed, for example, using a console (which can be any console that can be directly or indirectly connected to environment 1641). In this way, environment 1641 can act as a log server, with one or more of environments 1602, 1603, and 1604 writing events to the log server. Master controller 1601 can then access log server 1641 to monitor events on environments 1602, 1603, and 1604 without maintaining a direct connection to such environments 1602, 1603, and 1604 as described below. Environment 1641 can also be selectively disconnected from master controller 1601 and can be configured to only read from other environments 1602, 1603, and 1604.
[0316] The main controller 1601 may be configured to Figure 16CAs shown, a computer disconnected from any of its environments 1602 , 1603 , 1604 is also able to monitor some or all of its environments 1602 , 1603 , 1604 . Figure 16C The in-band management connection 270 between the master controller 1601 and the environments 1602, 1603, 1604 is shown to have been disconnected, which may help protect the master controller 1601 in the event that the environments 1602, 1603, 1604 are compromised. Figure 16C As shown, even if the in-band connection 270 between the master controller 1601 and environment 1602 is disconnected, the out-of-band connection 260 between the master controller 1601 and an environment such as 1602 can still be maintained. Furthermore, environment 1641 can have a selectively enabled or disabled connection to the master controller 1601. The master controller 1601 can configure monitoring as a separate application within environment 1641 that is isolated or isolated from environments 1602, 1603, and 1604. The master controller 1601 can use one-way communication for monitoring. For example, logs can be provided from environments 1602, 1603, and 1604 to environment 1641 via one-way communication. By using this one-way write and connection between the environment 1641 and the main controller 1601, even though there is no in-band connection 270 between the main controller 1601 and the environments 1602, 1603, 1604, the main controller 1601 can collect data and monitor the environments 1602, 1603, 1604 through the environment 1641, thereby reducing the risk of the environments 1602, 1603, 1604 damaging the main controller 1601. The access may be filtered or controlled and / or the access may be independent of the Internet. For example, Figure 16D As shown, if the in-band connection 270 between the main controller 1601 and the environment 1602 is connected, the main controller 1601 can control the network switch 1650 to disconnect the environment 1602 from the external network 1615, such as the Internet. When the environment 1602 is connected to the main controller 1601 via the in-band connection 270, the disconnection of the environment 1602 from the external network 1615 can provide enhanced security for the main controller 1601.
[0317] Therefore, it should be understood that 16B to 16DThe exemplary embodiment shows how a master controller can securely monitor environments 1602, 1603, 1604 while minimizing exposure to said environments 1602, 1603, 1604. Thus, master controller 1601 can disconnect itself (or at least disconnect itself from the in-band link) from environments 1602, 1603, 1604 while still maintaining a mechanism for monitoring said environments via the log server of environment 1641, to which environments 1602, 1603, 1604 may have one-way write privileges. Accordingly, if, in the process of reviewing the logs of environment 1641, master controller 1601 discovers that environment 1602 may have been compromised by malware, master controller 1601 can use SDN tools to isolate said environment 1602 so that only out-of-band connections 260 exist (e.g., see Figure 16C ). Additionally, controller 1601 may send notifications about possible problems to the administrator of environment 1602. The controller may also isolate the damaged environment 1602 by selectively disabling any connections (e.g., in-band management connections 270) between the damaged environment and any of the other environments 1603 and 1604. In another example, master controller 1601 may discover from logs that resources within environment 1603 are running too hot. This may allow the master controller to intervene in an application or service in environment 1603 and migrate it to a different environment (whether pre-existing or newly created).
[0318] The controller 1601 can also set up one or more similar systems according to the purchaser or user's request. Figure 16E As shown, a purchase application 1650 may be provided, for example, in a console or other location, allowing a purchaser to purchase a cloud, host, system environment, or application or request that the cloud, host, system environment, or application be provisioned for the purchaser. Purchase application 1650 may instruct controller 1601 to provision environment 1602. Environment 1602 may include controller 1601a, which will deploy or build an IT system, for example, by allocating or dispatching resources to environment 1602.
[0319] Figure 16FUser interfaces 1632, 1633, and 1634 are shown, usable when environments 1602, 1603, and 1604 each operate as a cloud and may or may not include a controller. Each of user interfaces 1632, 1633, and 1634 (corresponding to environments 1602, 1603, and 1604, respectively) can be connected via a main controller 1601, which manages the connection between the user interfaces and the environments. Alternatively or in addition, interface 1640a (which may take the form of a console) can be directly coupled to environment 1602, interface 1640b (which may take the form of a console) can be directly coupled to environment 1603, and interface 1640c (which may take the form of a console) can be directly coupled to environment 1604. A user can use one or more of these interfaces to access an environment or cloud regardless of whether the connection to main controller 1601 is detached, disconnected, or disabled.
[0320] Clone and back up systems for change management support: Some of the environments 1602, 1603, 1604 may be clones of typical setup software used by developers. The environments may also be clones of current working environments as a measure; for example, cloning an environment in another data center in a different location to reduce latency due to location.
[0321] It will therefore be appreciated that the master controller setting up systems and resources in separate environments or subsystems may allow for cloning or backing up portions of the IT system. This may be used in testing and change management as described herein. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, and / or other changes. The global rules may include a subset that includes backup rules that may be used in change management as described in various examples herein. It will therefore be appreciated that backup rules (examples of which are described elsewhere herein) may be used in change management. References Figure 21A -J describes in more detail an example of a system that implements backup rules.
[0322] According to an exemplary embodiment, an IT system or controller as described herein may be configured to clone one or more environments. The new or cloned environment may or may not include the same resources as the original environment. For example, it may be desirable or necessary to use a completely different combination of physical and / or virtual resources in a new or newly cloned environment. It may be desirable to clone an environment to a different location or time zone where usage can be managed optimally. It may be desirable to clone an environment to a virtual environment. In the process of cloning an environment, the global system rules 210 and global templates 230 of a controller or master controller may include information on how to configure and / or run various types of hardware. Configuration rules within the system rules 210 may indicate the placement and use of resources so that resources and applications are more optimal given the specific available resources.
[0323] The master controller structure provides its ability to set up systems and resources in separate environments or subsystems, provides a structure for cloning environments, provides a structure for creating development environments, and / or provides a structure for deploying a set of standardized applications and / or resources. Such applications or resources may include, for example, but are not limited to those that can be used for developing and / or running applications, or backing up parts of an IT system or restoring from backups of the IT system and other disaster recovery applications (e.g., a LAMP (apache, mysql, php) stack, a system containing a server running a web front end and react / redux and resources running node.js, as well as a mongo database and other standardized "stacks"). Sometimes, the master controller may deploy an environment that is a clone of another environment, and the master controller may obtain configuration rules from a subset of the configuration rules used to create the original environment.
[0324] According to an exemplary embodiment, change management of a system or a subset of a system can be accomplished by cloning one or more environments or the configuration rules or a subset of configuration rules of such environments. Changes may be required to, for example, make changes to code, configuration rules, security patches, templates, hardware changes, add / remove components and dependent applications, and other changes.
[0325] According to an exemplary embodiment, such changes to the system can be automated to avoid errors caused by directly manually entering changes. Changes can be tested by the user in a development environment before being automatically implemented on the live system. According to an exemplary embodiment, the live production environment can be cloned by using a controller to automatically power on, provision and / or configure an environment configured using the same configuration rules as the production environment. The cloned environment can be run and operated (while the backup environment can preferably be kept as a contingency in case the changes need to be undone). This can be done using the controller as described above with reference to Figures 1 to 16FThe description is done by creating, configuring, and / or provisioning a new system or environment using system rules 210, templates 230, and / or system state 220. The new environment can be used as a development environment to test changes that will later be implemented in a production environment. The controller can generate the infrastructure of such an environment from a software-defined structure into the development environment.
[0326] A production environment as defined herein means an environment used to operate the system, as opposed to an environment used only for development and testing, ie, a development environment.
[0327] When cloning a production environment, the infrastructure or cloned development environment is configured by the controller according to global system rules 210 and generated as a production environment. Changes to the development environment may be made to code, templates 230 (changing existing templates or changes related to the creation of new templates), security and / or application or infrastructure configuration. When new changes implemented in the development environment are ready through development and / or testing as needed, the system automatically makes the changes to the development environment before running or deploying it as a production environment. The new system rules 210 are then uploaded to the environment's controller and / or master controller, which applies the system rule changes to the specific environment. The system state 220 is updated in the controller, and additional or revised templates 230 may be implemented. Thus, the development environment and / or master controller can maintain complete system knowledge of the infrastructure and have the ability to recreate complete system knowledge of the infrastructure. As used herein, complete system knowledge may include, but is not limited to, system knowledge of resource status, resource availability, and system configuration. Complete system knowledge may be gathered by the controller from system rules 210, system status 220, and / or by querying resources using one or more in-band management connections 270, one or more out-of-band management connections 260, and / or one or more SAN connections 280. In particular, resources may be queried to determine resource, network, or application utilization, configuration status, or availability.
[0328] The cloned infrastructure or environment may be software defined via system rules 210; however, this need not be the case. The cloned infrastructure or environment may typically include or not include a front end or user interface, and one or more allocated resources, which may or may not include computing resources, networking resources, storage resources and / or application networking resources. The environment may or may not be arranged as a front end, middleware and database. A service or development environment may be started using the system rules 210 of the production environment. In particular, for cloning purposes, the infrastructure or environment allocated for use by the controller may be software defined. Therefore, the environment may be deployable via system rules 210 and cloneable by similar means. The cloned or development environment may be automatically set up by a local controller or a master controller using system rules 210 before or when a change is required.
[0329] Before the development environment is isolated from the production environment, data of the production environment may be written to a read-only data store so that the data will be used by the development environment during development and testing.
[0330] While the production environment is online, users or clients can make changes to the development environment and test those changes. While testing development and changes in the development environment, data in the data store can be modified. For volatile or writable systems, hot synchronization of the development environment's data with the production environment's data can also be used after the development environment is set up or deployed. Required changes to the system, applications, and / or environment can be made and tested in the development environment. Required changes can then be made to the system rules 210 script to create a new version for the environment or for the entire system and master controller.
[0331] According to another exemplary embodiment, the newly developed environment can then be automatically implemented as the new production environment while maintaining the previous production environment or keeping it in a fully functional state, so that recovery of the production environment in an earlier state is possible without losing a large amount of data. The development environment is then started using the new configuration rules within system rules 210, and the database is synchronized with the production database and switched to a writeable database. The original production database can then be switched to a read-only database. The previous production environment is preserved intact as a copy of the previous production environment for a desired period of time in case a recovery back to the previous production environment is needed.
[0332] An environment may be configured as a single server or instance that may include or contain physical and / or virtual hosts, networks, and other resources. In another exemplary embodiment, an environment may be a plurality of servers that contain physical and / or virtual hosts, networks, and other resources. For example, there may be a plurality of servers that form a load-balanced internet-facing application; and the servers may be connected to a plurality of API / middleware applications (the API / middleware applications may be hosted on one or more servers). The database of the environment may include one or more databases to which the API communicates queries in the environment. The environment may be constructed from system rules 210 in a static or volatile form. An environment or instance may be virtual or physical, or a combination of each.
[0333] Configuration rules for an application or a system within system rules 210 may specify various compute backends (e.g., bare metal, AMD epyc server, Intel Haswell on qemu / kvm), and may include rules for how to run an application or service on a new compute backend. Thus, for example, if there is a situation where the availability of resources for testing is reduced, an application may be virtualized.
[0334] Using and following the examples described in this article, the test environment can be deployed on virtual resources where the original environment uses physical resources. Figures 1 to 18B The controller described, and as further described herein, a system or environment can be cloned from a physical environment to an environment that may include or exclude virtual resources in whole or in part.
[0335] Figure 17A The illustrated system 100 includes an exemplary embodiment of a controller 1701 and one or more environments, such as 1702, 1703, and 1704. The system 100 can be a static system, i.e., a system in which active user data does not constantly change the state of the system or frequently manipulate data, such as a system that only hosts static web pages. The system can be coupled to a user (or application) interface 110.
[0336] The controller 1701 may be configured in a similar manner to the controllers 200 / 1401 / 1501 / 1601 described herein and may similarly include global system rules 210, controller logic 205, templates 230, and system state elements 220. The controller 1701 may be configured as described herein with reference to Figures 14A to 16F The controller 1701 may be coupled to one or more other controllers or environments in the manner described. The global rules 210 of the controller 1701 may include rules that may manage and control other controllers and / or environments. Such global rules 210, controller logic 205, system state 220, and templates 230 may be used to communicate with the controller 1701 in accordance with the present disclosure. Figures 1 to 16F Systems or environments are set up, provisioned, and deployed in a similar manner as described. Each environment can be configured using a subset of the global system rules 210 that define the environment's operations, including defining the environment's operations with respect to other environments.
[0337] The global system rules 210 may also include change management rules 1711. The change management rules 1711 include a set of rules and / or instructions that may be used when changes may be needed to the system 100, the global system rules 210, and / or the controller logic 205. The change management rules 1711 may be configured to allow users or developers to develop changes, test the changes in a test environment, and then implement the changes by automatically converting the changes to a new set of configuration rules within the system rules 210. The change management rules 1711 may be a subset of the global system rules 210 (e.g., Figure 17A), or the change management rules may be separate from the global system rules 210. The change management rules may utilize a subset of the global system rules 210. For example, the global system rules 210 may include a subset of environment creation rules configured to create a new environment. Change management rules 1711 may be configured to set up and use the system or environment configured and set up by the controller 1701 to copy and clone some or all aspects of the system 100. Change management rules 1711 may be configured to permit testing of new proposed changes to the system before implementation by using a clone of the system for testing and implementation. Change management rules 1711 may include or utilize the backup rules described below.
[0338] like Figure 17A The illustrated clone 1705 may include the rules, logic, applications, and / or resources of a specific environment or portion of system 100. Clone 1705 may include hardware similar to or different from system 100 and may or may not use virtual resources. Clone 1705 may be configured as an application. Clone 1705 may be configured and configured using configuration rules within system rules 210 of system 100 or controller 1701. Clone 1705 may or may not include a controller. Clone 1705 may include allocated networking resources, computing resources, application networks, and / or data storage resources, as described in more detail above. Such resources may be allocated under the control of controller 1701 using change management rules 1711. Clone 1705 may be coupled to a user interface that allows a user to make changes to clone 1705. The user interface may be the same as or different from user interface 110 of system 100. Clone 1705 may be used for the entire system 100, or for a portion of system 100, such as one or more environments and / or controllers. Clone 1705 may or may not be a complete copy of system 100. The clone 1705 can be coupled to the system 100 via an in-band management connection 270, an out-of-band management connection 260, and / or a SAN connection 280, which can be selectively enabled and / or completely disabled, and / or converted to a unidirectional read and / or write connection. Thus, the connection to the data in the clone environment 1705 can be changed to make the clone data read-only while the clone environment 1705 is isolated from the production environment during testing, or before the clone environment 1705 is ready to operate as the new production environment. For example, if the clone 1705 has a data connection to the environment 1702, this data connection can be made read-only for isolation purposes.
[0339] Optional backup 1706 may or may not be used for the entire system, or a portion of the system, such as one or more environments and / or controllers. Individual services may also be backed up when performing change management functions. For example, refer to Figure 21A-J, a backup of the service can be performed using backup rules as described below. Backup 1706 may include networking resources, computing resources, application networks, and / or data storage resources as described in more detail above. Backup 1706 may or may not include a controller. Backup 1706 may be a complete copy of system 100. Backup 1706 may be provided as an application or using hardware similar to or different from system 100. Backup 1706 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or disabled entirely, and / or converted to unidirectional read and / or write connections.
[0340] Figure 17B Diagram for use Figure 17A Example process flow for system change management for a clone and backup system. At step 1785, a user or management application initiates a change to the system. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, hardware changes, adding / removing components and / or dependent applications, and other changes. At step 1786, the controller 1701 initiates a change to the system. Figures 14A to 16F The environment is set up in the manner described to become a clone environment 1705 (where the clone environment may have its own new controller, or the clone environment may use the same controller as the original environment).
[0341] At step 1787, the controller 1701 may use the global rules 210 including the change management rules 1711 to clone all or part of one or more environments of the system (e.g., the "production environment") to the clone environment 1705 (e.g., where the clone environment 1705 may be used as the "development environment"). The backup rules 2104 may be used to extract data, where the extracted data may be used later for reference purposes. Figure 21A 1706 (with or without a controller) to back up the system and copy the templates 230, controller logic 205, and global rules 210.
[0342] After clone 1705 is generated from the production environment, it can be used as a development environment, where changes can be made to the clone's code, configuration rules, security patches, templates, and other changes. At step 1789, changes to the development environment can be tested before implementation. During testing, clone 1706 can be isolated from the production environment (system 100) or other components of the system. This can be achieved by causing controller 1701 to selectively disable one or more of the connections between system 100 and clone 1706 (e.g., by disabling in-band management connection 270 and / or application network connections). At step 1790, a determination is made as to whether the modified development environment is ready. If step 1709 determines that the development environment is not yet ready (a decision typically made by the developer), the process flow returns to step 1789 to make further changes to clone environment 1705. If step 1790 determines that the development environment is ready, the development and production environments can be switched at step 1791. That is, the controller transitions the development environment 1705 to the new production environment, and the previous production environment may remain unchanged until the transition to the development environment / new production environment is complete and satisfactory.
[0343] Figure 18A Another exemplary embodiment of a system 100 that may be configured and used in system change management is shown. Figure 18A In the example of , the system 100 includes a controller 1801 and one or more environments 1802, 1803, 1804, 1805. The system is shown with a clone environment 1807 and a backup system 1808. Backup and data recovery can be performed using the backup rules described elsewhere in this document. Figure 21A -J further describes an example of management of the backup system.
[0344] Controller 1801 is configured in a similar manner to controllers 200 / 1401 / 1501 / 1601 / 1701 described herein and may include global system rules 210, controller logic 205, templates 230, and system state 220 elements. Controller 1801 may be configured as described herein with reference to Figures 14A to 16F The controller 1801 may be coupled to one or more other controllers or environments in the manner described. The global rules 210 of the controller 1801 may include rules that may manage and control other controllers and / or environments. Such global rules 210, controller logic 205, system state 220, and templates 230 may be used to communicate with the controller 1801 in accordance with the present disclosure. Figures 1 to 17B Systems or environments are set up, provisioned, and deployed in a similar manner as described. Each environment can be configured using a subset of the global rules 210 that define the environment's operations (including defining the environment's operations with respect to other environments).
[0345] The global rules 210 may also include change management rules 1811. The change management rules 1811 may include a set of rules and / or instructions that may be used when changes to the system, global rules, and / or logic may be required. The change management rules may be configured to allow users or developers to develop changes, test the changes in a test environment, and then implement the changes by automatically converting the changes to a new set of configuration rules within the system rules 210. The change management rules 1811 may be a subset of the global system rules 210 (e.g., Figure 18A ), or the change management rules may be separate from the global system rules 210. The change management rules 1711 may use a subset of the global system rules 210. For example, the global system rules 210 may include a subset of environment creation rules configured to create a new environment. The change management rules 1811 may be configured to set up and use the system or environment set up and deployed by the controller 1801 to copy and clone some or all aspects of the system 100. The change management rules 1811 may be configured to permit testing of new changes proposed to the system before implementation by testing and implementation using a clone of the system. The change management rules 1811 may include or use backup rules as described elsewhere herein. The backup rules may use backup rules 2104 to extract data, where the extracted data may later be used as described in reference Figure 21A -J to restore the backup rules.
[0346] like Figure 18A The illustrated cloning environment 1807 may include a controller 1807a having rules, controller logic, templates, and system state data, and allocated resources 1820 that may be allocated to one or more environments and configured according to the controller's 1801 global system rules 210 and change management rules 1811. The backup system 1808 also includes a controller 1808a having rules, controller logic, templates, and system state data, and allocated resources 1821 that may be allocated to one or more environments and configured according to the controller's 1801 global system rules 210 and change management rules 1811. The system may be coupled to a user (or application) interface 110 or another user interface.
[0347] Clone environment 1807 may include the rules, logic, templates, system state, applications, and / or resources for a specific environment or portion of a system. Clone 1807 may include hardware similar to or different from system 100, and may or may not use virtual resources. Clone 1807 may be configured as an application. Clone 1807 may be set up and configured using configuration rules within system rules 210 of system 100 or the environment's controller 1801. Clone 1807 may or may not include a controller, and may share a controller with the production environment. Clone 1807 may include allocated networking resources, computing resources, application networks, and / or data storage resources, as described in more detail above. Such resources may be allocated under the control of controller 1801 using change management rules 1811. Clone 1807 may be coupled to a user interface that allows a user to make changes to clone 1807. This user interface may be the same as or different from user interface 110 of system 100.
[0348] The clone 1807 may be used for the entire system, or a portion of the system, such as one or more environments and / or controllers. In an exemplary embodiment, the clone 1807 may include a hot standby data resource 1820a coupled to the data resource 1820 of the environment 1802. The hot standby data resource 1820a may be used when setting up the clone 1807 and in testing changes. For example, as described herein with respect to Figure 18B As described, hot standby data resource 1820a can be selectively disconnected or isolated from storage resource 1820 during change management. Clone 1807 may or may not be a complete copy of system 100. Clone 1807 can be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which can be selectively enabled and / or completely disabled, and / or converted to a one-way read and / or write connection. Thus, the connection to volatile data in clone environment 1807 can be changed to make the clone data read-only while the clone environment 1807 is isolated from the production environment during testing, or until the clone environment is ready to operate as the new production environment.
[0349] When switching from an old production environment to a new one, controller 1801 can instruct the front-end, load balancer, or other applications or resources to point to the new production environment. Thus, when a change occurs, users, applications, resources, and / or other connections can be redirected. This can be accomplished, for example, through a variety of methods, including but not limited to: changing IP / IPOIb address lists, InfiniBand GUIDs, DNS servers, InfiniBand partitioning / OpenSM configurations; or changing software-defined networking (SDN) configurations, which can be accomplished by sending instructions to networking resources. The front-end, load balancer, or other applications and / or resources can point to systems, environments, and / or other applications, including but not limited to: databases, middleware, and / or other back-ends. Therefore, load balancers can be used in change management to switch from an old production environment to a new one.
[0350] Clones 1807 and backups 1808 can be set up and used in various aspects of managing system changes. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, hardware changes, adding / removing components and / or dependent applications, and other changes. Backups 1808 can be used for the entire system, or a portion of the system, such as one or more environments and / or controllers 1801. Backups 1808 may include networking resources, computing resources, application networks, and / or data storage resources as described in more detail above. Backups 1808 may or may not include controllers. Backups 1808 can be a complete copy of system 100. Backups 1808 may include the data needed to rebuild the system / environment / application from the configuration rules included in the backup, and may include the application data used. Backups 1808 can be set up as an application or using hardware similar to or different from system 100. Backup 1808 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or disabled, and / or converted to unidirectional read and / or write connections.
[0351] Figure 18B It is shown especially in Figure 18A Used when the system includes volatile data or when the database is writable Figure 18A
[00106] This example process flow for change management of a system in a system. Such a database may be part of the storage resources used by the environments in the system. At step 1870, the system (including the production environment) is deployed using global system rules.
[0352] At step 1871, the production environment is cloned to create a read-only environment using the global system rules 210 including the change management rules 1811 and the resource allocation made by the master controller 1801 or the controller in the clone environment. The clone environment is prohibited from writing to the system. The clone environment can then be used as a development environment.
[0353] At step 1872, hot spare 1820a is activated and assigned to clone environment 1807 to store any volatile data that has changed in system 100. The clone data is updated so that the new version of the development environment can be tested with the updated data. Hot sync data can be turned off at any time. For example, hot sync data can be turned off when testing a write from the old or production environment to the development environment.
[0354] At step 1873, the user can then use cloned environment 1807 as the development environment to process the changes. The changes to the development environment are then tested at step 1874. At step 1875, a determination is made as to whether the changed development environment is ready for production (typically, this determination is made by the developer). If step 1875 determines that the changes are not yet ready, the process flow may return to step 1873, where the user may back out and make additional changes to the development environment. If step 1875 determines that the changes are ready for production, the process flow proceeds to step 1876, where configuration rules are updated in the system or controller for the specific environment and the new, updated environment is deployed using these configuration rules.
[0355] At step 1877, the development environment (or new environment) can then be redeployed with the changes to the desired final configuration with the required resources and hardware allocations before execution. In the next step, at 1878, the write capability of the original production environment is disabled, and the original production environment becomes read-only. Although the original production environment is read-only, as part of 1878, any new data from the original production environment (or possibly the new production environment) can be cached and identified as transitional data. As an example, the data can be cached in a database server or other suitable location (e.g., a shared environment). The development environment (or new environment) and the old production environment are then switched at step 1879, so that the development environment (or new environment) becomes the production environment.
[0356] After this switch, the new production environment is made writable at step 1880. If the new production environment is deemed operational at step 1881, as determined by the developer, any data lost during the switch process (where such data had been cached at step 1878) can be reconciled with the data written to the new environment at step 1884. After this reconciliation, the change is completed (step 1885).
[0357] If step 1881 determines that the new production environment is not working (e.g., a problem is identified that requires the system to be restored to the old system), the environment is switched back at step 1882, so that the old production environment becomes the production environment again. As part of step 1882, the configuration rules for the principal environment on controller 1801 are restored to the previous version that was used for the now restored production environment.
[0358] At step 1883, changes to the database can be determined, for example, using cached data; and the data can be restored to the old production environment using the old configuration rules. To support step 1883, the database can maintain a log of changes made to it, allowing step 1883 to determine which changes may need to be undone. A backup database can be used to cache the data described above, with the cached data tracked and clocked; the clock can be restored to determine which changes have been made. Snapshots and logs can be used for this purpose.
[0359] After restoring the cached data at 1883, the process may return to step 1871 if necessary to start over.
[0360] The example change management system discussed herein can be used, for example, when upgrading, adding, or removing hardware or software, when patching software, when detecting a system failure, when migrating hosts during a hardware failure or detection, for dynamic resource migration, for changing configuration rules or templates, and / or for making any other system-related changes. The controller 1801 or system 100 can be configured to detect a failure and, upon detection of a failure, automatically implement change management rules or existing configuration rules to other hardware available for the system or controller. Examples of fault detection methods that can be used include, but are not limited to, checking hosts, querying applications, and running various tests or test suites. The change management configuration rules described herein can be implemented upon detection of a failure. Upon detection of a failure, such rules can trigger the automatic generation of a backup environment, automatic migration of data or resources, implemented by the controller. The selection of backup resources can be based on resource parameters. Such resource parameters can include, but are not limited to, usage information, speed, configuration rules, and data capacity and usage.
[0361] As described herein, any time a change occurs, the controller creates a log of the change and what was actually performed. For security or system updates, the controller described herein can be configured to automatically power on and off according to configuration rules and update the IT system status. The controller might shut down resources to save power. It might power on or migrate resources at different times for different efficiencies. During migrations, configuration rules are followed, and the environment or system can be backed up or replicated. If a security breach exists, the controller can isolate and shut down the affected area.
[0362] Configure and control service dependencies Figure 19A Illustrations for this article Figure 1-1 8, wherein the system 100 is augmented with related services (or applications), as shown by corresponding service modules 1901 and 1902 on one or more resources 1910. Service modules 1901 and 1902 may take the form of computer executable code that provides services such as authentication, email, webmail, web services, middleware, databases, and / or other services. For each reference, service 1901 may be referred to as service A, and service 1902 may be referred to as service B.
[0363] Figure 19A The system can be connected to the external network 1980 and / or the application network 390, and the connection can be based on Figures 13A to 13E Services 1901, 1902 are configured by the controller 200 as resources or applications, as described herein with reference to Figure 1 to Figure 1 8 as described in the Examples.
[0364] Services 1901 and 1902 can be controlled by controller 200 and can also interoperate via a common API 1903. Services 1901 and 1902 can use common API 1903 to resolve dependencies. For example, suppose a web application requires an HTTP server. A service with Apache or Nginx can have a "web server common API" that enables the server to serve the webapp's content and proxy information back to the application. Services 1901, 1902, and API 1903 can be directly coupled to controller 200 via a management network or via management connections (e.g., 260 and / or 270), or any other network connection between controller 200, services 1901, common API 1903, and services 1902.
[0365] Universal API 1903 can operate on or respond to services 1901, 1902, controller 200, or other resources in system 100. Service A in service module 1901 can be a dependency service, configured to be called by dependent service B in service module 1902 via universal API 1903 to perform one or more functions. A dependency service is a service that satisfies the dependency of another service (in this case, the other service is the dependent service). A dependency service can also be an optional dependency service.
[0366] Services 1901 , 1902 may be configured or created by the controller 200 to securely interoperate with each other.
[0367] Figure 19A Reference for examples of services and controllers interacting with each other Figure 19B . A service may be started, for example, by the controller 200 using configuration rules as described herein (see 19.1). The controller 200 resolves dependencies as described in the figures herein (see 19.2). A service may have a set of dependencies listed in its specification. As an example, this may be done via a JSON specification of the service. The system may also use dependency resolution similar to how a package manager works and provide the user with a method for satisfying dependencies. As another example, the system may provide the user with options for installing a dependency service or using / selecting an existing dependency service. Dependency service B makes a call to dependency service A via a universal API 1903 (see 19.3). This call may be a call to configure service A to support service B or to use some functionality of service A. Universal API 1903 translates dependency service A (1901) to instruct dependency service A to run a command (see 19.4). The translation may be done by one of the services or by the controller 200 making an API call (wherein an API function may call another API function on a different API).
[0368] According to some exemplary embodiments described herein, additional security can be provided for a system having a controller 200 and dependencies and dependent services 1901 and 1902. This additional security is useful when multiple services are simultaneously connected to an in-band management connection 270 and can communicate directly with each other. This additional security can be provided based on and / or using controller global system rules 210, logic 205, templates 230, and / or system state 220 during configuration, reconfiguration, and / or operation. In some exemplary embodiments, services 1901 and 1902 communicate via the in-band management connection 270 or other network or interconnect, which can provide additional security. According to some exemplary embodiments, a dependent service 1901 is configured to require the controller to authenticate the dependent service 1902. This can include authenticating the identity of the dependent service 1902 or the identity of the service running a command on an API. According to some exemplary embodiments, the dependent service is configured to request permission from the controller 200 to execute one or more functions, tasks, or a combination thereof for the specific dependent service(s). The dependency service may also or alternatively be provided with permissions or a set of permissions by the controller 200 when configuring or reconfiguring the dependency service. The permissions set in the dependency service may also be updated. For example, when adding a dependency service, the permission set may be updated.
[0369] exist Figure 19CThe illustrated flow illustrates an example of authentication and permissions provided between services. At step 19.11, the controller 200 provides a key or key pair to the service during configuration, thereby enabling authentication between the service and the controller 200. This step can be performed for each service in the system. The dependent service requests authentication from the controller, and the dependent service verifies the capabilities of the dependent service's request. The identification of the service and the data transmitted to and from the service can be verified, for example, through mutual TLS authentication, public key authentication, other forms of encryption, any network-based authentication technology (including but not limited to VLANs, VXLANs, partitions, etc.), and / or a combination thereof. Virtual networks and partitions can be used to divide a network into smaller networks, such as InfiniBand partitions. This can result in a situation where ports in so-called partitions 4 and 15 can only communicate on partitions 4 and 15. According to some variations, the controller 200 can serve as a key distribution center while maintaining authentication or authorization within service modules 1901 and 1902 independent of the controller 200. During configuration of a service, the controller may provide a key or key pair to the service, which may include a public key and / or private key that can be used to authenticate directly or indirectly between the service and the controller 200. According to one example, the controller 200 may act as the public key and delete or disable the provided private key. As another example, the service may generate its own keys in such a way that the public key from the service can be verified, identified and / or authenticated by the controller 200. In this further example, because the controller 200 provides the service with an initial public key, the service has the ability to send a trusted public key to the controller 200 without the controller 200 knowing the current private key of this trusted public key; and the service authenticates with its initial key pair to share the new public key with the controller 200.
[0370] At step 19.12, the dependent service (service B) calls the dependent service (service A) through the universal API 1903 to perform the function. The dependent service (service A) then authenticates the dependent service (service B) through the controller 200. According to one example, the dependent service can contact the controller 200, and the controller 200 can use the public key provided by the dependent service to authenticate the requesting dependent service. As above, the dependent service can obtain the public key at step 19.11. The dependent service can also create a new key pair (public + private key) and use the old key pair to prove that its new public key is authentic, because the controller 200 (in the case where the controller creates the public and private keys) knows that the old public key is trustworthy. The dependent service (service B) can authenticate the dependent service in a similar manner. (19.13) At step 19.14, the dependent service (Service A) may also establish permission to execute the function of the dependent service (Service B) before executing the function. For example, the dependent service may establish permission by querying the controller 200 if permission is available. As another example, the permission may be established by the controller 200 loading a permission manifest onto the dependent service (Service A).
[0371] Figure 19D An example of an enhanced security method for use with interoperable services is shown. At step 19.21, for example, the controller 200 and / or the template 230 described herein may be used to create a dependent service B. At step 19.22, when the dependent service B is run, as described herein for various embodiments (e.g., see Figures 13A-13E ), controller 200 can verify and / or disable connections to external network 1910 and / or application network 390. For example, practitioners may find it useful to disable management connections when developing services for networks such as the Internet. As explained above, this provides increased isolation and security. In this case, the cloud API (or in other cases, the out-of-band management connection 260) can be used to trigger the in-band management connection 270. At step 19.23, dependent service B executes an API command to request a service or function from dependent service A. This step can also be accomplished by dependent service B querying controller 200, and controller 200 executing the command via the general API (see 1903). At step 19.24, dependent service A verifies the identifier of dependent service B and the permissions to execute service B's service or function. As an example, this step 19.24 can be performed by service A verifying service B's permissions. As another example, this step 19.24 can be performed by service A verifying whether it has permission to provide services to service B. In either case, the service being modified by another service can ensure that the other service allows these modifications. If authorized and allowed, the dependent service A can then run the service or function specified in the command. At step 19.26, the management connection, such as out-of-band management 260, in-band management 270, or SAN 280 can be disconnected for additional security, as described herein with reference to Figures 13A-13E At step 19.27, if the connection to the external network 1980 and / or the application network 390 was disabled at 19.22, it may be re-enabled.
[0372] Figure 19E The diagram shows reference Figures 19A-19DAn example system 100 of the described system includes a set of purge rules 1904 included in the controller 200. The purge rules 1904 can be implemented as its own set of rules within the controller 200, or it can be implemented as global system rules, controller logic 205, templates 230, or a combination thereof. The purge rules 1904 include a set of instructions and rules to be followed when deleting a service. As an example, the purge rules 204 can be included in a template 230 used to set up a service, where the relevant rules for the service are loaded onto the service during setup or used to generate service specific purge rules. For example, a mail service can cause a DNS record to be added to the DNS service. If the mail service is deleted, the DNS service can remove the DNS record from the mail service.
[0373] Cleanup rules 1904 may be used to identify modifications made by a dependent service to a dependent service to enable deletion, removal, and / or revocation of those modifications when the dependent service is deleted or disabled. Figure 19F An example process flow for creating a purge rule is shown. For example, Figure 19F Illustrate how, when a dependent service is deleted, the removed modifications can be verified, for example logged and / or tracked.
[0374] like Figure 19F As shown, at step 19.31, a command is issued from the dependent service (Service B) to API 1903 to call the dependent service (Service A) to perform a function. At step 19.32, the dependent service (Service B) is verified and Figures 19A-19D Confirm the permissions as described in . If the dependent service (service A) is modified or is to be modified in the execution function, at step 19.33, the corresponding cleanup rules and / or cleanup commands corresponding to the API command are recoverably added to the dependent service (service B), the dependent service (service A) or one or more of the controller 200 (or associated with its cleanup rules). The cleanup rules or cleanup commands can identify modifications for subsequent cleanup in the case of deleting the dependent service (service A). According to different exemplary embodiments, the dependent service or dependent services have associated cleanup rules. The cleanup rules can also be configured to modify the connection between services. The dependent service can have such a cleanup rule that can correspond to each dependent service with which it has a relationship. The cleanup rules can also be generated by the logged API command. Each cleanup step can be executed once at the time obtained from the API command login.
[0375] When a dependent service is to be deleted, changed and / or modified, then a purge rule can be used. Purge rules can also be used when a dependent service has made changes to a dependent service. Figure 19GAs shown in the exemplary process flow in FIG, a determination is made at step 19.41 to delete a service. At step 19.42, the controller logic 205 looks at the dependent services of the service being deleted ("deleting service"). These dependencies may be recursive and may be discovered through recursive dependency resolution. Dependent services may be found in the example herein. Figures 2A-2K The dependent service is identified in the service template 230 or the system state 220 or other related database. If step 19.43 concludes that a dependent service exists, the controller can find alternative paths to satisfy the dependency relationship (see 19.44) (e.g., in Figure 2K ). If the controller 200 does not identify an alternative path to satisfy the dependency (see 19.45), the user / administrator can be notified to resolve the issue (see 19.46); for example, by adding a new dependent service, deleting the service, or similarly deleting the dependent service. If the controller identifies an alternative way to satisfy the dependency at step 19.44, the controller 200 can change the dependency and update and / or reconfigure the corresponding controller components (templates 230, rules 210, logic 205, system state 220, etc.) as well as the dependent service and / or dependent service components. Cleanup rules can then be followed at step 19.47. If step 19.43 concludes that there are no dependent services, the process flow can also proceed to step 19.47, where cleanup rules can be followed. As described, cleanup rules identify modifications made to the dependent service by the dependent service. Therefore, at step 19.47, these cleanup rules can be processed to enable the deletion, removal, and / or reversal of these modifications. As part of step 19.47, files created by the deleting service during use are removed from the dependent service.
[0376] Provide storage resources for computing resources: Figure 20A An exemplary system is illustrated that includes a controller 200 in which one or more computing resources host one or more services that utilize memory in one or more storage resources 410 . Figure 21A As further shown, if a physical service host is allowed to talk to the SAN 280, it is expected that such physical service host will only access remote storage that it is authorized to use. Accordingly, the system is preferably able to prevent bad actors 2002 from gaining unauthorized access to system resources, such as storage resources 410 (see Figure 20A ).
[0377] according to Figure 20BIn the example process flow described in FIG. 20 , controller 200 provides storage credentials to computing resources (see 20.10). By way of example, storage credentials may take the form of, but are not limited to, a password, a passphrase, a Challenge Handshake Authentication Protocol (CHAP) key, an encryption key, a certificate, or a combination thereof. CHAP is an authentication technology used for remote storage, such as iSCSI / iSER or other technologies. CHAP keys can be used as passwords for SANs. The provisioning at step 20.10 can be performed in any of a number of ways. For example, when creating a service image, controller 200 may include storage resource connection information with the service image. The computing resource may also query controller 200. Alternatively, controller 200 may provide information to computing resources during the boot process (this may be done after creating / provisioning storage resources, if they are created on demand). All of this information may be located in a database, and controller 200 may extract the storage credential information from the database, or the computing resource may request the storage credential information from the controller by performing a data query or making a database query API call. The computing resource may then use the storage credentials to connect to, log in to, or communicate with the storage resource.
[0378] According to Figure 20C Alternatively or additionally, in the example process flow described in [ 20.20 ], the SAN connection 280 between the compute resource 310 and the storage resource 410 can be disabled, as described in various embodiments herein (see 20.20). The compute resource 310 can then be provisioned with the storage resource 410 on a specific isolated connectivity network, including but not limited to VLANs, VXLANs, and InfiniBand partitions. Accordingly, the controller 200 can pair the compute resource with the storage resource and place them on the same network or fabric. For example, a port c...
Claims
1. An information technology (IT) computer system comprising: Controller; a computing resource for connecting to the controller; as well as storage resources used by the computing resources; wherein the controller is configured to provide storage credentials for the storage resource to the computing resource; and Wherein, the computing resource is configured to connect to the storage resource, log into the storage resource, and / or communicate with the storage resource based on the storage credentials.
2. The system according to claim 1, wherein: The controller is further configured to: (1) pair the computing resources and the storage resources; and (2) place the paired computing resources and storage resources on the same network or fabric.
3. The system according to claim 2, wherein: The controller is further configured to: (1) disable a storage area network (SAN) connection between the computing resource and the storage resource; and (2) make the storage resource available to the computing resource on an isolated connection network, thereby placing the paired computing resource and storage resource on the same network or fabric.
4. The system according to claim 3, wherein: Isolated connectivity networks include vlans, vxlans, and / or InfiniBand partitions.
5. The system according to any one of claims 1 to 4, wherein: The stored credentials include a password or a passphrase.
6. The system according to any one of claims 1 to 5, wherein: The stored credentials include the chap key.
7. The system according to any one of claims 1 to 6, wherein: The stored credentials include an encryption key.
8. The system according to claim 7, wherein: The computing resource is configured to calculate login credentials for the storage resource from an encryption key based on an encryption technique.
9. The system according to any one of claims 1 to 8, further comprising a plurality of computing resources and a plurality of storage resources; and wherein, The controller is configured to provide storage credentials for the storage resources to the computing resources such that a plurality of computing resources are paired with different ones of the storage resources.
10. An information technology (IT) method for use with a computer system comprising a controller, a computing resource connected to the controller, and a storage resource for use by the computing resource, the method comprising: The controller provides storage credentials for the storage resource to the computing resource; as well as The computing resource connects to, logs into, and / or communicates with the storage resource based on the storage credentials.