Modular and expandable high performance computing system

WO2025153598A3PCT designated stage Publication Date: 2025-08-28BF EXAQC AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/050996
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-16
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Organizations face challenges in designing and deploying high-performance computing (HPC) systems due to complexity, high costs, and the need for specialized components, which often result in inefficient resource utilization and prolonged deployment times, while maintaining data privacy and security is a concern when using third-party infrastructure.

Method used

A modular and expandable HPC system comprising a general-purpose processor module, an accelerator module, and a storage module, allowing for customizable and scalable configurations tailored to specific enterprise needs, with a comprehensive software stack for efficient management and integration.

Benefits of technology

The system provides flexible, cost-effective, and efficient HPC solutions that can be scaled to meet evolving demands, ensuring data security and compliance, with reduced deployment time and enhanced performance for AI-related tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025050996_28082025_PF_FP_ABST
    Figure EP2025050996_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The proposed invention is a modular and expandable high-performance computing system that addresses the existing challenges related to the design and deployment of high-performance computing (HPC) systems. The invention provides a cost-effective and flexible solution by incorporating modular components, including a general-purpose processor module, an accelerator module, and a storage module. These modules can be expanded or upgraded individually or in combination to meet evolving computational needs. The modular design allows for customization, scalability, and easy integration, while the comprehensive software stack enhances system management, programming flexibility, and development environment. The invention provides a versatile and efficient solution for researchers, organizations, and industries to harness the full potential of HPC resources.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Modular and Expandable High Performance Computing System

[0002] The present invention relates to a high-performance computing system, particularly to a modular and expandable high-performance computing system.

[0003] Background of the Invention

[0004] Artificial intelligence (Al) has emerged as a transformative technology that is set to revolutionize various aspects of our lives, ranging from business and science to society and art. At the heart of this revolution lies the development and deployment of high-performance computing systems to support the training and execution of advanced Al models, particularly in the context of Large Language Models (LLM) and Large Multimodal Models (LMM), and their refinement through the integration of application-specific enterprise data.

[0005] Artificial intelligence, defined as the ability of machines to mimic human intelligence, has gained significant momentum in recent years. This revolution is fueled by advancements in machine learning algorithms and the availability of huge amounts of data. Consequently, the scope of Al applications has expanded and is expected to further permeate various industries, driving innovation, and transforming traditional processes. High-performance computing (HPC) systems are the backbone of Al revolution, as they provide the computational power necessary to train and execute complex machine learning models. These systems typically include clusters of interconnected computers that work together to process vast amounts of data and perform massive calculations in parallel. Utilizing HPC systems enables researchers and data scientists to accelerate the training process, achieve higher performance, and effectively tackle challenging Al tasks.

[0006] In recent years, Large Language Models (LLM) have emerged as a crucial pillar of Al research. LLMs, such as OpenAI's GPT-4, are trained on massive amounts of text data to generate human-like text responses. These models have far-reaching potential, aiding in areas such as language translation, content generation, and even dialogue systems. However, the training of LLMs requires substantial computing power and abundant training data, making them highly dependent on high-performance computing systems. Expanding upon LLMs, Large Multimodal Models (LMM) are designed to process and understand multimodal inputs, combining textual, visual, and auditory information. With the rise of multimedia data on platforms like social media, LMMs have gained immense interest due to their ability to comprehend and generate content that goes beyond just text. By training on robust HPC architectures, LMMs become more capable of performing tasks such as complex image recognition, natural language processing in combination with visual or auditory context, and more.

[0007] While pretraining LLMs and LMMs using public datasets has yielded impressive results, refining these models using application-specific enterprise data further enhances their performance for specific use cases. Enterprises possess vast amounts of domain-specific data, including customer interactions, industry-specific jargon, and business processes. By integrating this data during fine-tuning, high-performance computing systems empower organizations to create tailored Al solutions that drive operational efficiency, personalized user experiences, and accurate decision-making.

[0008] Organizations possess invaluable proprietary data, including sensitive customer information, trade secrets, and confidential business processes. Leveraging this data to train Al models allows organizations to extract valuable insights, enhance decision-making, and gain a competitive edge. However, doing so in-house can be a daunting task due to the need for substantial computational resources and specialized expertise. However, using third-party infrastructure for training and fine-tuning Al models might not be a solution for many enterprises.

[0009] One of the primary concerns when leveraging third-party infrastructure is the potential compromise of data privacy and security. Entrusting proprietary data to external parties raises apprehensions regarding unauthorized access, data breaches, and compliance with regulatory requirements such as General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). Moreover, the risk of intellectual property theft must also be addressed. Many organizations therefore prefer to retain complete control and oversight of their data, fearing potential mismanagement or misuse by third-party providers.

[0010] This leads to the need for organizations to provide their own infrastructure. By establishing their own infrastructure for Al training, organizations can retain complete control and oversight of their data. This allows them to ensure compliance with privacy requirements and implement robust security measures tailored to their specific needs. With complete ownership of the infrastructure, organizational data remains within its boundaries, reducing the risk of unauthorized access or data leakage. However, many organizations face almost insurmountable problems when it comes to planning and creating their own infrastructure, since designing and building a High- Performance Computing (HPC) system is a complex endeavor due to the need for highly specialized components to work in concert to achieve optimal performance. Each aspect of an HPC system, from compute nodes and accelerators, interconnects, storage, cooling and power infrastructure, to job scheduling software, presents its own set of challenges.

[0011] Compute nodes must integrate general-purpose processors with accelerators like GPUs to balance the workloads that require different processing capabilities. This requires detailed knowledge of application requirements and a keen understanding of hardware capabilities. The inclusion of accelerators adds another layer of complexity, involving software compatibility and programming models.

[0012] The interconnects must facilitate high-speed, low-latency communication between nodes for efficient parallel computing. The choice of topology and technology has a profound impact on the system's ability to handle data-intensive tasks.

[0013] Storage subsystems must offer high bandwidth and capacity while maintaining data integrity and availability. Designing a storage hierarchy that can serve high-speed access for active data while also accommodating large volumes of less frequently accessed data is challenging and critical for system performance.

[0014] Cooling and power are intertwined, as the more powerful a system, the more heat it generates and the more energy it consumes. Engineering solutions that provide the necessary cooling capacity without introducing prohibitive power costs or environmental concerns adds another layer of complexity.

[0015] Finally, software ecosystems that include job schedulers and resource management tools must be robust and efficient to ensure fair and optimal use of resources. This involves scheduling algorithms and policy settings that can deal with a wide range of task priorities and resource demands.

[0016] One approach to building HPC systems is to implement the modular system architecture as described in US 10,142,156 by inventor Thomas Lippert. This architecture allows for a more flexible and scalable design, enabling the system to adapt to changing computational needs and requirements. In a modular HPC system, the computing resources are divided into independent modules that can be easily interconnected and configured as needed. The main advantage of modular HPC systems is their ability to provide communication flexibility and efficient data exchange between computation nodes and accelerators. This flexibility is achieved by incorporating independent accelerator units that perform specific computational tasks. These accelerators are coupled to a communication infrastructure which allows for faster and more efficient data transfer, resulting in improved system performance.

[0017] Additionally, modular HPC systems offer dynamic coupling of accelerators to computation nodes at runtime. This means that the system can dynamically assign accelerators to computation nodes based on a predetermined assignment metric. The assignment metric is managed by a resource manager, which can establish both static assignments at the start of a computation task and dynamic assignments during task processing. This dynamic assignment capability enables the system to optimize resource utilization and workload distribution, leading to improved efficiency and scalability.

[0018] Another key feature of modular HPC systems is fault tolerance. In the event of a booster failure, the system can automatically redistribute the workload to other available boosters, ensuring uninterrupted computation. This fault tolerance mechanism enhances system reliability and robustness, which is critical for mission-critical applications and time-sensitive tasks.

[0019] Even though modular systems have been suggested for High-Performance Computing (HPC) systems, they are still predominantly individually designed and built. This approach has a few drawbacks. When each HPC system is uniquely designed and built, it requires significant financial resources. The hardware needs to be procured separately, resulting in higher costs compared to standardized, off-the-shelf systems. Moreover, individual designs often necessitate specialized expertise, leading to increased labor costs. Furthermore, designing and building a custom HPC system can be a time-consuming process. Each component needs to be carefully selected, tested, and integrated into the overall system. Additionally, tailored hardware configurations might require longer lead times for manufacturing and delivery. Consequently, this can delay the deployment of HPC resources for research or other time-sensitive activities. Further, tailor-made solutions may lack the flexibility to adapt to evolving needs. As technology advances and new research requirements emerge, the individually designed HPC systems may not accommodate necessary upgrades or changes easily. As demands grow and more computational power is required, it could be challenging to expand an individually designed HPC system. Scaling individual components might be limited by compatibility issues or architectural constraints, leading to inefficient resource utilization. With individually designed HPC systems, maintenance and support can become more complex. As each component may come from different vendors, managing service agreements with multiple parties can be challenging.

[0020] Brief Summary of the Invention

[0021] To address these issues and provide a more flexible and cost-effective solution, a modular and expandable HPC system, also referred to as computing system, is proposed. The proposed modular and expandable HPC system is designed to meet immediate computational demands while offering the capacity for future expansion. It consists of a set of modular components including a general-purpose processor module, an accelerator module, and a storage module.

[0022] The general-purpose processor module forms the core of the HPC system and is responsible for executing various general computing tasks. It comprises standard CPUs and memory modules that provide the necessary computational power. The design allows for expanding the processor module.

[0023] The accelerator module, which operates independently or in conjunction with the general- purpose processor module, enhances the system's performance for specific computeintensive applications. Specialized accelerators, such as GPUs or FPGAs, are integrated into this module to offload specific computational tasks, enabling faster processing and improved efficiency. Like the general-purpose processor module, the accelerator module is configured to be expanded to accommodate evolving needs.

[0024] To support data-intensive workloads, a storage module is included in the modular and expandable HPC system. It houses high-capacity storage devices capable of handling large volumes of data. As data storage requirements increase, additional storage modules can be attached, ensuring scalability and uninterrupted accessibility. Moreover, advancements in storage technology can be easily incorporated, thereby keeping the system up-to-date in terms of both capacity and performance.

[0025] In addition to addressing future demands, one of the key advantages of the modular and expandable HPC system is its ability to offer systems of different sizes and capacities with ease. This feature allows for the customization of HPC solutions based on the specific requirements of different enterprises or industries. For a small enterprise, a 1-Rack System can be configured using a basic modular HPC system. This compact solution provides sufficient computational power and storage capacity to meet the needs of small-scale operations. By selecting and attaching a specific number of modular and expandable systems, small enterprises can create an HPC system tailored to their requirements without excessive investment or complex integration processes.

[0026] To cater to the needs of medium-sized enterprises, a 5-Rack System can be realized by combining five small enterprise solutions, which could be installed in half a container of a Modular HPC Data Centre (MDC). The ability to aggregate multiple basic modular HPC systems allows for increased computational power and storage capacity. As the demand for resources grows, medium enterprises can easily expand their HPC capabilities by incrementally attaching additional units. This flexibility enables them to align computational resources with evolving requirements efficiently.

[0027] For larger and more resource-intensive industries, such as research institutions or government bodies, a modular and expandable HPC system can be scaled up to fill multiple containers within a Modular HPC Data Centre (MDC). For instance, by using the basic building blocks, a system filling 5 containers can be created. This industry-level solution provides extensive computational power and storage capability, allowing for high- performance computing on a massive scale. The ability to expand the system by attaching more modular units offers seamless scalability to meet changing computational demands.

[0028] In addition to providing overall system scalability and customization, the modular and expandable HPC systems offer the advantage of expanding individual modules according to specific needs. This flexibility allows for targeted scaling of components without the necessity of expanding the entire system. For example, the accelerator module can be easily expanded by attaching additional accelerators, such as a rack full of accelerator blades, without affecting the general-purpose processor module or the storage module. Moreover, other modules can also be expanded individually or in combination, providing a highly adaptable and efficient solution. As the computational demands of specific workloads increase, attaching additional accelerators, such as GPUs or FPGAs, offers enhanced processing power for those tasks. By expanding the accelerator module independently, organizations can optimize their system to achieve higher performance and efficiency without the need to invest in additional resources for other modules. With individual module expansion, the general-purpose processor module and storage module can remain unchanged if they provide sufficient processing power and storage capacity for current requirements. This flexibility allows organizations to allocate resources strategically and efficiently. Instead of expanding all modules simultaneously, they can prioritize their investments and scale up specific components as needed, optimizing costeffectiveness without sacrificing performance.

[0029] The ability to expand individual modules or combination of modules caters to different workload requirements. For instance, an organization primarily focused on data-intensive tasks may choose to expand the storage module to accommodate increasing data volumes without expanding other components unnecessarily. Conversely, an organization with compute-intensive workloads may decide to expand the accelerator module to boost performance without enlarging the general-purpose processor module or the storage module. This tailored expansion approach enables organizations to scale their computational resources precisely to meet their unique workload demands.

[0030] As technology evolves rapidly, the expansion of individual modules allows organizations to stay abreast of advancements. This future-proofing capability allows organizations to extend and adapt their HPC resources seamlessly, maximizing their investments in a modular and expandable architecture.

[0031] The advantages of modular and expandable HPC systems extend beyond simply accommodating future expansion needs. Their modular design allows for the easy configuration of systems with varying sizes and capacities. By selecting and attaching a particular number of basic modular systems, enterprises can create solutions specific to their needs, whether they operate on a small scale, require moderate resources, or necessitate the immense computational power of an industry-level solution. This customizability ensures that HPC systems can be tailored to suit the requirements of different enterprises and industries, while still enjoying the benefits of scalability, cost-efficiency, and reduced time-to- production offered by modular designs.

[0032] By adopting a modular and expandable approach, this system reduces costs compared to tailor-made solutions. Standardized components and the ability to incrementally expand the system lower initial investment. Additionally, users can strategically invest in expansions and upgrades as needed, optimizing cost-effectiveness over time. The modular and expandable nature of this system reduces the time required for deployment. With pre-designed modules readily available, integration and testing become simpler and more streamlined. This drastically shortens the time-to-production, enabling researchers and organizations to access computational resources promptly.

[0033] The basic modular and expandable HPC system offers unparalleled flexibility. It allows for easy expansion by attaching additional basic modular systems or specific extensions. Users can scale their computational power incrementally, keeping pace with evolving demands without disrupting existing operations. Furthermore, the modular and expandable design provides the ability to adapt to future hardware advancements.

[0034] The basic modular and expandable HPC system presents an innovative approach to overcome at least some of the drawbacks associated with individually designed systems. By providing a flexible, cost-effective, and expandable solution, this approach addresses the high costs and extended time-to-production challenges. With its ability to adjust to increasing computational demands through incremental expansion or specific extensions, this system empowers researchers and organizations to harness the full potential of HPC resources in a scalable and efficient manner.

[0035] When considering the installation of HPC solutions for different enterprise sizes, the 1-rack solution for small businesses can be installed in a server room or similar on-site facilities. However, it is advantageous to opt for a Modular HPC Data Centre (MDC) when deploying medium or large enterprise solutions such as the 5-rack system.

[0036] A Modular HPC Data Centre (MDC) refers to an infrastructure solution that utilizes containers to house both the computing and support equipment of the supercomputer. The enterprise only needs to provide the structural foundation and basic infrastructure like electricity and water supply via well-defined interfaces.

[0037] Containers within the MDC can be allocated to accommodate individual instances of the basic building blocks. For example, a container may host a general-purpose processor module, an accelerator module, or a storage module. This arrangement provides flexibility and modular scalability, allowing each container to function as an independent computational unit.

[0038] Containers can also be configured to house a combination of the basic building blocks along with module-specific expansions. For instance, a container may include a general-purpose processor module and be expanded with additional accelerators or storage devices. This configuration provides a universal and scalable solution, offering enhanced computational capabilities with the support of specialized modules.

[0039] To cater to needs, containers within the MDC can be exclusively designated to host expansions of one specific kind. For example, a container may be dedicated to accommodating additional accelerators to bolster compute-intensive workloads or to integrating extra storage capacity to handle large-scale data processing. This approach allows organizations to target their expansions precisely, optimizing the performance of specific modules without affecting others.

[0040] The containerized standardization inherent in MDCs facilitates the deployment process. Once the containers are ready, they can be easily transported to the site and quickly integrated into the MDC. This standardized approach reduces installation time and complexity, ensuring faster time-to-production for the HPC system.

[0041] The independent operation of parts within separate containers enhances fault tolerance and reduces the impact of hardware failures. In the event of a component failure in one container, the other containers can continue to function independently, minimizing downtime and ensuring uninterrupted operation of critical applications.

[0042] Maintaining an HPC system can be complex, but MDCs simplify maintenance tasks. With containers housing specific hardware modules, maintenance or upgrades can be performed on individual containers without affecting the overall system operation. This compartmentalized approach streamlines maintenance processes, ensuring efficient and seamless management of the HPC infrastructure.

[0043] As technology evolves, hardware components may require upgrading or replacement. With the modular design of MDCs, specific containers can be easily upgraded or replaced without disrupting the entire system. This flexibility enables organizations to keep their HPC infrastructure up to date, leveraging the latest technologies without significant interruptions or investments.

[0044] The modular and expandable high-performance computing system offers a complete software stack that encompasses various essential components to facilitate seamless training, data management, curation, and inference services. The software stack can be categorized into four key categories for efficient operation and control: Management software components encompass the overarching management of the modular and expandable high-performance computing, based on the Modular System Architecture (MSA). These components handle system-level functionalities, such as resource allocation, task scheduling, and load balancing. They ensure optimal utilization of computational resources, enabling efficient training and overall system performance.

[0045] An Al training software is specifically designed to cater to the training on the HPC system. It includes frameworks, libraries, and tools that facilitate machine learning model development, parameter tuning, and model optimization. This software enables the efficient execution of complex training algorithms and models, leveraging the computational power of the HPC system.

[0046] A data management software addresses the specific requirements of the data module within the HPC system. It consists of tools and protocols for data ingestion, preprocessing, storage, curation, and quality control. This software manages large-scale data sets and enables efficient data access and manipulation, ensuring data integrity and enhancing the overall training process.

[0047] An inference services software is dedicated to performing inference on the HPC system. It provides the necessary infrastructure and APIs to deploy and operate machine learning models in production environments. This software handles real-time or batch processing of input data, leveraging the trained models to make predictions or generate insights.

[0048] Additionally, it incorporates feedback mechanisms to support continuous training, allowing the models to improve and adapt over time.

[0049] The modular and expandable HPC system's complete software stack is designed to work together seamlessly, ensuring compatibility, and streamlined integration between different components. This integration simplifies the setup and configuration process, reducing the time and effort required to deploy the system. The software stack is optimized to leverage the full potential of the HPC system's computational resources. With specifically tailored software for each module and functionality, such as Al training, data management, and inference services, the system can efficiently execute resource-intensive tasks, leading to improved performance and faster results.

[0050] The comprehensive software stack provides a range of tools and frameworks that simplify and automate various aspects of the machine learning lifecycle. This streamlines the development, deployment, and management of machine learning models, enabling researchers and data scientists to focus more on their core tasks and enhance productivity. The software stack is designed to support the modular and expandable nature of the HPC system, allowing for easy scalability and adaptability to changing requirements. Additional modules or expansions can be seamlessly integrated into the existing software stack, ensuring a flexible and future-proof solution.

[0051] The modular and expandable high-performance computing (HPC) system is equipped with advanced software tools and frameworks that enhance system management, programming flexibility, version control, containerization, and development environment. These software components offer a range of advantages for efficient management and utilization of the HPC system.

[0052] The ParaStation software by ParTec AG provides comprehensive system control capabilities for the modular and expandable HPC system. It offers efficient task scheduling, resource allocation, and load balancing, maximizing the utilization of computational resources. With ParaStation, the system can effectively manage job submissions, monitor performance, and ensure optimal execution of computing tasks.

[0053] ParaStation software also encompasses system management functionalities for the HPC system. It enables centralized administration of nodes, user management, and system diagnostics, providing administrators with powerful tools for system monitoring, troubleshooting, and maintenance. This streamlines system management tasks, ensuring smooth operation and efficient utilization of resources.

[0054] Python preferably serves as a versatile programming language for the modular and expandable HPC system. Its simplicity, rich ecosystem of libraries, and extensive community support make it an ideal language for developing and executing various computational tasks. Python enables researchers and developers to write efficient, scalable, and easily maintainable code, enhancing productivity and facilitating seamless integration with other system components.

[0055] Git may advantageously be used as a robust version control system, ensuring effective project management, collaboration, and code versioning within the HPC system. With Git, developers can track changes, manage branches, and coordinate collaborative efforts, enhancing code quality, reproducibility, and facilitating seamless project updates. The HPC system leverages Docker and Kubernetes to enable efficient containerization and orchestration of applications and software components. Docker simplifies the deployment and portability of applications, ensuring consistency across different environments. Kubernetes provides orchestration capabilities for managing containers at scale, enabling automated scaling, fault tolerance, and efficient resource allocation.

[0056] The system supports popular development environments such as Jupyter Notebooks and integrated development environments (IDEs) tailored for Python. Jupyter Notebooks offers an interactive and collaborative workspace for code experimentation, data exploration, and visualization. IDEs provide advanced coding features, debugging capabilities, and project management tools, enhancing the development experience for researchers and developers.

[0057] The ParaStation software stack includes modular and parallel libraries designed to optimize performance for specific computational tasks. These libraries provide high-performance computing capabilities, enabling efficient parallelization of algorithms and easing the utilization of the HPC system's resources. Leveraging the modular and parallel libraries provided by ParaStation reduces code complexity, enhances performance, and improves time-to-solution.

[0058] By incorporating these software tools and frameworks, the modular and expandable HPC system offers numerous advantages, including streamlined system control, efficient system management, programming flexibility, version control, containerization, and development environment support. These advantages enhance productivity, simplify software development and deployment, optimize resource utilization, and facilitate seamless collaboration within the HPC system ecosystem.

[0059] The presently described expandable computing system is designed specifically for high- performance computing applications, with a focus on Al-related tasks. It consists of three main components: a cluster module, a booster module, and a storage module. These components are interconnected through a communications infrastructure, allowing for efficient data transfer and collaboration.

[0060] The cluster module is made up of multiple computation nodes that collectively handle the processing tasks. These computation nodes are designed to work together, dividing the workload among themselves to achieve better performance and speed. Examples of computation nodes could be powerful processors, such as multi-core CPUs, capable of executing complex calculations and algorithms. The booster module consists of additional nodes that enhance the computing power of the system. These booster nodes are specifically designed to tackle demanding Al-related tasks and provide additional processing capabilities. They can be equipped with specialized hardware accelerators, like GPU, FPGA or ASIC, that are optimized for Al computations or a combination of multi-core CPU and GPU, such as Nvidia's Grace Hopper Superchip. The booster nodes work independently, or in conjunction with the computation nodes in the cluster module to augment the system's performance.

[0061] The storage module plays a crucial role in the system by providing a centralized storage infrastructure. It securely stores and manages the data required for computation and analysis. The storage module can be implemented using technologies like solid-state drives (SSDs), hard disk drives (HDDs), or even distributed storage systems such as network- attached storage (NAS) or storage area networks (SANs). It ensures the availability and accessibility of data to the computation and booster nodes.

[0062] The communications infrastructure is responsible for establishing connections between the computation nodes, booster nodes, and storage module. It enables communication and data transfer among the modules within the system. The infrastructure includes interfaces that allow expanding the cluster module and booster module, facilitating scalability and future enhancements.

[0063] For example, the communications infrastructure can utilize high-speed interconnect technologies like InfiniBand or Ethernet for fast and reliable communication between the modules. The interfaces for expanding the cluster module and booster module can be implemented using industry-standard protocols such as PCI Express (PCIe) or Ethernet, allowing the incorporation of additional computation or booster nodes as needed.

[0064] Due to the design, the system can be easily expanded by adding more computation nodes or booster nodes, allowing for increased computing power and performance. The expandable nature of the system enables customization to meet specific computing requirements. Depending on the workload, the number of computation nodes and booster nodes can be adjusted accordingly. With specialized booster nodes designed for Al tasks, the system can efficiently handle complex machine learning algorithms and neural network training, whereas the storage module ensures organized and accessible storage of data, facilitating quick retrieval and sharing of data between computation and booster nodes. The communications infrastructure enables smooth and fast data transfer between the computation nodes, booster nodes, and storage module, minimizing latency and maximizing system efficiency, and the expandability allows for incremental upgrades as needed, avoiding complete system replacement when additional computing power is required.

[0065] The expandable computing system preferably comprises a communications infrastructure that includes two separate networks: the first communications network connecting the computation nodes in the cluster module and the second communications network connecting the booster nodes in the booster module. These networks are distinct but are interconnected through an interface that enables communication between them.

[0066] The first communications network connects the computation nodes within the cluster module. It allows seamless communication and data transfer between the computation nodes, thereby facilitating collaborative processing and workload distribution. The network can be implemented using high-speed interconnect technologies like InfiniBand or Ethernet, which provide low-latency and high-bandwidth communication. For example, the computation nodes may be interconnected using InfiniBand links, allowing for fast and efficient communication.

[0067] The second communications network connects the booster nodes within the booster module. This network is specifically designed to handle the communication needs of the booster nodes, which are responsible for executing specialized tasks related to Al computation. It may utilize separate physical links or a separate logical network within the system. The booster nodes might require faster communication due to the complexity of the Al-related tasks they perform.

[0068] The interface between the first and second communications networks enables communication and coordination between the computation nodes in the cluster module and the booster nodes in the booster module. This interface provides a bridge for exchanging data and instructions between the two networks. It ensures that the computation nodes can efficiently access the resources and capabilities of the booster nodes when necessary, enhancing the overall performance of the system.

[0069] For example, the interface may use a combined protocol stack that supports both the first and second communications networks, allowing seamless data transfer between the two. It can be implemented using a network switch or router that connects the two networks together, facilitating communication and information exchange. Separating the communication networks for the computation nodes and booster nodes allows for efficient and optimized communication. The computation nodes can focus on their internal coordination and workload distribution, while the booster nodes can handle their specialized tasks, reducing potential bottlenecks.

[0070] The use of separate networks for the cluster module and booster module, along with the expansion interfaces, allows for easy scalability and expansion of the system. Additional computation nodes and booster nodes can be added independently as per the requirements, providing flexibility in system design and customization.

[0071] In one embodiment of the expandable computing system, a third communications network is provided to connect the first communications network and the second communications network. This additional network provides a dedicated pathway for communication between the computation nodes in the cluster module and the booster nodes in the booster module.

[0072] The third communications network ensures efficient and direct communication between the computation nodes and the booster nodes, enabling seamless data transfer and coordination. It can be implemented using various interconnect technologies, such as Ethernet, InfiniBand, or custom networking solutions designed for high-performance computing environments.

[0073] In another embodiment, the storage module is communicatively connected with the first communications network. This configuration allows the computation nodes in the cluster module to access the storage module directly, without the need for additional intermediate networks or interfaces.

[0074] By connecting the storage module to the first communications network, the system provides fast and efficient access to data for computation and analysis. This direct connectivity reduces latency and enables seamless data retrieval and sharing between the computation nodes. It allows for high-speed data exchange, which is essential for Al-related tasks that often require large volumes of data.

[0075] In an alternative configuration of the expandable computing system, the storage module is communicatively connected with the third communications network, rather than the first communications network. By connecting the storage module to the third communications network, the system introduces a separate pathway specifically dedicated to storage-related communication. This configuration offers several technical advantages and functional aspects: The storage module can benefit from a direct connection to the third communications network, which is designed to handle storage-related data transfer. This dedicated pathway ensures efficient and high-speed data transfer between the storage module and the computation nodes or booster nodes, minimizing latency and maximizing data access and retrieval speed.

[0076] Separating the storage module from the first communications network and connecting it to the third communications network creates a more streamlined network communication architecture. The first communications network remains dedicated to the computation nodes, allowing for parallel processing and efficient resource allocation. The third communications network is responsible for storage-related communication, ensuring a dedicated and optimized pathway for data transfer to and from the storage module. Further, separating the storage-related communication onto the third communications network enhances performance isolation within the system. The computational workload of the computation nodes and the specialized Al operations of the booster nodes can operate independently without affecting storage performance. This isolation allows the computation and booster nodes to efficiently utilize their respective networks while avoiding contention or congestion caused by storage-related data traffic.

[0077] A network topology refers to the arrangement or structure of nodes (devices or computers) and the links (connections) between them in a computer network. Different network topologies exist, each defining how nodes are connected and how data is transmitted. One example of a network topology is the DragonFly+ topology. In DragonFly+, nodes are organized into groups, and these groups are interconnected by global lines. Each group consists of several nodes, which are typically physically close to each other. The groups are then connected to each other through global lines. In this topology, communication can occur within a group or between groups. When nodes within a group need to communicate, they can directly exchange data through internal links within their group. On the other hand, if nodes in different groups wish to communicate, they send the data through the global lines connecting their respective groups.

[0078] The DragonFly+ topology offers advantages in terms of scalability and performance. By organizing nodes into groups, it reduces the overall communication overhead, as nodes within a group can communicate at lower latencies. Moreover, the global lines connecting the groups ensure efficient communication between different groups, enabling high-speed data transfer.

[0079] In a preferred configuration of the expandable computing system, the first network is configured with a network topology that corresponds to at least one group (e.g., a group in a DragonFly+ network) of a network topology defining groups and global lines connecting these groups, such as the DragonFly+ network. The interfaces for expanding the cluster module then provide external communication for at least one global line.

[0080] In another preferred configuration of the expandable computing system, the second network is configured with a network topology that corresponds to at least one group (e.g., a group in a DragonFly+ network) of a network topology defining groups and global lines connecting these groups, such as the DragonFly+ network. The interfaces for expanding the booster module then provide external communication for at least one global line. Like the first network, this configuration organizes the booster nodes into groups and establishes global lines for communication between these groups. Some examples of network topologies that can be used for the first and the second network include, a mesh topology, a tree topology or a hypercube topology.

[0081] Configuring the first and or the second network with a network topology that includes groups and global lines provides benefits such as improved group communication, scalability, fault tolerance, reduced network congestion, and support for parallel processing. These aspects enhance the performance, efficiency, and flexibility of the expandable computing system, particularly in high-performance computing environments focused on Al-related tasks.

[0082] In another preferred embodiment of the expandable computing system, a management node is included to provide access to management functions provided by software running on the system. This management node serves as a centralized control point for overseeing and managing the entire expandable computing system.

[0083] The management node is typically a dedicated computer or server that interacts with the software running on the expandable computing system. It allows administrators or users to monitor and control various aspects of the system's operation. This can include tasks such as resource allocation, workload management, performance monitoring, system configuration, software updates, and overall system health assessment. The inclusion of a management node in the expandable computing system provides valuable functionality for remote monitoring and control, resource allocation, system configuration, performance monitoring, software updates, and scalability management. These aspects enhance the overall efficiency, flexibility, and manageability of the system, supporting smooth operation and effective administration in advanced high-performance computing environments.

[0084] In the preferred embodiment, the expandable computing system is communicationally connected with another expandable computing system through the respective interfaces used for expanding the cluster module and the booster module. This connection allows for the integration and collaboration of multiple expandable computing systems, creating a larger and more powerful computing infrastructure.

[0085] The communication between the expandable computing systems can be established using dedicated high-speed network connections, such as fiber optic links or high-performance interconnect technologies like InfiniBand. These connections provide low-latency and high- bandwidth communication channels, enabling efficient exchange of data and coordination between the interconnected systems. Connecting multiple expandable computing systems through the interfaces for expanding the cluster module and booster module offers advantages such as improved scalability and consolidated management. This interconnected infrastructure leverages the strengths of individual systems while providing the possibility to expand the system to a more powerful computing environment for high-performance computing, particularly for Al-related tasks.

[0086] In a preferred configuration, the expandable computing system is set up so that the first network is configured to form a complete network topology when a predetermined number of expandable computing systems are communicationally connected. This topology thereby formed is preferably a high-speed network topology, such as a DragonFly network topology or a DragonFly+ network topology.

[0087] The Dragonfly network topology connects compute elements in groups through a full graph. Inner-group structures like a full graph, generalized hypercube, or Fat-Tree can be used. Traditional Dragonfly uses the full graph, while the innovative Dragonfly+ (DF+) with support from InfiniBand employs the Fat-Tree. Dragonfly+ is more scalable, accommodating more hosts with the same switch radix. It offers improved worst-case throughput and better switch buffer utilization compared to traditional Dragonfly. Configuring the expandable computing system with a DragonFly or DragonFly+ network topology forming a complete network when a predetermined number of systems are interconnected offers advantages such as minimal network congestion, enhanced performance, and efficient routing even when the system is fully expanded. Brief Description of the Drawings

[0088] Fig. 1 provides a visual representation of an expanded high-performance computing system with system and module expansions.

[0089] Fig. 2 illustrates a front view of a rack, specifically configured to house an expandable high- performance computing system.

[0090] Fig. 3 depicts a layout of five connected expandable computing systems.

[0091] Detailed Description of the Invention

[0092] Fig. 1 provides a visual representation of an expanded high-performance computing system with system and module expansions. The diagram shows a first expandable computing system 102 consisting of a booster module 104, a cluster module 106, and a storage module 108. These modules are communicatively connected through a communications infrastructure (not shown).

[0093] The first expandable computing system 102 functions as the central unit of the system, driving the overall high-performance computing capabilities. The booster module 104 is an integral component of the expandable computing system 102, responsible for enhancing the system's performance. It typically comprises dedicated booster nodes equipped with specialized hardware accelerators, such as GPUs, FPGAs or ASICs. These accelerators expedite Al-related computations and optimize resource utilization.

[0094] The cluster module 106 is another crucial building block of the expandable computing system 102. It consists of computation nodes that work in tandem to execute various computational tasks. These computation nodes can be powerful processors, such as multi-core CPUs, capable of executing complex calculations and algorithms.

[0095] The storage module 108 forms an essential part of the expandable computing system 102, providing a centralized storage infrastructure. It securely stores and manages the data required for computation and analysis. The storage module can utilize various storage technologies, such as solid-state drives (SSDs), hard disk drives (HDDs), or distributed storage solutions like network-attached storage (NAS) or storage area networks (SANs). The diagram does not explicitly depict the communications infrastructure, but it serves as the backbone for connecting the booster module 104, cluster module 106, and storage module 108. The communications infrastructure typically includes high-speed interconnect technologies like InfiniBand or Ethernet (not shown), enabling efficient data transfer and communication between the modules.

[0096] A second expandable computing system 110 is depicted, which expands the first expandable computing system 102. Like the first system 102, the second expandable computing system 110 consists of a booster module 104a, a cluster module 106a, and a storage module 108a. These modules 104a, 106a, 108a are communicatively connected to each other through a communications infrastructure (not shown). Further, the booster module 104a of the second expandable computing system 110 is communicatively connected to booster module 104 of the first expandable computing system 110. The cluster module 106a of the second expandable computing system 110 is communicatively connected to cluster module 106 of the first expandable computing system 110. The storage module 108a of the second expandable computing system 110 is communicatively connected to storage module 108 of the first expandable computing system 110.

[0097] The second expandable computing system 110 functions independently but can integrate and collaborate with the first expandable computing system 102, enabling expanded computing capabilities and resource sharing.

[0098] The booster module 104a within the second expandable computing system is equipped with specialized hardware accelerators, such as GPU, FPGA or ASIC, which enhance the system's performance for Al-related tasks. These accelerators expedite computations and optimize resource utilization, like the booster module 104 in the first expandable computing system 102.

[0099] The cluster module 106a of the second expandable computing system consists of computation nodes that work collaboratively to execute various computational tasks. These nodes can include powerful processors, such as multi-core CPUs, capable of handling complex calculations, like the cluster module 106 in the first expandable computing system 102.

[0100] The storage module 108a in the second expandable computing system serves as a centralized storage infrastructure for securely managing the data required for computation and analysis. This storage module can incorporate various technologies, such as solid-state drives (SSDs), hard disk drives (HDDs), or distributed storage systems such as network- attached storage (NAS) or storage area networks (SANs), like the storage module 108 in the first expandable computing system 102.

[0101] While not shown in the diagram, the communications infrastructure facilitates seamless communication and data transfer between the modules in the second expandable computing system 110, and to the respective modules in the first expandable computing system 102. This infrastructure typically includes high-speed interconnect technologies like InfiniBand or Ethernet, ensuring efficient and reliable communication.

[0102] Further, Fig. 1 also depicts additional expansions within the computing system 100. The first expandable computing system 102 can not only be expanded with the second expandable computing system 110 but also allows for individual module expansions. Each module can be independently expanded with respective modules, namely, booster module 104b, cluster module 106b, and storage module 108b.

[0103] These expanded modules are communicatively connected through a communications infrastructure (not shown). The communications infrastructure facilitates data transfer and communication between the expanded modules within each module and between individual modules, such as modules of the first expandable computing system 102 and the second expandable computing system 110.

[0104] The communications infrastructure, though not visible in the drawing, serves as the communication backbone. It enables seamless data transfer and coordination between the expanded modules and across the entire expandable computing system as visualized with the grouping of booster modules 112, the grouping of cluster modules 114 and the grouping of storage modules 116. This allows an expansion within the computing system 100 beyond the integration of the first expandable computing system 102 with the second expandable computing system 110, namely, the expansions of the individual modules, including booster module, cluster module, and storage module, with respective counterparts, along with the communicative connections between them. These expansions enhance the system's overall capabilities, offering customization, scalability, and increased performance in a flexible and adaptable high-performance computing environment.

[0105] Fig. 2 illustrates a front view of a rack 200, specifically configured to house an expandable high-performance computing system. The rack serves as the outer container or chassis that accommodates various components, such as rackmount servers, switches, Power Distribution Units (PDUs), and cabling.

[0106] The rack is designed to provide infrastructure for the compute system, including features like warm water-cooling solutions and power supply management. These elements ensure the efficient operation and performance of the computing system within the rack 200.

[0107] The drawing depicts a rack with 42 slots, referred to as bays, where each bay has a height of 1 U. In the context of rackmount cases, 1U corresponds to 1.75 inches of vertical space. The 42 bays allow for the installation of various rack mount cases, accommodating the compute system's hardware.

[0108] The rack is equipped with 25 rack mount cases 204, each housing four GPU-CPU processors, such as Nvidia's Grace Hopper Superchips, including local memory and corresponding network interface cards. These rack mount cases deliver high-performance computing capabilities with advanced processing power and memory resources.

[0109] The network interface cards of each rack mount case are connected through multiple switches provided by a 2U switch unit 206. The switches enable efficient and reliable network connectivity between the rack mount cases, facilitating data transfer and communication within the high-performance computing system.

[0110] The combined ensemble of the rack mount cases 204 and the switch unit 206 forms the booster module within the expandable computing system.

[0111] Further shown are the specific components provided in the rack 200 to form a cluster module in an expandable computing system. The cluster module comprises five rack mount cases 208 and a 2U switch unit 214.

[0112] The rack mount cases 208 are designed to fit within the rack and house the key components of the cluster module. Each rack mount case accommodates four CPU multi-core processors, local memory, and four network interfaces. The CPU multi-core processors contribute significant processing power to the cluster module, allowing for intensive computations.

[0113] The network interfaces are directly connected (not shown) to a part of the 2U switch unit 214.

[0114] The switch unit forms the backbone of the cluster module, facilitating communication and data transfer between the rack mount cases. The switch unit comprises networking hardware and ports that enable seamless connectivity and efficient network operations within the cluster.

[0115] Further, in the lower part of the rack two 4U (7-inch high) storage units 216, 218 are provided. These storage units serve as the implementation of the storage module within the expandable computing system. The storage units 216, 218 are dedicated components specifically designed to provide storage capabilities for the system. They are utilized for securely storing and managing the data required for computation and analysis. The 4U height denotes their vertical space within the rack enclosure. These storage units 216, 218 are interconnected through a third switch unit 220. The switch unit acts as an interconnect, enabling seamless communication and data transfer between the booster module and the cluster module within the expandable computing system, by connecting to the switch unit 206 and switch unit 214. This interconnect facilitates efficient coordination and cooperation between the computational and storage elements of the system.

[0116] A central aspect of the expandability of the expandable computing system is the provision of the interfaces for connecting additional racks housing expandable computing systems. Two interfaces are shown. Interface 218 is provided by the switching unit 206. It facilitates the connection between the booster module of the expandable computing system and one or more booster modules in further expandable computing systems (not shown). The interface 218 consists of a predetermined number of individual connections, denoted by 'n'. These connections allow for the expansion and integration of booster modules across multiple racks.

[0117] Interface 220 is provided by the switching unit 214. It enables the connection between the cluster module of the expandable computing system and one or more cluster modules in additional expandable computing systems. Like Interface 218, Interface 220 also comprises a predetermined number of individual connections, represented by 'm'. These connections enable the expansion and integration of cluster modules in multiple racks.

[0118] These interfaces serve as essential components for the expandability of the overall computing system, allowing for the connection and collaboration of various modules across different racks of expandable computing systems. The switching unit 206 can be exemplified by a high-performance network switch capable of providing fast and reliable connections. Interface 218 can utilize industry-standard protocols such as Ethernet or InfiniBand to establish communication between the expandable computing systems. The switching unit 214 can also be a high-performance network switch with similar capabilities. Interface 220 may utilize the same or different protocols as Interface 218, depending on the specific requirements and technologies employed in the expandable computing systems.

[0119] Additionally, a management node 222 is provided as part of a 2U unit. The management node serves as a central control point for overseeing and managing various aspects of the expandable computing system. It provides access to management functions and software running on the system, allowing administrators or users to monitor and control system operations, resource allocation, performance, and other management tasks.

[0120] In Fig. 3 a layout 300 of five connected expandable computing systems is depicted. The layout showcases the network configuration and connections between the systems.

[0121] A first network 301 is depicted, representing the network responsible for connecting the booster nodes within the expandable computing systems. The first network follows the DragonFly+ network topology, which in the shown configuration has five groups 302, 304, 306, 308, 310 connected according to this configuration.

[0122] Each group 302, 304, 306, 308, 310 is represented by a distinct region in the layout. Within each group, individual booster nodes are connected, exemplified by circles extending from the groups, e.g., 312. These booster nodes provide specialized processing capabilities and contribute to the overall computational power of the expandable computing systems.

[0123] In the layout shown, there are five groups representing the DragonFly+ network topology, but it is important to note that the number of groups can vary. Depending on the system requirements, the number of groups can range from 20 to 64 groups, allowing for further scalability and expansion.

[0124] In each group, there are "global connections" that link to each of the other groups. These connections are represented by lines connecting the groups, e.g., line 314, ensuring efficient communication and data transfer across the five expandable computing systems.

[0125] In this specific configuration, each expandable computing system, mounted in its respective rack as shown in Fig. 2, includes one group of the DragonFly+ network. The interfaces provided by each expandable computing system, implemented as switch connections within the rack-mounted system, are represented by small circles, such as interface 316. These interfaces facilitate connections between the systems and contribute to the overall interconnectivity. The four pairs of dotted lines in the drawing outline the boundaries of the five individual systems, clearly demarcating the separate configurations. Between each pair of dotted lines, network cables, either copper-based or optical, establish the connections between the racks. Those global lines which continue from the group switch but do not have a circle (interface) placed on a dotted line do not have a connection to the respective rack. They just passing by.

[0126] There is a corresponding second network 317, representing the network responsible for connecting the cluster nodes within the expandable computing systems. The second network also follows the DragonFly+ network topology, which is in the shown configuration has five groups 322, 324, 326, 328, 330 connected according to this configuration.

[0127] Each group 322, 324, 326, 328, 330 is represented by a distinct region in the layout. Within each group, individual cluster nodes are connected, exemplified by circles extending from the groups, e.g., 332. These cluster nodes provide specialized processing capabilities and contribute to the overall computational power of the expandable computing systems.

[0128] In the layout shown, there are five groups representing the DragonFly+ network topology, but it is important to note that the number of groups can vary. Depending on the system requirements, the number of groups can range from 20 to 64 groups, allowing for further scalability and expansion.

[0129] In each group, there are "global connections" that link to each of the other groups. These connections are represented by lines connecting the groups, e.g., line 334, ensuring efficient communication and data transfer across the five expandable computing systems.

[0130] In this specific configuration, each expandable computing system, mounted in its respective rack as shown in Fig. 2, includes one group of the DragonFly+ network. The interfaces provided by each expandable computing system, implemented as switch connections within the rack-mounted system, are represented by small circles, such as interface 336. These interfaces facilitate connections between the systems and contribute to the overall interconnectivity. Between each pair of dotted lines, network cables, either copper-based or optical, establish the connections between the racks. Those global lines which continue from the group switch but do not have a circle (interface) placed on a dotted line do not have a connection to the respective rack. Each expandable computing system within its rack includes a booster module group, such as group 302, a cluster module group 322, and a storage module 340, for the first system.

[0131] There are connections between the respective groups of the booster module and the cluster module within each expandable computing system, illustrated by slightly thicker lines such as line 318, for the first system.

[0132] Additionally, there is a connection between the respective group of the cluster module and the respective storage module within each expandable computing system, represented by a line, such as line 349.

[0133] In this specific configuration, the first expandable computing system is mounted in rack 350 and includes booster module group 302, connected to cluster module group 322 through line 318, and connected to storage module 340 through line 349. This applies accordingly to the connected expandable computing systems in racks 352, 354, 356, 358.

[0134] It is important to note that in the present configuration each expandable computing system provides four interface connections from the booster module to other booster modules of other systems (for global lines) and four interface connections from the cluster module to other cluster modules (for global lines).

[0135] In the shown full configuration, each expandable computing system is interconnected with the otherwise identical systems, forming a complete high-performance computing system 360, where the network of all booster modules forms a complete DragonFly+ network topology and where the network of all cluster modules forms a complete DragonFly+ network topology. This configuration maximizes performance and facilitates highly efficient data transfer, communication, and collaboration among the distributed computing systems. The complete high-performance computing system 360 consists of multiple expandable computing systems distributed across racks 350, 352, 354, 356, 358.

Claims

Claims1. An expandable computing system for high performance computing, particularly for Al related tasks, comprising: a cluster module (106) including a plurality of computation nodes (332), a booster module (104) including plurality of booster nodes (312), a storage module (108), a communications infrastructure (301 , 317, 318, 349) communicationally connecting the plurality of computation nodes (332), the plurality of booster nodes (312) and said storage module (108), wherein said communications infrastructure (301 , 317, 318, 349) includes at least one interface (336) for expanding said cluster module (106) and at least one interface (316) for expanding said booster module (104).

2. The expandable computing system according to claim 1 , characterized in that said communications infrastructure (301 , 317, 318, 349) includes a first communications network (317) communicationally connecting said plurality of computation nodes (332), a second communications network (301) communicationally connecting said plurality of booster nodes (312), and an interface communicationally (318) connecting said first communications network (317) and said second communications network (301), wherein said at least one interface (336) for expanding said cluster module (104) is provided by said first communications network (317) and said at least one interface (316) for expanding said booster module (106) is provided by said second communications network (301).

3. The expandable computing system according to claim 2, further comprising a third communications (318) network communicationally connecting said first communications network (317) and said second communications network (301).

4. The expandable computing system according to one of the preceding claims, characterized in that said storage module (108, 340) is communicationally connected with said first communications network (317).

5. The expandable computing system according to claim 3, characterized in that said storage module (108, 340) is communicationally connected with said third communications network (318).

6. The expandable computing system according to one of the preceding claims, wherein said first network (317) has the network topology that corresponds to at least one group of a network topology defining groups and global lines connecting said groups, wherein said at least one interface (336) for expanding said cluster module (104) provides external communication for at least one global line.

7. The expandable computing system according to one of the preceding claims, wherein said second network (301) has the network topology that corresponds to at least one group of a network topology defining groups and global lines connecting said groups, wherein said at least one interface (316) for expanding said booster module (106) provides external communication for at least one global line.

8. The expandable computing system according to one of the preceding claims, further comprising a management node (222) for accessing management functions provided by software running on said expandable computing system (100).

9. The expandable computing system according to one of the preceding claims, characterized in that at least a second expandable computing systems (110) is communicationally connected with said expandable computing system (100) via the respective ones of said at least one interface (336) for expanding said cluster module (104) and said at least one interface (316) for expanding said booster module (104).

10. The expandable computing system according to claim 9, if reference is made to claim 2, wherein said first network (317) is configured such that if a predetermined number of said expandable computing systems (350, 352, 354, 356, 358) are communicationally connected, the first networks (317) of all said predetermined number of said communicationally connected expandable computing systems (350, 352, 354, 356, 358) together form a complete network topology, such as a DragonFly network topology or a Dragon Fly+ network topology.

11. The expandable computing system according to claim 9 or 10, if reference is made to claim 2, wherein said second network (301) is configured such that if a predetermined number of said expandable computing systems (350, 352, 354, 356, 358) are communicationally connected, the first networks (301) of all said predetermined number of said communicationally connected expandable computing systems (350,352, 354, 356, 358) together form a complete network topology, such as a DragonFly network topology or a Dragon Fly+ network topology.

Citation Information

Patent Citations

  • Computer Cluster Arrangement for Processing a Computation Task and Method for Operation Thereof

    US20130282787A1