Server system, resource scheduling method for server system, and chip and chiplet
By introducing a multi-dimensional computing resource pool, a data storage resource pool, and a control module into the server system, and utilizing a cache consistency bus and a switching module to achieve multi-dimensional computing power collaboration, the problem of poor scalability of multi-dimensional computing power fusion schemes is solved, and the scalability of computing power and system computing performance are improved.
Patent Information
- Application Number
- PCT/CN2025/083537
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2025-03-19
- Publication Date
- 2025-11-06
AI Technical Summary
Existing multi-computing power fusion solutions suffer from poor scalability of computing power, making it difficult to achieve collaborative computing across multiple instruction set architectures.
By introducing a multi-dimensional computing resource pool, a data storage resource pool, a switching module, and a control module into the server system, and utilizing a cache consistency bus and a switching module to achieve multi-dimensional computing power collaboration, the control module performs computing power scheduling and dynamic allocation of data storage resources. By adopting hardware decoupling and elastic combination design, the pooling of multi-dimensional computing units and system resources is realized.
It improves the scalability of computing power, enables flexible combination and collaborative computing of diverse computing resources, meets diverse application needs, and enhances system computing performance and efficiency.
Smart Images

Figure CN2025083537_06112025_PF_FP_ABST
Abstract
Description
Server system, resource scheduling method of server system, chip and core
[0001] Cross-reference to related applications
[0002] The present application claims priority from the Chinese patent application No. 202410538814.4 filed on April 30, 2024, and entitled "Server system, resource scheduling method of server system, chip and core", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] Embodiments of the present application relate to the field of computers, in particular to a server system, a resource scheduling method of the server system, a chip and a core. BACKGROUND
[0004] Computing power refers to the ability of a computer device or a computing / data center to process information, and more broadly, it can refer to computing resources, such as general-purpose computing power, heterogeneous computing power. In related technologies, in order to fuse multiple heterogeneous computing power, a general-purpose computing power core is usually used as a general-purpose processing unit and a scheduling core, and a heterogeneous computing power core is used as a slave core to realize acceleration computing for specific applications. These heterogeneous fusion processors adopt a master-slave working mode, and the heterogeneous computing power slave cores have the same architecture, and through a unified operating system, a given workload is scheduled.
[0005] However, when multiple computing power cores fuse more heterogeneous computing power units, there are parts of instruction sets that are strongly related within these computing power units, and the difference between the instruction sets of the cores is relatively large, making it difficult to realize collaborative computing of multiple instruction set architecture cores. As can be seen, the multi-computing power fusion scheme in related technologies has the problem of poor scalability of computing power. SUMMARY
[0006] Embodiments of the present application provide a server system, a resource scheduling method of the server system, a chip and a core, to at least solve the problem of poor scalability of computing power in the multi-computing power fusion scheme in related technologies.
[0007] According to an aspect of the embodiments of the present application, a server system is provided, comprising: a multi-element computing resource pool, a data storage resource pool, a switching module and a control module, wherein the multi-element computing resource pool comprises a general computing resource pool and a heterogeneous computing resource pool, the general computing resource pool comprises a group of general computing units, the heterogeneous computing resource pool comprises a group of heterogeneous computing units, the multi-element computing resource pool is connected with the data storage resource pool through a cache coherence bus inside the server system via the switching module; the data storage resource pool contains data storage resources allowed to be shared by multi-element computing resources in the multi-element computing resource pool, wherein the multi-element computing resources in the multi-element computing resource pool comprise the general computing units in the general computing resource pool and the heterogeneous computing units in the heterogeneous computing resource pool; the control module is configured to perform computing power scheduling on the multi-element computing resources in the multi-element computing resource pool and dynamically allocate the data storage resources in the data storage resource pool to the multi-element computing resources in the multi-element computing resource pool.
[0008] According to another aspect of the embodiments of the present application, a resource scheduling method of a server system is also provided, comprising: the server system comprises: a multi-element computing resource pool, a data storage resource pool, a switching module and a control module, wherein the multi-element computing resource pool comprises a general computing resource pool and a heterogeneous computing resource pool, the general computing resource pool comprises a group of general computing units, the heterogeneous computing resource pool comprises a group of heterogeneous computing units, the multi-element computing resource pool is connected with the data storage resource pool through a cache coherence bus inside the server system via the switching module; the data storage resource pool contains data storage resources allowed to be shared by multi-element computing resources in the multi-element computing resource pool, wherein the multi-element computing resources in the multi-element computing resource pool comprise the general computing units in the general computing resource pool and the heterogeneous computing units in the heterogeneous computing resource pool; the method comprises: performing computing power scheduling on the multi-element computing resources in the multi-element computing resource pool and dynamically allocating the data storage resources in the data storage resource pool to the multi-element computing resources in the multi-element computing resource pool by the control module.
[0009] According to still another aspect of the embodiments of the present application, a chip is provided, comprising: multi-element computing resources, on-chip memory and a control core interconnected through an on-chip cache coherence bus of the chip, wherein the multi-element computing resources comprise general computing resources and heterogeneous computing resources, the general computing resources comprise a group of general processors, the heterogeneous computing resources comprise a group of heterogeneous accelerator processors; the on-chip memory is allowed to be shared by the general processors in the general computing resources and the heterogeneous accelerator processors in the heterogeneous computing resources; the control core is configured to perform computing power scheduling on the general processors in the general computing resources and the heterogeneous accelerator processors in the heterogeneous computing resources.
[0010] According to another aspect of the embodiments of the present application, a core particle is provided, which is interconnected by a set of cores and a memory controller in the core particle through an on-chip cache coherence bus, wherein the set of cores includes a control core and a plurality of processor cores, and the plurality of processor cores includes a set of general-purpose processor cores and a set of heterogeneous accelerated processor cores, the cores in the set of cores have private caches and share a designated cache and a memory controlled by the memory controller; the control core is configured to perform computing power scheduling on the general-purpose processor cores in the set of general-purpose processor cores and the heterogeneous accelerated processor cores in the set of heterogeneous accelerated processor cores.
[0011] According to the present application, the server internal cache coherence bus interconnection and switching are used as the core, the computing power and resources are pooled, the multi-computing units and system resources are decoupled (the switching module and the control module can be replaced as needed) and flexibly combined, and the server system includes a multi-computing resource pool, a data storage resource pool, a switching module, and a control module. The multi-computing resource pool includes a general-purpose computing resource pool and a heterogeneous computing resource pool, the general-purpose computing resource pool includes a set of general-purpose computing units, the heterogeneous computing resource pool includes a set of heterogeneous computing units, the multi-computing resource pool is connected to the data storage resource pool through the cache coherence bus inside the server system via the switching module; the data storage resource pool contains data storage resources that can be shared by the multi-computing resource pools (such as general-purpose computing units and heterogeneous computing units) in the multi-computing resource pool. At the same time, the control module is used to perform computing power scheduling on the multi-computing resources in the multi-computing resource pool and dynamically allocate the data storage resources in the data storage resource pool to the multi-computing resources in the multi-computing resource pool. Since the multi-computing power is cooperated through the cache coherence bus inside the server system via the switching module, and the expansion of computing power can be performed with the expansion capability of the switching module, the technical effect of improving the scalability of computing power can be achieved, thereby solving the problem of poor scalability of multi-computing power in related technologies and achieving the technical effect of improving the scalability of computing power. BRIEF DESCRIPTION OF DRAWINGS
[0012] FIG. 1 is a hardware structure block diagram of an optional server device according to an embodiment of the present application.
[0013] FIG. 2 is a structure block diagram of an optional server system according to an embodiment of the present application.
[0014] FIG. 3 is a structure block diagram of another optional server system according to an embodiment of the present application.
[0015] FIG. 4 is a schematic diagram of optional memory decoupling and pooling of a one-machine multi-core system according to an embodiment of the present application.
[0016] Figure 5 is a schematic diagram of an optional one-chip multi-core system multi-host shared storage resource pooling according to embodiments of the present application.
[0017] Figure 6 is a schematic diagram of an optional one-chip multi-core system top-level architecture according to embodiments of the present application.
[0018] Figure 7 is a schematic diagram of an optional one-chip multi-core single-node server according to embodiments of the present application.
[0019] Figure 8 is a schematic diagram of an optional multi-core computing module according to embodiments of the present application.
[0020] Figure 9 is a schematic diagram of an optional one-chip multi-core standardized interface according to embodiments of the present application.
[0021] Figure 10 is a schematic diagram of an optional one-chip multi-core multi-node modular server according to embodiments of the present application.
[0022] Figure 11 is a schematic diagram of an optional system-level one-chip multi-core fusion architecture server according to embodiments of the present application.
[0023] Figure 12 is a schematic diagram of an optional memory resource pool top-level architecture according to embodiments of the present application.
[0024] Figure 13 is a schematic diagram of an optional CXL bus switching unit according to embodiments of the present application.
[0025] Figure 14 is a schematic diagram of an optional memory resource management engine according to embodiments of the present application.
[0026] Figure 15 is a schematic diagram of a resource management module asset information display and management module according to embodiments of the present application.
[0027] Figure 16 is a schematic diagram of a management system dynamically allocating and scheduling resources according to embodiments of the present application.
[0028] Figure 17 is a schematic diagram of an optional management control system according to embodiments of the present application.
[0029] Figure 18 is a schematic diagram of an optional management software according to embodiments of the present application.
[0030] Figure 19 is a schematic diagram of an optional one-chip multi-core system fusion management according to embodiments of the present application.
[0031] Figure 20 is a schematic diagram of an optional multi-core processor DVFS voltage regulation method according to embodiments of the present application.
[0032] Figure 21 is a schematic diagram of an optional multi-level power consumption management mechanism according to embodiments of the present application.
[0033] Figure 22 is a schematic diagram of an optional one-machine multi-core multi-core cooperative and unified scheduling operating system architecture according to an embodiment of the present application.
[0034] Figure 23 is a schematic diagram of an optional one-machine multi-core system resource scheduling design scheme according to an embodiment of the present application.
[0035] Figure 24 is a schematic diagram of an optional one-machine multi-core system resource scheduling allocation process according to an embodiment of the present application.
[0036] Figure 25 is a schematic diagram of an optional two-dimensional grid structure according to an embodiment of the present application.
[0037] Figure 26 is a schematic diagram of an optional three-dimensional grid structure according to an embodiment of the present application.
[0038] Figure 27 is a schematic diagram of an optional resource scheduling method of a server system according to an embodiment of the present application.
[0039] Figure 28 is a structural block diagram of an optional chip according to an embodiment of the present application.
[0040] Figure 29 is a structural block diagram of another optional chip according to an embodiment of the present application.
[0041] Figure 30 is a structural block diagram of an optional core particle according to an embodiment of the present application.
[0042] Figure 31 is a structural block diagram of another optional core particle according to an embodiment of the present application. DETAILED DESCRIPTION
[0043] Hereinafter, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with embodiments.
[0044] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and in the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0045] As a new generation of productivity for the development of digital economy, computing power can inject new momentum for promoting the digital transformation of various industries. Computing power refers to the ability of computer equipment or computing / data center to process information, and more broadly, it can refer to computing resources, such as general-purpose computing power, heterogeneous computing power, etc. Here, general-purpose computing power can be CPU (Central Processing Unit), and heterogeneous computing power can be GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), etc.
[0046] With the accelerated development of core technologies of digital economy such as cloud computing, artificial intelligence, and big data, new scenarios and demands are emerging, driving the rapid growth of computing power scale and application scenarios. The computing industry is facing the trend and challenge of diversification, massification, and ecological decentralization. Diversified scenarios require more efficient, diverse, and flexible computing power support. However, there is still a huge gap from chip to computing power, and the value of multi-algorithm (such as general-purpose computing power, heterogeneous computing power, etc.) has not been fully released. In the face of the challenge of multi-algorithm, co-constructing an algorithmic industry system and releasing the value of multi-algorithm are important technical directions for the construction of new algorithmic infrastructure.
[0047] There are great differences in technical solutions, system architecture, software platform, hardware equipment, and service guarantee among different computing power platforms. For example, X86 architecture is the mainstream CPU architecture of data centers, with a mature ecosystem, complex instruction sets, and suitable for complex infrastructure environments of heavy-load businesses. However, the increasingly complex instruction system reduces CPU efficiency and versatility, and increases system complexity. For example, ARM (Advanced RISC Machines) architecture CPU has more cores per unit area, suitable for low-power, high-concurrency computing businesses (such as cloud gaming, distributed databases, containerized deployment, distributed storage, etc.). For example, Power (Performance Optimization With Enhanced Reduced Instruction Set Computing) architecture has performance advantages in high-end server fields, leading in software and hardware reliability, availability, and maintainability, and suitable for high-performance computing and other key host business scenarios.
[0048] In view of the characteristics of the CPU platform and the diversified application scenarios, the current response measures at the hardware design level are mainly to customize servers of different specifications and different models. However, each platform CPU will develop multiple generations, and each generation of CPU has multiple product forms such as single, double, and four-way, and has multiple configurations of memory, storage, heterogeneous accelerators, and manager devices. The design of platform design combination is numerous. The hardware tightly coupled design method brings many difficulties to the resource utilization, management and operation and maintenance of the data center, resulting in problems such as energy waste, upgrade and expansion cost, and affecting the full release of system performance.
[0049] On the other hand, with the development of computer applications such as artificial intelligence and big data, the demand for computing power is growing rapidly, and general-purpose CPUs encounter more and more performance bottlenecks in mass processing calculations and mass data / pictures, such as low computing parallelism, low bandwidth, high latency, etc. The traditional server architecture with CPU as the computing core has been unable to meet today's needs, so the design of the new generation of server architecture faces new challenges. In order to efficiently process various computing tasks, different types of heterogeneous acceleration processors such as GPU, FPGA, ASIC, etc. have emerged. The computing architecture is gradually breaking through the architecture system centered on CPU, and the heterogeneous computing architecture of "CPU + heterogeneous accelerator" has become the mainstream server architecture in the current artificial intelligence, high-performance computing and other scenarios. Different heterogeneous acceleration processors have their own characteristics and are suitable for different computing models. Among them, the combination of CPU + GPU is more suitable for general-purpose scenarios and is the current mainstream choice. The combination of CPU + ASIC / FPGA is developing simultaneously due to its excellent performance in special scenarios such as inference training, and showing a hundred flowers blooming situation.
[0050] However, the current heterogeneous computing acceleration chips are facing challenges such as non-uniform hardware system interface, non-uniform interconnection protocol, and incompatible software ecology, which seriously restrict the deployment and application development of heterogeneous computing power infrastructure. The multi-element computing power needs to continue to innovate from "usable" to "good use" and create value.
[0051] In order to improve the system computing power, the related art adopts a homogeneous multi-core design, and meets the computing power requirement through multi-core parallel operation, but the homogeneous multi-core has no advantage in diversified demand scenarios. In the face of diversified application scenarios and the increasing demand for heterogeneous computing power, some processors implement heterogeneous multi-cores with the same instruction set (the multi-cores have different register types, register quantities, register descriptions, etc.), and the multi-cores implement functions such as data flow monitoring, performance monitoring, delay and bandwidth tuning, and power management through dynamic sharing units. The task scheduler statically schedules or dynamically schedules and allocates the processor used by the task according to the performance requirement of each task and the processing capacity of the core. However, these chips still have similar core architectures with the same instruction set, and cannot meet the demand of the multi-heterogeneous and different instruction set architecture core (processor core) collaborative work scenario.
[0052] In terms of multi-heterogeneous fusion, the related art usually adopts a general-purpose processor (CPU) as a general-purpose processing unit and a scheduling core, and a heterogeneous accelerator core as a slave core to implement application-oriented acceleration calculation. These heterogeneous fusion processors adopt a master-slave working mode, and the heterogeneous computing power slave core has the same architecture. Through a unified operating system, a given workload is scheduled, and the calculation performance and efficiency can be optimized. However, when more heterogeneous computing power units are fused in the multi-heterogeneous core, such as GPUs, FPGAs, ASIC field special accelerators of different manufacturers, there are parts with strong correlation between instruction sets in these computing chips, and the instruction sets between cores are quite different, making it difficult to realize collaborative computing of multiple instruction set architecture cores.
[0053] In addition, the related art is increasingly fragmented in terms of field / scene, and the engine architecture of heterogeneous processors of different manufacturers is increasingly diversified, and it is also necessary to build an efficient, standard, and open ecological system to fully utilize multi-heterogeneous computing and multi-instruction set processors, meet diversified application requirements, and improve system efficiency.
[0054] In order to at least partially solve the above technical problems, in the embodiment, a high-performance and high-efficiency server is provided through "system-component-chip-software" multi-level collaborative design with application as the guide and system architecture design as the core, unified management and scheduling of multi-heterogeneous computing power are realized, the system advantage of multi-heterogeneous computing power is fully played, the computing power is easy to use, and the popularization and application of artificial intelligence technology can be promoted, and the development of computing power can be accelerated. Here, through the above multi-level collaborative design, a multi-core computing fusion architecture can be realized through one machine and multiple cores. Here, the multiple cores can refer to multiple chips, and one machine and multiple cores refer to a server host with multiple chips.
[0055] According to an aspect of the embodiments of the present application, a server system is provided, which can correspond to a server device. FIG. 1 is a hardware structure block diagram of an optional server device according to an embodiment of the present application. As shown in FIG. 1, the server device can include one or more (only one is shown in FIG. 1) processors 101 (the processor 101 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 102 for storing data, wherein the server device can further include a transmission device 103 for communication function and an input and output device 104. Those skilled in the art can understand that the structure of FIG. 1 is only schematic, which does not limit the structure of the server device. For example, the server device can further include more or less components than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1.
[0056] The memory 102 can be configured to store computer programs, for example, software programs of application software and modules, such as a computer program corresponding to the resource scheduling method of the server system in the embodiments of the present application. The processor 101 can execute various function applications and data processing, i.e., implement the above method, by running the computer programs stored in the memory 102. The memory 102 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 102 can further include a memory remotely arranged with respect to the processor 101, which can be connected to the server device through a network. Examples of the network can include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0057] The transmission device 103 is configured to receive or send data via a network. An optional example of the network can include a wireless network provided by a communication service provider of the server device. In one example, the transmission device 103 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 103 can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet in a wireless manner.
[0058] The server system in the embodiments of the present application can be as shown in FIG. 2, which can include a multi-element computing resource pool 201, a data storage resource pool 202, a switching module 203 and a control module 204.
[0059] Here, the multi-computing resource pool 201 is connected with the data storage resource pool 202 via the exchange module 203 through a cache coherence bus (internal bus) inside the server system. The multi-computing resource pool 201 includes a general computing resource pool including a group of general computing units and a heterogeneous computing resource pool including a group of heterogeneous computing units. In some embodiments, the general computing units in the multi-computing resource pool can be central processing units (CPUs), and the heterogeneous computing units in the multi-computing resource pool can include at least one of a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and other types of general computing units or heterogeneous computing units. The cache coherence bus is a special hardware mechanism that can maintain cache data consistency of each processor core in a multi-processor system through the internal cache coherence bus of the server.
[0060] The data storage resource pool 202 contains data storage resources that can be shared by the multi-computing resources in the multi-computing resource pool 201. The data storage resources can be system resources for data storage, can be shared memory, storage devices, and other system resources capable of data storage. In addition to the data storage resources in the data storage resource pool 201, the multi-computing resources in the multi-computing resource pool 201 can contain private data storage resources, such as private caches, which are not limited in this embodiment. Here, the multi-computing resources in the multi-computing resource pool 201 can include general computing units in the general computing resource pool and heterogeneous computing units in the heterogeneous computing resource pool.
[0061] The control module 204 is configured to perform computing power scheduling on the multiple computing resources in the multiple computing resource pool 201, and dynamically allocate data storage resources in the data storage resource pool 202 to the multiple computing resources in the multiple computing resource pool 201. Here, the exchange module 203 and the control module 204 can be obtained by modularizing the server host through a board-level decoupling manner (referring to the decoupling manner adopted at the level of a printed circuit board, and the printed circuit board can be a server mainboard), which can include one or more controllers, such as a management controller. Correspondingly, the multiple computing resource pool 201 can be integrated as a multiple computing module, and the data storage resource pool 202 can be integrated as a data storage module. Through the above-mentioned modular design, the hardware decoupling of heterogeneous computing power can be realized; the multiple computing power collaboration can be realized through the open consistency interconnection and exchange; and the system resources such as data storage resources are shared by heterogeneous cores to realize the computing power and resource pooling. It should be noted that the resource module decoupling is realized by physical form, and each module can be replaced individually as needed. Traditional servers are designed as a whole and cannot be decoupled.
[0062] The exchange module 203 can include a group of exchange units, and the exchange unit can include one or more exchange chips. Different types of data storage resources or other system resources can be connected to the multiple computing resource pool 201 through different exchange units. For example, the data storage resource in the data storage resource pool 202 is a CXL (Compute Express Link) memory, and the corresponding exchange chip can be a CXL exchange chip.
[0063] It should be noted that the multiple computing power collaboration can be realized through the cache consistency bus in the server system via the exchange module 203, and the expansion of computing power can be realized by means of the expansion capability of the exchange module 203, so as to achieve the technical effect of improving the scalability of computing power. The multiple computing module. The control module can schedule at least one of the general computing unit and the heterogeneous computing unit, and the multiple computing resource to which the data storage resource is dynamically allocated can be at least one of the general computing unit and the heterogeneous computing unit.
[0064] The server system in the embodiment includes: a multi-element computing resource pool, a data storage resource pool, a switching module, and a control module. The multi-element computing resource pool includes a general computing resource pool and a heterogeneous computing resource pool. The general computing resource pool includes a group of general computing units, and the heterogeneous computing resource pool includes a group of heterogeneous computing units. The multi-element computing resource pool is connected to the data storage resource pool through the switching module via a cache coherence bus inside the server system. The data storage resource pool contains data storage resources that are allowed to be shared by the multi-element computing resources in the multi-element computing resource pool. The multi-element computing resources in the multi-element computing resource pool include the general computing units in the general computing resource pool and the heterogeneous computing units in the heterogeneous computing resource pool. The control module is configured to perform computing power scheduling on the multi-element computing resources in the multi-element computing resource pool and dynamically allocate the data storage resources in the data storage resource pool to the multi-element computing resources in the multi-element computing resource pool. The problem of poor scalability of computing power in the multi-element computing power fusion solution in the related art is solved, and the scalability of computing power is improved.
[0065] In some example embodiments, as shown in FIG. 3, the server system further includes an I / O (Input / Output) device module 301. The I / O device module 301 is connected to the multi-element computing resource pool 201 through the switching module 203 via a cache coherence bus inside the server system. The I / O device module is also referred to as an I / O module.
[0066] The data storage resource pool 201 can include a memory resource pool and a storage resource pool. The memory resource pool can include a group of shared memories, and the storage resource pool can include a group of storage devices. The memory resource pool is integrated into a memory module, and the storage resource pool is integrated into a storage device module. The switching module 203, the I / O device module 301, the memory module, and the storage device module can all be hardware modules. Different hardware modules can be connected to other modules through corresponding interfaces (which can be standardized interfaces). After the standardized interface, resource module decoupling can be realized in a physical form, and each module can be more conveniently replaced individually as needed.
[0067] The multi-element computing resource pool can be integrated into a multi-element computing module. The multi-element computing module can be a multi-core computing module, which can be a computing module containing multiple chips. Only core components can be reserved inside the multi-core computing module (for example, a CPU and a memory on a server motherboard can be used as core components to constitute the multi-core computing module). The multi-core computing module can be replaced with different processor platforms and can be compatible with different systems such as single-channel and double-channel systems.
[0068] In some embodiments, the multi-element computing module can be connected to other modules in the server system except the multi-element computing module through a set of preset interfaces, including at least one of the following: a switching module, a control module, and different modules. The preset interfaces used for connecting different modules can be the same or different.
[0069] For example, the overall architecture of a one-machine multi-core system design is shown in FIG. 4, which realizes hardware decoupling of heterogeneous computing power; through open consistency interconnection and switching, multi-element computing power collaboration is realized; heterogeneous cores share memory and I / O resources, realizing computing power and resource pooling; a design management and scheduling unit (an example of the aforementioned control module) is used to realize system startup and dynamic scheduling of resources on demand, and a system management unit is used to realize management functions such as system resource, topology, power supply and power consumption management, and state monitoring. At the software level, a unified framework of strong correlation between the underlying hardware is formed, and intercommunication between heterogeneous cores (kernels) is realized at the operating system level, so that the operating system still has a global state under different ISA (Instruction Set Architecture) core running instances; the upper application is not aware of the underlying hardware, and the application can automatically select the optimal mapping and perform computing power allocation. Here, in FIG. 4, the multi-core computing module can include a plurality of computing units (which can include general-purpose computing units or heterogeneous computing units), the switching module can include a plurality of switching units, and the computing units and the switching units are connected through a cache coherence bus (internal bus), and the memory module (global memory module) includes a set of shared memory (fine-grained shared memory), which is connected to the switching units through a memory controller via the cache coherence bus.
[0070] At the system level, memory is a core component of a computer system, and memory capacity and bandwidth directly affect system computing performance. Currently, the local memory of a computer system is limited by parallel buses and system space, and cannot be expanded on demand and dynamically allocated, resulting in the expansion of memory bandwidth and capacity being unable to match the rapid growth of CPU core numbers. With the advent of serial cache coherence buses, low-latency access paths and cache coherence guarantees are provided for multi-core shared memory. System-level shared memory can be based on serial cache coherence buses and high-performance switching to realize low-latency, high-bandwidth, and high-capacity memory resource remote expansion and physical pooling, dynamically allocate remote memory and share fine-grained address space through software definition, and meet the needs of diversified computing power for shared memory through software and hardware adaptation and optimization, reduce the loss of business migration between different computing platforms, and improve overall computing efficiency, as shown in FIG. 5.
[0071] At the system level storage layer, the current host system performance expansion is limited by the system interconnection bandwidth, the performance mismatch between the storage at each level, the low utilization rate of I / O resources, the implementation of storage device resource decoupling in one machine with multiple cores, the realization of storage device resource pooling through high-performance exchange of internal buses and software-defined system design, the removal of the binding relationship between the physical link layer of the storage device and the computing unit, the flexible adjustment of the port configuration and resource allocation path of the exchange network, the fine-grained division of the shared resource pool, and the realization of on-demand elastic allocation and multi-host sharing of storage device resources, as shown in FIG. 6. In FIG. 6, the CPU in the general computing module can be a CPU from different platforms, the switch module can include a group of switch chips (Switch, i.e., SW), and the shared memory and the memory controller (Mem.Ctrl) can be connected using the DDR (Double Data Rate) technology.
[0072] Through the embodiment, the multi-element computing resource, the memory resource, and the storage device resource are modularly designed, and meanwhile, the multi-element computing module is connected to other modules through a standard interface, so that the migration of the multi-element computing resource between different platforms can be improved, and the flexibility of the server system construction can be improved.
[0073] In some example embodiments, the server system further includes: a heat dissipation module configured to dissipate heat from at least part of the devices in the server system; a power supply module configured to supply power to the server system; and a network module configured to connect the server system to a network.
[0074] To improve the security of the server operation, the server system further includes: a heat dissipation module configured to dissipate heat from at least part of the devices in the server system. To ensure the normal operation of the server system, the server system further includes: a power supply module configured to supply power to the server system. To enable the server system to be connected to a network, the server system further includes: a network module configured to connect the server system to a network. Correspondingly, the other modules described above further include at least one of: a heat dissipation module, a power supply module, and a network module.
[0075] Here, the heat dissipation module can be a heat dissipation fan, the server system can include a baseboard management controller, and the heat dissipation fan can be an intelligent peripheral inserted into the baseboard management controller. At least part of the modules in the server system described above can be located on the baseboard management controller.
[0076] For example, as shown in FIG. 7, for a one-machine multi-core single-node server decoupling design, the server host modular design is implemented through the multi-core computing module, the storage module (storage device module), the management module (which can be the aforementioned control module, or another module different from the control module), the I / O device module, and the like. The standardized design and interconnection between different modules are implemented through the replacement of the multi-core computing module to realize the flexible switching of different processor platforms, and the on-demand design, asynchronous upgrade, and agile delivery of the host system. The schematic diagram of the multi-core computing module is shown in FIG. 8, which can be adapted to various processor platforms (i.e., computing platforms, CPU platforms). Here, the CPU and the memory on the server mainboard can be used as the core to constitute the multi-core computing module, and the CPU can be different CPU platforms such as X86, ARM, MIPS (Microprocessor without interlocked piped stages), Power, and the like. The multi-core computing module provides standard interfaces such as PCIe (Peripheral Component Interconnect express), CXL, power supply, management, and the like, and is connected with each unit of the single-node server in FIG. 7.
[0077] Here, the asynchronous upgrade mode of the single-node server can be the replacement of the storage device module, the multi-core computing module, the heat dissipation module, the management module, the I / O device module, and the like. After the standardized definition, compared with the existing server, the one-machine multi-core system architecture can realize the decoupling design of the hardware, and can realize the asynchronous upgrade and on-demand design through the replacement of each unit.
[0078] For the one-machine multi-core elastic architecture, through the hardware decoupling and the standardized interface definition, the modular design of the heterogeneous computing unit and the storage resource (storage device resource), the I / O resource (I / O device resource), and the memory resource is realized, the host system is composed of unified management, network, power supply, and heat dissipation, and the multi-element computing power is flexibly provided by the software-defined system design. The standard definition top-level architecture is shown in FIG. 9. The MPPM (Multi-platform Processing Module) is consistently interconnected with the following modules through the standard interface: SDM (Storage Device Module), IODM (I / O Device Module), GMM (Global Memory Module), and the like, and the resource sharing and dynamic allocation are realized through the SM (Switching Module).
[0079] In addition, as shown in FIG. 9, the UCM (Unified Management Module) is responsible for multi-element computing core scheduling and cooperation. The entire system is integrated through UNM (Unified Network Module), UMM (Unified Management Module), UPM (Unified Power Module), and UTM (Unified Thermal Module) and other basic modules, and the overall system-level one-machine multi-core system architecture is realized.
[0080] Referring to FIG. 9, the interfaces used by the multi-core computing module to connect with each module are respectively: SDMI (Storage Device Module Interface), IODMI (I / O Device Module Interface), GMMI (Global Memory Module Interface), UCMI (Unified Management Module Interface), UMMI (Unified Management Module Interface), and UNMI (Unified Network Module Interface).
[0081] Here, the one-machine multi-core design decouples and flexibly combines multi-element computing modules and system resources at the server system architecture level, realizes differentiated design and integration of underlying hardware through standardized interface definition and system unified management. The system design takes high-performance server internal cache consistency bus interconnection and switching as the core, realizes multi-element computing power cooperative computing and resource on-demand allocation through software-defined system design, realizes multi-element computing power fusion and load balancing through operating system and software cooperative innovation. Through software and hardware cooperative design, chip, component, system, and software fusion and efficient cooperative work are realized. Facing the characteristics of heterogeneous multi-element computing power, the one-machine multi-core system can quickly and efficiently combine multi-element computing power, flexibly adapt and optimize, and fully exert the advantages of multi-element computing power; under the condition that the performance of key devices and components is limited, the system performance is improved through system architecture innovation, and the data center bottleneck problem is alleviated. The above scheme can be diversified to the data center level multi-element heterogeneous fusion modular design.
[0082] It should be noted that the exchange module can be a server internal bus exchange module such as PCIe / CXL, which is set as an extension of CPU resources; the unified control module is responsible for multi-element computing core scheduling and cooperation, which is consistent with the aforementioned management scheduling unit, and the unified control module is standardized.
[0083] Through the embodiment, the multi-element computing module is connected with the heat dissipation module, the power supply module and the network module through a standard interface, which can ensure the normal operation of the multi-element computing module and improve the safety and networking function of the multi-element computing module.
[0084] In some example embodiments, the number of multi-element computing modules is multiple, and the multiple multi-element computing modules are connected with the host of the server system in a card type, and different multi-element computing modules are connected through the cache coherence bus inside the server system. The corresponding server can be a one-machine multi-core multi-node modular server.
[0085] An example of a one-machine multi-core multi-node modular server is shown in FIG. 10, which has the advantages of convenient deployment, flexible expansion, high density, higher performance power ratio, etc. The multi-node modular server form is adapted to multiple computing platforms through card type multi-core computing modules, and the system is integrated through standardized interfaces and high-speed signal backboard design, resource module decoupling, unified management and heat dissipation. The multi-core computing modules can maintain global cache coherence through an open internal cache coherence bus.
[0086] A one-machine multi-core system-level server host is shown in FIG. 11, which faces diversified application requirements, changes the CPU-centered design, and through key technologies such as hardware decoupling and consistent high-speed interconnection, takes the internal bus exchange as the core, decouples and restructures the whole server, realizes the collaborative computing of multiple general-purpose processor platforms and DPU (Data Processing Unit), GPU, FPGA and other heterogeneous acceleration units at the system level, realizes the hardware decoupling, pooling and asynchronous iteration of large-scale computing resources, memory resources, heterogeneous acceleration resources and storage resources. Through software-defined system design, the power resource collaborative dynamic scheduling is realized. The host and resource pool in the whole system are integrated through unified management, heat dissipation and power supply to form a high-performance host with heterogeneous high computing power, I / O resources and memory resources that can be expanded on demand, and flexible scheduling and allocation of resources. The whole system can flexibly provide multi-element computing power in multiple application scenarios, realize application demand-oriented computing under the condition of limited key components and devices, and integrate multi-element computing power through system-level architecture key technology innovation to improve computing performance.
[0087] Through the embodiment, the plurality of multi-element computing modules are connected to the host of the server system in a plug-in manner, and the different multi-element computing modules are connected through the cache consistency bus inside the server system, so that the adaptation capability between the multi-element computing modules and different computing platforms can be improved, and the convenience and reliability of data processing can be improved.
[0088] In some example embodiments, similar to the foregoing embodiments, the data storage resource pool includes a memory resource pool and a storage resource pool, wherein the memory resource pool includes a group of shared memories, and the storage resource pool includes a group of storage devices. The control module is further configured to dynamically allocate the shared memories in the memory resource pool and the storage devices in the storage resource pool to the multi-element computing resources in the multi-element computing resource pool.
[0089] For example, for a one-machine multi-core server system, through system decoupling design and standardized definition, the computing power of a data center with multiple platforms and multiple instruction set heterogeneous cores is fused, resources are flexibly allocated on demand, system unified scheduling and management are implemented, and diversified scene requirements are met. The one-machine multi-core can serve as a standard of data center computing power supply unit, fully releasing the value of multi-element computing power from chip to system design, and supporting sustainable expansion of large-scale data centers.
[0090] Through the embodiment, the memory resources and the storage resources are flexibly allocated on demand by the control module, so that diversified scene requirements can be met.
[0091] In some example embodiments, the server system further includes a memory controller, and the group of shared memories are connected to one or more memory controllers, and each shared memory is connected to only one memory controller. In some embodiments, the group of shared memories can include a group of CXL memories, and the group of CXL memories are memories connected to the memory controllers. Correspondingly, in the switching module, the switching unit connected to the memory controller to which the group of CXL memories belong is a CXL bus switching unit. The group of CXL memories are pooled to obtain a CXL memory pooling system, which is constructed based on a CXL bus type network. To meet the demand of multi-scene large memory applications, the CXL memory pooling system takes the internal CXL bus switching unit as the core, decouples and reconstructs the server storage architecture, realizes memory resource reorganization and asynchronous iteration, and maximizes the expansion of system memory; through software-defined system design, the system topology scheme can be flexibly configured and adjusted, the memory resources can be expanded on demand, the utilization rate of memory resources in the data center can be improved, and the space and energy consumption of the data center storage cabinet can be reduced. The top-level architecture of the CXL memory resource pool can be as shown in FIG. 12.
[0092] In FIG. 12, the memory resource pool unit (i.e., the memory resource pool) is connected to the CXL bus switching unit through the CXL interconnection bus, the CXL bus switching unit is connected to the general computing unit through the CXL interconnection bus, and the general computing unit includes a CPU core and private memory.
[0093] In some embodiments, the CXL memory pooling system is based on the design of the CXL2.0 switch (an example of the switching unit), realizes multi-host sharing of memory resources and flexible allocation of resources. The general computing module can be adapted to various platform servers, and the CXL bus switching module (as shown in FIG. 13) supports the CXL2.0 protocol and provides 32 high-speed interfaces externally, each of which can support 32GT / s, x16 bandwidth, and supports arbitrary configuration of uplink and downlink. The memory resource pool module (i.e., the memory module) can support the CXL2.0 protocol and can externally expand 64 DDR5 (DDR 5 Synchronous Dynamic Random Access Memory) RDIMM (Registered Dual-Inline-Memory-Modules) modules. All units can be interconnected through CDFP (400 Gigabit Dual-Flex Parallel) high-speed cables.
[0094] Exemplarily, a set of specifications of the CXL bus switching module can be as follows: 19-inch 2U standard chassis, depth: 700mm; width: 435mm; height: 87mm, supporting L-shaped bracket mounting; single CXL Switch 256lane (logical channel or channel), 16X16 port (port), configurable uplink and downlink, the whole machine can support 2 CXL Switches, configurable as 16-way uplink and 16-way downlink; the whole machine can output 32 CDFP x16 ports, supporting CXL 2.0; using the CPU module for in-band management, and the baseboard management controller for system management; supporting 2 CRPS (Common Redundant Power Supplies) power supplies, 1+1 redundancy; fan supporting N+2 rotor redundancy, supporting hot plug, supporting power consumption monitoring; external cable: CDFP DAC (Direct Attach Cable) cable interconnection.
[0095] In some embodiments, a set of CXL memories is divided into a plurality of memory partitions, and a memory partition in the plurality of memory partitions allows allocation to a plurality of computing resources in a plurality of computing resource pools. The switching module includes a set of CXL switching chips, wherein the plurality of computing resources in the plurality of computing resource pools are connected to memory controllers to which the set of CXL memories belong via the set of CXL switching chips through a cache coherence bus inside the server system.
[0096] Here, the memory resources in the memory resource pool can be managed by a memory resource management engine, which belongs to the control module or the management module. It should be noted that the control module or the management module can be one or more processor cores on the server, which can be located on the baseboard management controller or independent of the baseboard management controller. Regardless of the form or variation, as long as it can conform to the system architecture in the embodiment.
[0097] In the high-performance CXL switching module, the CXL switching chip realizes the resource decoupling of the memory pooling system, the memory resource management engine uniformly manages the CXL switching chip and provides an interactive interface, so that the CXL devices (CXL memories) in the memory resource pool can be uniformly managed, and the identification, management and dynamic allocation of memory resources are completed. An example of the overall technical scheme of the memory resource management engine can be as shown in FIG. 14. Referring to FIG. 14, the memory resource management engine can realize the following functions: memory resource identification and asset information display; device partitioning and dynamic allocation information.
[0098] For memory resource identification and asset information display, the memory resource management engine calls a special communication protocol interface to obtain all memory information from the CXL switching chip to complete memory resource identification, and the upper-layer management software provides a network API (Application Programming Interface) interface. The memory resource management engine not only provides the network API interface, but also provides a visual interface, and a user accesses the management engine through a browser to view a memory resource list and a system topology. The visual interface provides two display modes: list display and graphical display. The list display mainly includes: a device memory resource list (including device asset information, device partition information, device memory allocation relationship, etc.), a computing unit list (including computing unit basic information, computing unit location information, computing unit resource list, etc.), and a port resource list (such as port basic attributes, port configuration information, port reserved resource information, etc.). The graphical display optional function is as shown in FIG. 15, and the graphical interface shows the connection relationship between the host, the switch, the memory, the device and the control unit.
[0099] The memory resource mainly includes the vendor information, device type, device partition information, and allocation status of the memory resource. In the port resource list part, the basic attributes of the port, the port resource reservation state, and the port allocation relationship are mainly included. The memory resource management engine obtains and displays the information of the memory controller, which mainly includes: the location of the memory controller; the memory resource status allocated by the memory controller; and the device link health status of the memory controller. The management engine provides the graphical display and management of the memory resource through the WEB GUI (Graphical User Interface) interface. The user can complete the dynamic allocation of the resource in the interconnection topology view through dragging and other forms, and the device link is monitored in real time.
[0100] For the device partition and dynamic allocation information, the memory resource management engine is connected to each memory controller through the network and interconnected with the resource management controller on each memory controller. The resource management software on each memory controller is responsible for monitoring the memory resource usage of the memory controller, and the memory resource management system engine is responsible for the memory resource monitoring and scheduling of the whole system, as shown in FIG. 16.
[0101] Exemplarily, referring to FIG. 16, the extended memory of a certain computing unit (general computing unit or heterogeneous computing unit) in a long-term idle state can be recycled to the resource pool, and the memory resource management engine controls to start the extended memory hot removal process; the computing unit extended memory hot removal process is ended, and the management engine is notified to update the system state; the memory resource management engine allocates the memory to another computing unit (general computing unit or heterogeneous computing unit) and prepares the hot addition process; the computing unit hot addition extended memory process is ended, and the memory resource management engine is notified to update the system state.
[0102] The device partition information is the division of the memory resource. One memory controller can be connected with multiple memories, and all the memory resources under the same memory controller can be divided into several intervals, that is, the different memory bars under the same memory controller in the system can be allocated to the same or different computing units. The division of the memory can be perceived and monitored by the resource management system, and the resource management system provides visual interactive adjustment of the device partition.
[0103] The dynamic allocation of the memory resource means that the memory resource management engine can dynamically migrate the memory from a certain computing unit to another computing unit, and the whole operating system remains online, without the need for the restart of the computing unit, which does not affect the running business, and the whole process can be monitored in real time on the WEB interface of the management engine. The above function is based on the hot plug basic attribute of the CXL device, and does not need the additional function support of the memory.
[0104] Through the embodiment, by means of the CXL memory (CXL device) hung under the memory controller, the multi-computing resources in the multi-computing resource pool are connected with the memory controller to which the CXL memory belongs via a set of CXL switching chips and a set of CXL memories through the cache consistency bus inside the server system, so that the convenience of CXL memory control can be improved.
[0105] In some example embodiments, the server system further comprises a management module, wherein the management module is configured to manage at least one of the following of the server system: device management, in-band management, out-of-band management, security management, and the management module and the control module are the same module or different modules. The management module can be the aforementioned unified management module, and the control module can be the same module or a different module.
[0106] Here, the management module can belong to a management control system, which is mainly responsible for the management control of the whole system. From the management level, the substrate controller of the high-performance CXL switching unit is taken as the core, and the whole system centralized management control is realized through the network management controller and the substrate controller in the high-density memory expansion unit. The management control system architecture is shown in FIG. 17, and the management control system comprises a TOR (Top of Rack) network switch, a management controller, a high-performance CXL switching unit and a high-density memory expansion unit connected through a LAN (Local Area Network).
[0107] The management controller can include a processor, a PCH (Platform Controller Hub) ME (Management Engine), a CPLD (Complex Programmable Logic Device), a BIOS (Basic Input / Output System) ROM (Read-Only Memory), a baseboard management controller, wherein the processor is connected with the PCH ME through a PECI (Platform Environment Control Interface) bus, the BIOS ROM is connected with the PECI bus and the baseboard management controller through an SPI (Serial Peripheral Interface) bus, the CPLD is connected with the baseboard management controller through an I2C (Inter-Integrated Circuit) bus or a JTAG (Joint Test Action Group), and the PCH ME is connected with the baseboard management controller through an LPC (Low Pin Count), an ESPI (Enhanced Serial Peripheral Interface), a VGA (Video Graphics Array) or a USB (Universal Serial Bus).
[0108] The high-performance CXL switching unit can include a CXL memory switching chip, a CPU and a baseboard management controller, wherein the CXL memory switching chip and the CPU are connected through an out-of-band management bus, and the CPU and the baseboard management controller are connected through an SGMII (Serial Gigabit Media Independent Interface).
[0109] The high-density memory expansion unit can include a memory DIMM / E 3, a memory control chip, and a baseboard management controller, wherein the memory DIMM (Dual In-line Memory Module) / E 3 (ECC, Error Checking and Correcting) and the memory control chip are connected through a DDR or I2C bus, and the memory control chip and the baseboard management controller are connected through a DOE (Data Object Exchange) special command interface of the I2C bus.
[0110] The management controller, the high-performance CXL switching unit, and the baseboard management controller of the high-density memory expansion unit can be connected with a voltage sensor, a temperature sensor, a power consumption sensor, or other sensors through an out-of-band management bus.
[0111] The management control system mainly includes node management, memory monitoring management, and fault diagnosis function modules, and a software architecture diagram is shown in FIG. 18. The functions of the management controller include health state monitoring, component fault detection, running state management, component fault diagnosis, asset management, component fault reporting, etc. The functions of the high-performance CXL switching unit include whole machine topology discovery, memory resource topology, heat dissipation management, whole machine running state management, memory asset management, log management, whole machine system health state monitoring, overall memory fault management, firmware management, etc. The functions of the high-density memory expansion unit include in-place memory monitoring, memory fault detection, memory sensor monitoring, memory fault diagnosis, memory asset information monitoring, memory fault reporting, etc.
[0112] For node management, the management modules are interconnected through Ethernet under the whole machine framework of the memory-pooled storage server system, and the baseboard management controller of the high-performance CXL switching module monitors the upstream controller and the downstream memory expansion unit, to realize whole machine topology discovery, node running state, and whole machine system health state monitoring.
[0113] For memory monitoring management, the baseboard management controller sends a DOE special command interface to the memory controller through an out-of-band management bus, to obtain in-place, asset information, sensor, and fault diagnosis information of each memory, to form a whole resource topology of the pooled memory. A visual page is designed to centrally display the whole memory pool resource state, including the monitoring state and resource allocation utilization of all memories.
[0114] For the diagnosis of pooled memory failure, unlike the traditional local memory, the extended memory is decoupled from the host through the CXL memory controller, which as the master unit of memory read and write, masters the memory correction information based on ECC. The baseboard controller and the CXL memory controller establish a dedicated DOE communication protocol to collect memory failure information and realize accurate fault detection at the memory bank level.
[0115] Through this embodiment, the management, security and control functions are separated from the computing unit, which can be compatible with different computing power platforms and management platforms, and improve the compatibility of multi-element computing modules.
[0116] In some example embodiments, similar to the foregoing embodiments, the server system further comprises: an input / output device module, wherein the input / output device module is connected to the multi-element computing resource pool through the cache coherence bus inside the server system via the switching module; a heat dissipation module configured to dissipate heat for at least part of the devices in the server system; a power supply module configured to supply power to the server system; and a network module configured to perform network connection of the server system. As has been described, no further elaboration is made here.
[0117] In some embodiments, the management module supports multiple interfaces for managing different modules. Correspondingly, the multi-element computing resources in the multi-element computing resource pool, the shared memory in the memory resource pool, the storage devices in the storage resource pool, the switching module, the input / output device module, the heat dissipation module, the power supply module and the network module are respectively connected to the corresponding interfaces of the management module through at least one of the following buses: system management bus, power management bus, high-speed peripheral component interconnect bus, computing quick link bus, universal asynchronous receiver / transmitter bus, integrated circuit bus, serial peripheral interface bus.
[0118] For example, in terms of management, as shown in FIG. 19, a one-machine multi-core system design can decouple the computing unit and the management module and define them in a standardized manner, separating the management, security and control functions from the computing unit. This can be compatible with different computing power platforms and management platforms, support multi-interface multi-core unified management, and meet the needs of different application scenarios. Each module can be connected to the unified management module through at least one of the following buses: SMBUS (System Management Bus), PMBUS (Power Management Bus), PCIe bus, CXL bus, UART (Universal Asynchronous Receiver / Transmitter) bus, I2C bus, SPI bus.
[0119] Through the embodiment, the management module supports multiple interfaces to manage different modules (i.e., supports multiple interfaces and unified management of multiple cores), which can meet the needs of different application scenarios.
[0120] In some example embodiments, the management module can be configured to perform multiple management functions, which can include but are not limited to at least one of the following: topology identification of the server system to obtain system topology information of the server system, wherein the system topology information includes node reference information of system nodes in the server system, the system nodes in the server system correspond to system devices in the server system, and the node reference information includes at least one of the following: node type, power-on state, and health status; management of device asset information of the system devices of the server system; power-on and power-off control of the system devices of the server system to make the system devices of the server system power on and power off in a preset order; automatic detection of a reset signal for system reset during power-on of the server system or host restart of the server system.
[0121] Here, the management module can realize computing power allocation and management, voltage regulation and power consumption management of the multi-core computing module, and ensure efficient system design; realize monitoring of resource utilization rate, I / O throughput, and resource health status, realize collaborative power-on and power-off, centralized management of resources and topology, and ensure system availability; realize on-demand rapid deployment and automatic management of hardware resources, monitoring and fault management of resource key information, intelligent positioning and recovery of faults, and ensure the reliability of the computing system.
[0122] Regarding power consumption management, the DVFS (dynamic voltage frequency scaling) technology based on voltage and frequency regulation has become mainstream in the design of power consumption management of processors and their systems. The basic working principle of DVFS is to monitor the working load of the processor to determine its different requirements for computing power, dynamically adjust the voltage and frequency of each module in each processor according to a preset control strategy, reduce the voltage and frequency while ensuring performance, and thus achieve the purpose of reducing the power consumption of the processor.
[0123] Figure 20 is three kinds of multi-core processor DVFS voltage regulation mode. Respectively, off-chip regulation, on-chip global regulation and on-chip independent control regulation. Off-chip regulation technology has high power conversion efficiency, and the stabilization time is long, usually tens of microseconds, and the regulation efficiency is low. On-chip regulation can shorten the voltage stabilization time to tens to hundreds of nanoseconds, but on-chip regulation needs to integrate the voltage regulation module inside the processor, which increases the chip area and improves the design complexity. With the increase of the number of multi-core processor cores and the development of heterogeneous core integration technology, the chip area overhead and design difficulty brought by independent control of each core gradually increase, so on the basis of independent control and global control, the partition control regulation technology is proposed. Partition control regulation is to divide the multi-core processor into multiple regions, and each region executes the same voltage and frequency control strategy to balance the design overhead and power consumption.
[0124] For system-level power management, a multi-level power management mechanism is usually built using a combination of software and hardware, as shown in Figure 21. System-level multi-level power management can be divided into three levels: high-performance infrastructure design and support technology, low-power compilation technology, and fine-grained power operation management. High-performance infrastructure design and support technology refers to the design of power supply and heat dissipation modules at the system-level hardware level to achieve high conversion efficiency and efficient heat dissipation. Low-power compilation optimization management is an energy optimization at the bottom level of the processor. By optimizing the compilation of instructions, the energy consumption of the processor in fetching instructions and decoding can be reduced. Fine-grained power management is a software-level mechanism that uses the fine-grained resource allocation, release, and loading mechanism of the operating system to sense the working state of each computing unit, control idle computing units to enter sleep mode, and control them to switch to working state when needed, thereby achieving energy consumption control.
[0125] The one-machine multi-core system management module can be responsible for unified management, provide operation and maintenance capabilities through standardized service interfaces, and realize integrated monitoring, fault early warning, and visual management. The one-machine multi-core system fusion management architecture is shown in Figure 16. Its functions can include but are not limited to at least one of the following: topology identification, supporting topology view, and viewing node summary information such as node type, power-on state, and overall health status; asset information management, supporting hierarchical device information viewing, and viewing device detailed information such as device asset information and high-speed interface connection state through Web / Redfish pages; collaborative control, supporting centralized power-on and power-off control, and each unit powering on and off in sequence; supporting system reset function, automatically detecting reset signals for reset during power-on process and host restart process; supporting reset control after resource reallocation, etc.
[0126] Through the topology identification, asset information management, and collaborative control performed by the management module, the management capability of the management module for the server system can be improved, and the controllability of the system can be improved.
[0127] In some example embodiments, the operating system of the server system includes a plurality of operating system kernels, an operating system kernel of the plurality of operating system kernels corresponds to a plurality computing resources in the plurality computing resource pool, and the operating system kernel of the plurality of operating system kernels is compiled and run under a corresponding instruction set architecture instruction set. The server system performs computing power scheduling for a running application in the operating system of the server system by specifying system software.
[0128] In order to fully exert the performance of heterogeneous multi-core and adapt to various application scenarios, it is necessary to ensure that the application can be freely migrated between multiple cores. However, migration of the application between different instruction set heterogeneous multi-cores is a problem to be solved. Load migration and resource scheduling are applied in the following scenarios: the operating system performs load balancing, and the process is migrated to a static or lightly loaded core; according to different load types, the process is migrated to a core of a different instruction set architecture; the power consumption state changes, and some process migration needs to be handled to achieve power consumption control; when a core is overheated, the process is migrated to another core, and the like.
[0129] In this embodiment, the operating system of the server system (which can be an operating system on a baseboard management controller) includes a plurality of operating system kernels (kernels), an operating system kernel of the plurality of operating system kernels corresponds to a plurality computing resources in the plurality computing resource pool, and each operating system kernel is compiled and run under a corresponding ISA instruction set. The server system performs computing power scheduling for a running application in the operating system of the server system by specifying system software, and the computing power scheduled by the specified system software is a plurality of computing resources in the plurality computing resource pool. Here, the entire system performs on-demand balanced scheduling of computing power for applications through unified system software, which can reduce the difficulty of migration of the application on different instruction set architecture cores and improve the possibility of migration of the application on different instruction set architecture cores.
[0130] For example, for a one-machine multi-core system, the operating system includes multiple kernels, each kernel is compiled and run under a specific ISA instruction set, and the entire system performs on-demand balanced scheduling of computing power for applications through unified system software.
[0131] Through this embodiment, the server system performs on-demand balanced scheduling of computing power for applications through unified system software, which can ensure the rationality of computing power scheduling and improve the possibility of migration of the application on different instruction set architecture cores.
[0132] In some example embodiments, the server system further comprises a scheduling core, and the system software is configured to perform computing power scheduling for running applications in an operating system of the server system through a scheduling system on the scheduling core, the operating system on the plurality of operating system cores is a private operating system, and the plurality of operating system cores are connected through an inter-core communication bus, and the inter-core communication bus is configured to perform inter-core communication of the plurality of operating system cores.
[0133] In this embodiment, at the operating system level, communication and cooperation among the plurality of operating system cores are implemented, so that the entire system still has a global state under different ISA core running instances, and through a unified scheduling core and application of a corresponding scheduling technology, migration of an application between different instruction set heterogeneous multi-cores can be guaranteed.
[0134] For example, referring to FIG. 22, at the operating system level, communication and cooperation among the kernels are implemented, so that the entire system, including different private operating systems (OSs), still has a global state under different ISA core running instances, the global memory allows a plurality of computing cores to access and share data, and each computing core can have its own private memory. On this basis, through a unified scheduling core and application of an advanced scheduling technology, an application program is not aware of the underlying hardware, the application can automatically select the optimal mapping through cooperation of the operating system and the scheduling system, and performance of a multi-instruction set heterogeneous multi-core system is improved.
[0135] Through this embodiment, computing power scheduling is performed for an application based on a unified scheduling core, rationality and timeliness of computing power scheduling can be improved, and migration of an application between different instruction set heterogeneous multi-cores can be guaranteed.
[0136] In some example embodiments, in a one-machine multi-core system design, the discovery, management and flexible adjustment of system resources are critical. The control module is the core management unit of dynamic resource adjustment, and controls the high-performance switching unit to realize automatic discovery of resource topology and flexible automatic switching of resources. The optional functions are as follows: resource identification and display, providing a network interface and a visual Web interface to display the resource list and topology, including heterogeneous computing resource unit information, I / O port information, etc.; dynamic allocation of resources, through techniques such as hot removal, hot insertion and hot reset of devices, to realize dynamic allocation and adjustment of heterogeneous computing units at a level of seconds; load balancing, based on an optimized scheduling algorithm for load balancing, to dynamically balance and allocate physical resources, realize fine-grained storage resource on-demand allocation, and maximize the release of heterogeneous computing power; expert template, according to business requirements, resource state and performance indicators, etc. Parameters, to realize expert template description and dynamic switching interface of heterogeneous computing power resources, so that applications can apply for resources according to the expert template, and at the same time provide a network interface and a visual Web interface to realize visual application of the expert template. Here, the expert template describes the rules for resource switching, i.e. how many resources are configured in what situation, and the rules are defined in advance. The dynamic switching interface refers to the dynamic switching interface of the system according to the expert template after the scene is changed.
[0137] In some embodiments, the control module is further configured to receive a resource dynamic scheduling request sent by a resource scheduling requester, wherein the resource scheduling requester is a general computing unit in the multi-element computing resource pool, and the resource dynamic scheduling request is used to request to adjust the resource state of a system resource of a specified resource type. The system resource of the specified resource type is at least one of the following: a heterogeneous computing unit in the multi-element computing resource pool, and a data storage resource in the data storage resource pool. In response to the received resource dynamic scheduling request, the resource state of a target system resource is adjusted, wherein the resource type of the target system resource is the specified resource type.
[0138] In the embodiment, the control module can schedule the system resources of the server system based on the received resource dynamic scheduling request. Accordingly, the control module is further configured to receive a resource dynamic scheduling request sent by a resource scheduling requestor. The resource scheduling requestor is a general computing unit in the multi-computing resource pool, and the resource dynamic scheduling request is used to request to adjust the resource state of the system resource of a specified resource type. The system resource of the specified resource type is at least one of the following: a heterogeneous computing unit in the multi-computing resource pool, and a data storage resource in the data storage resource pool. If the resource dynamic scheduling request requests to adjust the resource state of the specified system resource, the resource dynamic scheduling request can carry the resource information of the specified system resource (for example, release the specified system resource); if the resource dynamic scheduling request does not request to adjust the resource state of the specified system resource, the resource dynamic scheduling request can carry the type information of the specified resource type (for example, apply for allocation of a heterogeneous computing unit). Here, the resource state of the system resource in the server system can include multiple types, such as an idle state, an occupied state, and the like.
[0139] In some embodiments, the control module is further configured to adjust the resource state of a target system resource in response to the received resource dynamic scheduling request, wherein the resource type of the target system resource is the specified resource type. Similar to the foregoing embodiment, the target system resource can be a specific system resource specified in the resource dynamic scheduling request, or can be a system resource determined based on the specified resource type after receiving the resource dynamic scheduling request.
[0140] Through the embodiment, the control module adjusts the resource state of the system resource of the specific resource type based on the received resource dynamic scheduling request, which can improve the convenience of system resource scheduling.
[0141] In some example embodiments, the resource dynamic scheduling request is a resource allocation request, which is used to request to allocate the system resource of the specified resource type to the resource scheduling requestor. In addition to indicating the specified resource type, the resource amount of the requested system resource and other requirements required to be met by the system resource can also be indicated.
[0142] Correspondingly, the control module is further configured to determine the system resource allocated to the resource scheduling requestor from the unallocated system resource of the specified resource type in the server system to obtain a target system resource, and allocate the target system resource to the resource scheduling requestor in response to the received resource allocation request. Here, the resource state of the unallocated system resource can be an idle state, and the resource state of the allocated system resource can be an occupied state.
[0143] Through the embodiment, the convenience of system resource allocation can be improved by allocating a specific resource type of system resource for a general computing node in response to a received resource allocation request.
[0144] In some example embodiments, the control module is further configured to send a resource configuration request to the switching module to configure an association between the resource scheduling requestor and the target system resource in the switching module.
[0145] In the embodiment, the general computing unit in the multi-element computing resource pool is connected with the data storage resource in the data storage resource pool and the heterogeneous computing unit in the multi-element computing resource pool through the switching module, and the switching module stores an association between the general computing unit connected thereto and the system resource allocated for the general computing unit. Storing the above association in the switching module not only facilitates data forwarding, but also facilitates network topology construction by the management module and the like.
[0146] To ensure the timeliness and accuracy of the association stored in the switching module, the control module is further configured to, in a case where the target system resource is allocated to the resource scheduling requestor, send a resource configuration request to the switching module to configure an association between the resource scheduling requestor and the target system resource in the switching module.
[0147] Through the embodiment, the timeliness and accuracy of information configuration can be improved by the control module configuring the corresponding association in the switching module based on the allocated system resource.
[0148] In some example embodiments, the resource dynamic scheduling request can be a resource release request for requesting to release the target system resource allocated for the resource scheduling requestor. The resource release request can carry resource information of the target system resource so as to enable the control module to determine the target system resource to be released. The resource information of the target system resource can be a resource identifier of the target system resource or information capable of identifying the target system resource.
[0149] Correspondingly, the control module is further configured to, in response to the received resource release request, send a resource location acquisition request to the switching module to acquire the resource physical location of the target system resource. Here, the switching module can determine the resource physical location of the target system resource based on the system resource or other information connected to the port thereof, and return the resource physical location of the target system resource to the control module.
[0150] The control module is further configured to send a reset signal to the target system resource according to the obtained resource physical location, so as to reset the target system resource. Here, the resetting of the target system resource can include breaking the association with the resource scheduling requester, and can further include requesting other invalid information (i.e., information related to the source scheduling requester and irrelevant to subsequent resource allocation).
[0151] Through the embodiment, the resource physical location of the system resource to be released is obtained through interaction with the exchange module, and the resource release is performed based on the obtained resource physical location, so that the accuracy of the resource release can be improved.
[0152] In some example embodiments, the target system resource is further configured to, in response to the received reset signal, perform a reset operation after a process running on the target system resource ends, and send an indication message to the control module to indicate that the reset of the target system resource is completed after the reset is completed.
[0153] In the embodiment, for the target system resource, after receiving the reset signal, the reset operation can be directly performed. At this time, if there is a running process, the execution of the reset operation can cause the interruption of the running process, thereby causing a system error.
[0154] In order to improve the safety and reliability of the system operation, the target system resource can perform a reset operation after a process running on the target system resource ends, and send an indication message to the control module to indicate that the reset of the target system resource is completed after the reset is completed. The target system resource after the reset can be allocated to other computing units.
[0155] For example, for a multi-core system, the exchange unit is a key infrastructure connecting the multi-element computing module and the system resource pool. The system realizes resource topology automatic discovery, resource flexible scheduling by controlling the high-performance exchange unit, and provides a standard interface for the data center monitoring and management platform (the monitoring and management platform is a higher level management platform than the aforementioned management platform, and is a data center level management platform) to realize resource centralized scheduling and management. The overall hardware and software design scheme of the system resource scheduling is shown in FIG. 23. Thanks to the powerful resource expansion capability of the high-performance exchange unit (the capability of the exchange module is expansion), the high-performance exchange unit is taken as the core to realize the identification of the multi-element computing unit and the device in the multi-core system server system, and the display of the topology information and asset information.
[0156] Referring to FIG. 23, a multi-core system can include an interface layer, a function layer, a driver layer and a hardware layer. In the hardware layer, a multi-element computing resource pool includes a general computing unit resource pool (i.e., a general computing resource pool) and a heterogeneous computing unit resource pool (i.e., a heterogeneous computing resource pool), wherein the general computing unit resource pool includes general computing units 1 to n1, and the heterogeneous computing unit resource pool includes heterogeneous computing units 1 to n2. A data storage resource pool includes a memory resource pool and a storage resource pool (i.e., a storage device resource pool), wherein the memory resource pool includes memories 1 to n3, and the storage resource pool includes storage devices 1 to n4. The general computing unit resource pool, the heterogeneous computing unit resource pool, the memory resource pool and the storage resource pool are connected through a switching unit (a high-performance switching unit).
[0157] The switching unit is connected to the function layer through a plurality of interfaces of the driver layer, which can include but are not limited to a UART bus and an I2C bus. The function layer can have at least one of the following functions: port configuration, resource identification, resource information display, resource information management, resource dynamic allocation, load balancing, fault monitoring and log management. The interface layer can be connected to upper-layer management software and a GUI interface through a Restful (Representational State Transfer) API (Application Programming Interface) or a GUI interface.
[0158] Here, the system resource scheduling and allocation process is shown in FIG. 24. The system management interface is used to perform operations such as on-demand allocation, dynamic scaling and resource release on general, heterogeneous and other multi-core computing modules and system resources. Users can match system resources with computing power according to load requirements, or match heterogeneous computing power with general computing power. When the application stops, the user can release the resources back to the resource pool (for example, by sending a command through a Restful interface or a GUI interface, and the system releases the resources) to facilitate efficient resource circulation and full utilization.
[0159] Referring to FIG. 24, a general computing unit can initiate a resource dynamic scheduling request, for example, to release a certain heterogeneous computing unit or other system device and allocate it to other computing units to ensure that the related processes of the device are ended. The unified management and control module responds to the received resource dynamic scheduling request to hot remove the device, and the management module sends a request to the switching unit to obtain the physical location of the device, thereby determining the physical location of the device. The management unit sends a reset signal to the device, the device completes the reset, and sends a reset completion indication to the management module. The management module can hot allocate the reset completed device to other computing units.
[0160] Through the embodiment, the system resources to be released are reset after the process running thereon ends, which can ensure the safety and reliability of system running.
[0161] In some example embodiments, the interconnection topology between the multiple computing resources in the pool of multiple computing resources can be configured as needed. For example, the interconnection network structure of the multi-core computing module of the one-machine multi-core system can adopt a Bus bus structure or a crossbar bus structure when the number of multiple computing resources is small, or can select a Mesh structure when the number of multiple computing resources is large or considering the scalability of the architecture later. That is, the interconnection topology between the multiple computing resources in the pool of multiple computing resources is a mesh structure. The mesh structure adopted by the interconnection topology between the multiple computing resources in the pool of multiple computing resources can be one or more, and can include but is not limited to at least one of the following: a two-dimensional mesh structure, a three-dimensional mesh structure.
[0162] As an optional implementation, the mesh structure adopted by the interconnection topology between the multiple computing resources in the pool of multiple computing resources can be a two-dimensional mesh structure (two-dimensional Mesh structure), for example, a cross mesh structure. The nodes of the cross mesh structure are the multiple computing resources (for example, processor cores) in the pool of multiple computing resources. The nodes of the cross mesh structure are only connected to the adjacent nodes in the same row or column.
[0163] For example, in the cross mesh structure, the nodes representing the processor cores are only connected to the adjacent nodes in the same row or column. This method has a simple topology, good scalability, and convenient path finding, and is an optional on-chip interconnection network structure. Therefore, the Mesh network structure is shown in FIG. 25. The processor cores are located at the network intersection node positions.
[0164] As another optional implementation, the mesh structure adopted by the interconnection topology between the multiple computing resources in the pool of multiple computing resources can be a three-dimensional mesh structure. The nodes of the three-dimensional mesh structure are the multiple computing resources in the pool of multiple computing resources. The nodes of the three-dimensional mesh structure are only connected to the adjacent nodes in the horizontal or vertical direction.
[0165] Here, when the number of processor cores is large, the same communication path between the processor cores at the edges of the two-dimensional Mesh structure (for example, the cross network structure described above) is long, and the transmission delay is greatly increased. To improve the transmission delay between the processor core nodes far apart in the two-dimensional Mesh structure, a three-dimensional Mesh structure can be extended on the basis of the two-dimensional Mesh structure, as shown in FIG. 26. The three-dimensional Mesh structure or even a higher-dimensional Mesh structure can greatly reduce the network diameter of the Mesh structure and reduce the transmission delay.
[0166] Through the embodiment, the interconnection topology between the multiple computing resources in the multiple computing resource pool adopts a two-dimensional grid structure, which can improve the simplicity and scalability of the network topology, and the interconnection topology between the multiple computing resources in the multiple computing resource pool adopts a three-dimensional grid structure, which can reduce the network diameter of the grid structure and reduce the transmission delay.
[0167] It should be noted that the above modules can be implemented by software or hardware, and for the latter, the following implementation manners can be used, but are not limited thereto: the above modules are located in the same processor; or the above modules are located in different processors in any combination.
[0168] According to another aspect of the embodiment, a resource scheduling method of a server system is also provided, which is executed by the server system in the foregoing embodiments and is used to implement the foregoing embodiments and optional implementation manners, which have been described and will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0169] FIG. 27 is a flow diagram of an optional resource scheduling method of a server system according to an embodiment of the present application. As shown in FIG. 27, the flow includes the following steps: in step S2701, the control module performs computing power scheduling on the multiple computing resources in the multiple computing resource pool and dynamically allocates the data storage resources in the data storage resource pool to the multiple computing resources in the multiple computing resource pool. The server system performs the above steps through the control module, and the manner can refer to the foregoing embodiments, which will not be repeated here.
[0170] Through the above steps, the control module obtained by modularization performs computing power scheduling on the multiple computing resources in the multiple computing resource pool and dynamically allocates the data storage resources in the data storage resource pool to the multiple computing resources in the multiple computing resource pool, thereby solving the problem of poor scalability of computing power in the related art multi-computing power fusion scheme and improving the scalability of computing power.
[0171] In some example embodiments, the power scheduling of the multi-computing resources in the multi-computing resource pool and the dynamic allocation of the data storage resources in the data storage resource pool to the multi-computing resources in the multi-computing resource pool by the control module comprises: receiving, by the control module, a resource dynamic scheduling request sent by a resource scheduling requester, wherein the resource scheduling requester is a general computing unit in the general computing resource pool or a heterogeneous computing unit in the heterogeneous computing resource pool, and the resource dynamic scheduling request is used to request scheduling of a system resource of a specified resource type, and the system resource of the specified resource type is at least one of the following: a general computing unit in the general computing resource pool, a heterogeneous computing unit in the heterogeneous computing resource pool, and a data storage resource in the data storage resource pool; and in response to the received resource dynamic scheduling request, scheduling, by the control module, a target system resource, wherein the resource type of the target system resource is the specified resource type.
[0172] In the present embodiment, the scheduling of the multi-computing resources and the data storage resources by the control module can be performed based on a resource dynamic scheduling request of a multi-computing resource (e.g., a general computing unit) in the multi-computing resource pool, and the resource dynamic scheduling request can be a resource allocation request, a resource release request, or other requests capable of scheduling at least one of the multi-computing resources and the data storage resources.
[0173] The control module schedules the target system resource in response to the received resource dynamic scheduling request, and the manner of scheduling is similar to that in the foregoing embodiments. The processing flow in the case where the resource dynamic scheduling request is a resource allocation request or a resource release request is similar to that in the foregoing embodiments. Details are not described herein.
[0174] Through the present embodiment, the resource state of the system resource of a specific resource type is adjusted by the control module based on the received resource dynamic scheduling request, which can improve the convenience of system resource scheduling.
[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and a general hardware platform as necessary, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk), and the storage medium can be a non-volatile storage medium, including a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the method of each embodiment of the present application.
[0176] According to still another aspect in the embodiments, there is also provided a chip that can be used to implement the aforementioned server system, for example, the aforementioned one-chip multi-core system, which has been described above and will not be repeated here.
[0177] FIG. 28 is a structural block diagram of an optional chip according to an embodiment of the present application. As shown in FIG. 28, the chip includes multi-element computing resources 2801, on-chip memory 2802, and a control core 2803. The multi-element computing resources 2801, the on-chip memory 2802, and the control core 2803 are interconnected through an on-chip cache coherence bus of the chip.
[0178] The multi-element computing resources 2801 include general-purpose computing resources and heterogeneous computing resources. The general-purpose computing resources include a group of general-purpose processors, and the heterogeneous computing resources include a group of heterogeneous acceleration processors. The general-purpose processors correspond to the aforementioned general-purpose computing units, and the heterogeneous acceleration processors correspond to the aforementioned heterogeneous computing units. The on-chip memory 2802 is allowed to be shared by the general-purpose processors in the general-purpose computing resources and the heterogeneous acceleration processors in the heterogeneous computing resources. The control core 2803 is configured to perform computing power scheduling on the general-purpose processors in the general-purpose computing resources and the heterogeneous acceleration processors in the heterogeneous computing resources.
[0179] In some embodiments, the chip further includes an input / output device, wherein the input / output device is allowed to be shared by the general-purpose processors in the general-purpose computing resources and the heterogeneous acceleration processors in the heterogeneous computing resources. The general-purpose processors in the general-purpose computing resources are central processing units, and the heterogeneous acceleration processors in the heterogeneous computing resources include at least one of the following: a graphics processing unit, a field-programmable gate array, and an application-specific integrated circuit.
[0180] For example, a one-chip multi-core chip module design is shown in FIG. 29, which realizes cooperation of general-purpose computing resources and heterogeneous acceleration computing resources at a micro scale, and the general-purpose computing resources and the heterogeneous acceleration computing resources share on-chip memory (memory on the chip). Each decoupling unit cooperates through an on-chip coherence interconnection bus to realize fusion of multi-core particles, and computing power scheduling is realized through a unified control core.
[0181] Through the chip provided in the embodiments, the chip includes multi-element computing resources, on-chip memory, and a control core interconnected through an on-chip cache coherence bus of the chip, wherein the multi-element computing resources include general-purpose computing resources and heterogeneous computing resources, the general-purpose computing resources include a group of general-purpose processors, and the heterogeneous computing resources include a group of heterogeneous acceleration processors; the on-chip memory is allowed to be shared by the general-purpose processors in the general-purpose computing resources and the heterogeneous acceleration processors in the heterogeneous computing resources; and the control core is configured to perform computing power scheduling on the general-purpose processors in the general-purpose computing resources and the heterogeneous acceleration processors in the heterogeneous computing resources, thereby solving the problem of poor computing power scalability in the multi-element computing power fusion scheme in the related art and improving the computing power scalability.
[0182] According to still another aspect in the embodiments, there is also provided a die, which can be used to implement the aforementioned server system, for example, the aforementioned one-machine multi-die system, which has been described and will not be repeated.
[0183] FIG. 30 is a structural block diagram of an optional die according to an embodiment of the present application. As shown in FIG. 30, the die includes a group of cores 3001 and a memory controller 3002. The group of cores 3001 and the memory controller 3002 are interconnected through an on-chip cache coherence bus within the die.
[0184] The group of cores 3001 includes a control core and a plurality of processor cores, wherein the plurality of processor cores includes a group of general-purpose processor cores and a group of heterogeneous accelerator processor cores, the cores in the group of cores have private caches (e.g., L1-level caches and L2-level caches), and share a designated cache (e.g., an L3-level cache) and a memory controlled by the memory controller. The control core is configured to perform power scheduling on the general-purpose processor cores in the group of general-purpose processor cores and the heterogeneous accelerator processor cores in the group of heterogeneous accelerator processor cores.
[0185] In some embodiments, the die further includes an I / O controller, wherein the I / O devices controlled by the I / O controller are shared by the general-purpose processors in the general-purpose computing resources and the heterogeneous accelerator processors in the heterogeneous computing resources. The general-purpose processor cores in the group of general-purpose processor cores are central processor cores, and the heterogeneous accelerator processor cores in the group of heterogeneous accelerator processor cores include at least one of the following: a graphics processing unit core, a field programmable gate array core, and an application-specific integrated circuit core.
[0186] It should be noted that the chip (chip module) in the foregoing embodiments is described in the dimension of a chip, and different computing cores are integrated in the chip. The dimension of the die is smaller, and the integration of the multi-element heterogeneous computing cores is performed on the die in the chip. The one-machine multi-die can be at the level of a chip, a server, or an entire rack, and one rack can be one server.
[0187] For example, the one-machine multi-die die design is shown in FIG. 31. Different cores are integrated in the die, and the multi-core has private caches and shares an L3-level cache, a memory, and I / O devices. The units are connected through the bus in the die, and can implement a plurality of system operation modes such as multi-core cluster, master-slave multi-core, multi-core cooperation, and unified scheduling. Different services are run on different cores and operating systems, and the system design is optimized for application requirements.
[0188] The chiplet provided by the embodiment includes a group of cores and a memory controller interconnected by an on-chip cache coherent bus, wherein the group of cores includes a control core and a multi-element processor core, the multi-element processor core includes a group of general-purpose processor cores and a group of heterogeneous accelerator processor cores, the cores in the group of cores have private caches and share a specified cache and a memory controlled by the memory controller; the control core is configured to perform computing power scheduling on the general-purpose processor cores in the group of general-purpose processor cores and the heterogeneous accelerator processor cores in the group of heterogeneous accelerator processor cores, thereby solving the problem of poor computing power scalability of the multi-element computing power fusion scheme in the related art and improving the computing power scalability.
[0189] The optional examples in the embodiment can refer to the examples described in the above embodiments and exemplary embodiments, and details are not described herein.
[0190] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be realized by general computing devices, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by program codes executable by the computing devices, so that they can be stored in storage devices and executed by the computing devices, and in some cases, the steps described herein can be executed in different order, or they can be manufactured into individual integrated circuit modules, or multiple modules or steps can be manufactured into a single integrated circuit module. Therefore, the present application is not limited to any specific combination of hardware and software.
[0191] The above is only an optional embodiment of the present application and is not used to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A server system, comprising: a plurality of computing resource pool, a data storage resource pool, a switching module and a control module, wherein the plurality of computing resource pool comprises a general computing resource pool and a heterogeneous computing resource pool, wherein the general computing resource pool comprises a set of general computing units, and the heterogeneous computing resource pool comprises a set of heterogeneous computing units, and the plurality of computing resource pool is connected to the data storage resource pool via the switching module through a cache coherent bus inside the server system; the data storage resource pool comprises data storage resources allowed to be shared by the plurality of computing resources in the plurality of computing resource pool, wherein the plurality of computing resources in the plurality of computing resource pool comprises the general computing units in the general computing resource pool and the heterogeneous computing units in the heterogeneous computing resource pool; the control module is configured to perform computing power scheduling on the plurality of computing resources in the plurality of computing resource pool, and dynamically allocate the data storage resources in the data storage resource pool to the plurality of computing resources in the plurality of computing resource pool. 2.The server system of claim 1, wherein the server system further comprises an input and output device module, wherein the input and output device module is connected to the plurality of computing resource pool via the switching module through the cache coherent bus inside the server system; the data storage resource pool comprises a memory resource pool and a storage resource pool, wherein the memory resource pool is integrated into a memory module, and the storage resource pool is integrated into a storage device module; the plurality of computing resource pool is integrated into a plurality of computing modules, wherein the plurality of computing modules are connected to other modules in the server system except the plurality of computing modules through a set of preset interfaces, and the other modules comprise at least one of the following: the switching module, the control module. 3.The server system of claim 2, wherein the server system further comprises: a heat dissipation module configured to dissipate heat from at least part of the devices in the server system; a power supply module configured to supply power to the server system; a network module configured to perform network connection of the server system; wherein the other modules further comprise at least one of the following: the heat dissipation module, the power supply module, and the network module. 4.The server system of claim 2, wherein the plurality of computing modules are in a number of multiple, and the plurality of computing modules are connected to a host of the server system in a plug-in card manner, and different plurality of computing modules are connected through the cache coherent bus inside the server system. 5.The server system of claim 1, wherein the data storage resource pool comprises a memory resource pool and a storage resource pool, wherein the memory resource pool comprises a set of shared memories, and the storage resource pool comprises a set of storage devices. The control module is further configured to dynamically allocate the shared memory in the memory resource pool and the storage device in the storage resource pool to the multi-element computing resource in the multi-element computing resource pool. 6.The server system of claim 5, wherein, The server system further comprises a memory controller, wherein the set of shared memory comprises a set of compute express link (CXL) memory, the set of CXL memory is a memory hung under the memory controller, the set of CXL memory is divided into a plurality of memory partitions, and a memory partition in the plurality of memory partitions is allowed to be allocated to the multi-element computing resource in the multi-element computing resource pool. The switching module comprises a set of CXL switch chips, wherein the multi-element computing resource in the multi-element computing resource pool is connected to the memory controller to which the set of CXL memory belongs via the set of CXL switch chips through a cache coherent interconnect for accelerators (CCIX) bus inside the server system. 7.The server system of claim 5, wherein, The server system further comprises: a management module, wherein the management module is configured to manage at least one of the following of the server system: device management, in-band management, out-of-band management, and security management, and the management module is the same module or a different module as the control module. 8.The server system of claim 7, wherein, The server system further comprises: an input / output device module, wherein the input / output device module is connected to the multi-element computing resource pool via the switching module through a CCIX bus inside the server system; a heat dissipation module configured to dissipate heat from at least part of the devices in the server system; a power supply module configured to supply power to the server system; a network module configured to perform network connection of the server system; wherein the multi-element computing resource in the multi-element computing resource pool, the shared memory in the memory resource pool, the storage device in the storage resource pool, the switching module, the input / output device module, the heat dissipation module, the power supply module, and the network module are respectively connected to a corresponding interface of the management module through at least one of the following buses: a system management bus (SMBus), a power management bus (PMBus), a peripheral component interconnect express (PCIe) bus, a CXL bus, a universal asynchronous receiver-transmitter (UART) bus, an inter-integrated circuit (I2C) bus, and a serial peripheral interface (SPI) bus. 9.The server system of claim 7, wherein, The management module is configured to perform topology identification on the server system to obtain system topology information of the server system, wherein the system topology information comprises node reference information of system nodes in the server system, the system nodes in the server system correspond to system devices in the server system, and the node reference information comprises at least one of the following: a node type, a power-on state, and a health status; manage device asset information of the system devices of the server system; perform power-on and power-off control on the system devices of the server system to make the system devices of the server system power on and power off in a preset order; and automatically detect a reset signal to reset the system during a server system power-on process or a host restart process of the server system.
10. The server system of claim 1, wherein the operating system of the server system comprises a plurality of operating system cores, an operating system core in the plurality of operating system cores corresponds to a plurality of computing resources in the multi-element computing resource pool, an operating system core in the plurality of operating system cores is compiled and run under a corresponding instruction set architecture instruction set, and the server system schedules computing power for a running application in the operating system of the server system by a specified system software, and the computing power scheduled by the specified system software is the plurality of computing resources in the multi-element computing resource pool.
11. The server system of claim 10, wherein the server system further comprises a scheduling core, the specified system software is configured to schedule computing power for the running application in the operating system of the server system by a scheduling system on the scheduling core, the operating system on the plurality of operating system cores is a private operating system, and the plurality of operating system cores are connected through an inter-core communication bus configured for inter-core communication of the plurality of operating system cores.
12. The server system of claim 1, wherein the control module is further configured to receive a resource dynamic scheduling request sent by a resource scheduling request party, wherein the resource scheduling request party is a general-purpose computing unit in the general-purpose computing resource pool or a heterogeneous computing unit in the heterogeneous computing resource pool, the resource dynamic scheduling request is used to request scheduling of a system resource of a specified resource type, and the system resource of the specified resource type is at least one of the following: a general-purpose computing unit in the general-purpose computing resource pool, a heterogeneous computing unit in the heterogeneous computing resource pool, and a data storage resource in the data storage resource pool; and in response to the received resource dynamic scheduling request, scheduling a target system resource, wherein the target system resource is of the specified resource type.
13. The server system of claim 12, wherein the resource dynamic scheduling request is a resource allocation request, and the resource allocation request is used to request allocation of the system resource of the specified resource type to the resource scheduling request party. The control module is further configured to determine, in response to the received resource allocation request, a system resource allocated to the resource scheduling requester from system resources of the specified resource type that are not allocated by the server system, to obtain the target system resource, and to allocate the target system resource to the resource scheduling requester.
14. The server system of claim 13, characterized in that, The control module is further configured to send a resource configuration request to the exchange module to configure an association between the resource scheduling requester and the target system resource in the exchange module.
15. The server system of claim 12, characterized in that, The resource dynamic scheduling request is a resource release request, wherein the resource release request is used to request release of the target system resource allocated to the resource scheduling requester; The control module is further configured to, in response to the received resource release request, send a resource location acquisition request to the exchange module to acquire a resource physical location of the target system resource, and to send a reset signal to the target system resource according to the acquired resource physical location to reset the target system resource.
16. The server system of claim 15, characterized in that, The target system resource is further configured to, in response to the received reset signal, perform a reset operation after a process running on the target system resource ends, and to send an indication message indicating that the target system resource is reset to the control module after the reset is completed.
17. The server system of claim 1, characterized in that, The interconnection topology between the multiple computing resources in the multiple computing resource pool adopts a mesh structure of at least one of: a cross mesh structure, wherein a node of the cross mesh structure is a multiple computing resource in the multiple computing resource pool, and the node of the cross mesh structure is connected only to adjacent nodes in the same row or column; a three-dimensional mesh structure, wherein a node of the three-dimensional mesh structure is a multiple computing resource in the multiple computing resource pool, and the node of the three-dimensional mesh structure is connected only to adjacent nodes in the horizontal or vertical direction.
18. The server system of any one of claims 1 to 17, characterized in that, The general-purpose computing unit in the multiple computing resource pool is a central processing unit, and the heterogeneous computing unit in the multiple computing resource pool includes at least one of a graphics processing unit, a field programmable gate array, and an application-specific integrated circuit.
19. A resource scheduling method for a server system, characterized in that, The server system comprises: a multi-element computing resource pool, a data storage resource pool, a switching module and a control module, wherein the multi-element computing resource pool comprises a general computing resource pool and a heterogeneous computing resource pool, the general computing resource pool comprises a group of general computing units, the heterogeneous computing resource pool comprises a group of heterogeneous computing units, the multi-element computing resource pool is connected with the data storage resource pool through the switching module via a cache coherence bus inside the server system; the data storage resource pool contains data storage resources allowed to be shared by multi-element computing resources in the multi-element computing resource pool, wherein the multi-element computing resources in the multi-element computing resource pool comprise general computing units in the general computing resource pool and heterogeneous computing units in the heterogeneous computing resource pool; The method comprises: Through the control module, the computing power of the multi-element computing resources in the multi-element computing resource pool is scheduled, and the data storage resources in the data storage resource pool are dynamically allocated to the multi-element computing resources in the multi-element computing resource pool.
20. The method of claim 19, wherein The computing power of the multi-element computing resources in the multi-element computing resource pool is scheduled through the control module, and the data storage resources in the data storage resource pool are dynamically allocated to the multi-element computing resources in the multi-element computing resource pool, comprising: Through the control module, a resource dynamic scheduling request sent by a resource scheduling requester is received, wherein the resource scheduling requester is a general computing unit in the general computing resource pool or a heterogeneous computing unit in the heterogeneous computing resource pool, the resource dynamic scheduling request is used to request to schedule system resources of a specified resource type, and the system resources of the specified resource type are at least one of the following: general computing units in the general computing resource pool, heterogeneous computing units in the heterogeneous computing resource pool, and data storage resources in the data storage resource pool; In response to the received resource dynamic scheduling request, the target system resources are scheduled through the control module, wherein the resource type of the target system resources is the specified resource type.
21. A chip, comprising: A plurality of computing resources, on-chip memory and a control core interconnected through an on-chip cache coherence bus of the chip, wherein The multi-element computing resources comprise general computing resources and heterogeneous computing resources, the general computing resources comprise a group of general processors, and the heterogeneous computing resources comprise a group of heterogeneous accelerator processors; The on-chip memory is allowed to be shared by general processors in the general computing resources and heterogeneous accelerator processors in the heterogeneous computing resources; The control core is configured to schedule the computing power of general processors in the general computing resources and heterogeneous accelerator processors in the heterogeneous computing resources.
22. A core particle, comprising: A group of cores and a memory controller interconnected through an on-chip cache coherence bus inside the core particle, wherein The set of cores includes a control core and a plurality of processor cores, wherein the plurality of processor cores includes a set of general-purpose processor cores and a set of heterogeneous accelerator processor cores, the cores in the set of cores have private caches, and share a designated cache and a memory controlled by the memory controller; The control core is configured to perform computing power scheduling on the general-purpose processor cores in the set of general-purpose processor cores and the heterogeneous accelerator processor cores in the set of heterogeneous accelerator processor cores.
Citation Information
Patent Citations
Resource sharing device, resource management device, and resource management method
CN115586964A
Converged architecture system, nonvolatile storage system and storage resource acquisition method
CN116185641A
Server system, resource scheduling method of server system, chip and chip grain
CN118210634A
Hybrid heterogeneous host system, resource configuration method and task scheduling method
US20160378548A1
Cited By
Heterogeneous resource scheduling method and device of cloud data center, medium and product
CN121433917A
Server task management system, method and equipment based on BMC (Baseboard Management Controller) collaboration and medium
CN121478362A
RISC-V multi-core heterogeneous platform intelligent load balancing method and system
CN121560575A