A resource management method, system, device and computer readable storage medium
Patent Information
- Application Number
- CN202310316633.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-03-24
AI Technical Summary
这就导致数据中心面临较大的数据负载
[0040]本申请提供的一种资源管理方法,应用于目标主机,获取数据处理请求;基于一致性互联协议,获取目标计算架构中第一数量个资源设备的目标资源,且第一数量的值大于等于2;对目标资源进行配置,得到目标资源的配置结果;按照配置结果,通过目标资源对数据处理请求进行处理;其中,目标计算架构中的资源设备间通过一致性互联协议相连通。本申请中目标计算架构中的资源设备间通过一致性互联协议相连通,这样,目标主机获取多个资源设备的目标资源后,可以统一对目标资源进行配置,并可以按照配置结果来应用多个目标资源对数据处理请求进行处理,实现计算架构中资源设备之间的高效协同。本申请提供的一种资源管理系统、设备及计算机可读存储介质也解决了相应技术问题。
Smart Images

Figure CN116302554B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a resource management method, system, device, and computer-readable storage medium. Background Technology
[0002] With the rapid development of technologies such as cloud computing, big data, 5G, and AR (Augmented Reality), and the widespread application of the Internet of Things, there is a growing demand for real-time data processing anytime, anywhere. Coupled with the direct impetus from the "new infrastructure" initiative, edge data centers have rapidly come to the forefront. In the 5G era, user terminals will form an incredibly tight cloud-edge-device architecture with cloud data centers and edge data centers.
[0003] Edge data centers were developed to support the deployment of new 5G services with lower latency. Due to the extremely high density of terminals supported by 5G, the resulting data volume is staggering. For application systems to truly benefit, edge data centers must be able to access, process, and communicate with end users very quickly—this is what edge data centers are required to do. This leads to significant data loads on data centers. Simultaneously, this has spurred the development of various AI chips with different specifications for cloud-edge-device scenarios, resulting in the diversification of heterogeneous AI devices. According to IDC predictions, the global AI chip market will reach $72.6 billion by 2025. Since 2019, the industry has witnessed a flourishing of AI chip products for cloud-edge-device scenarios. Against the backdrop of the future development of a nationwide integrated intelligent computing power ecosystem and the increasing demand for computing power from the AI industry, the collaborative integration of various heterogeneous computing powers that make up a single computing architecture has become an important development trend. Therefore, achieving efficient collaboration between resource devices in a computing architecture is of great significance to the development of current AI (Artificial Intelligence) systems.
[0004] In summary, how to achieve efficient collaboration between resource devices in a computing architecture is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this application is to provide a resource management method that can, to some extent, solve the technical problem of how to achieve efficient collaboration between computing devices in a computing architecture. This application also provides a resource management system, a device, and a computer-readable storage medium.
[0006] To achieve the above objectives, this application provides the following technical solution:
[0007] A resource management method, applied to a target host, includes:
[0008] Get data processing request;
[0009] Based on the consensus interconnection protocol, obtain the target resources of a first number of resource devices in the target computing architecture, and the value of the first number is greater than or equal to 2;
[0010] Configure the target resource to obtain the configuration result of the target resource;
[0011] According to the configuration results, the data processing request is processed through the target resource;
[0012] The resource devices in the target computing architecture are interconnected through the consensus interconnection protocol.
[0013] Preferably, the consistency interconnection protocol includes a first sub-protocol for transmitting physical signals to the resource device, a second sub-protocol for maintaining memory consistency in the resource device, and a third sub-protocol for the resource device to access memory consistency from the target host.
[0014] Preferably, the target computing architecture includes the first computing architecture to which the target host belongs;
[0015] The acquisition of the target resources of a first number of resource devices in the target computing architecture includes:
[0016] Obtain the target resources of the first number of resource devices in the target computing node to which the target host belongs, wherein the target computing node includes the computing node to which the target host belongs in the first computing architecture.
[0017] Preferably, acquiring the target resources of a first number of resource devices in the target computing architecture includes:
[0018] Acquire the target resources of the first number of resource devices in other computing nodes, wherein the other computing nodes include computing nodes in the first computing architecture other than the target computing node;
[0019] The resource devices in the target computing node are connected to the resource devices in the other computing nodes via a virtual high-speed bus bridge.
[0020] Preferably, the target computing architecture includes a second computing architecture in addition to the first computing architecture;
[0021] The acquisition of the target resources of a first number of resource devices in the target computing architecture includes:
[0022] Obtain the target resources of the first number of resource devices in the second computing architecture;
[0023] The first computing architecture and the second computing architecture are connected via a switch and interconnection devices. The interconnection devices include a virtual variable-capacity memory connected to the first computing architecture or the second computing architecture, a first virtual cache pool and an address translator connected to the virtual variable-capacity memory, a virtual memory-free accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memory-free accelerator.
[0024] Preferably, the target resource includes target memory resources;
[0025] The configuration of the target resource to obtain the configuration result of the target resource includes:
[0026] The target memory resource is split into target memory sub-resources;
[0027] The target memory sub-resources are uniformly addressed to obtain the configuration result.
[0028] Preferably, the resource device includes edge computing nodes and / or hybrid memory nodes and / or heterogeneous computer nodes and / or ASIC computing nodes;
[0029] The hybrid memory node includes a persistent memory resource pool consisting of CPUs and persistent memory; the heterogeneous computing node includes a computing resource module consisting of CPUs and GPUs; and the ASIC computing node includes a computing resource module consisting of heterogeneous computing resources of CPUs and ASICs.
[0030] A resource management system, applied to a target host, includes:
[0031] The first acquisition module is used to acquire data processing requests;
[0032] The second acquisition module is used to acquire the target resources of a first number of resource devices in the target computing architecture based on the consensus interconnection protocol, and the value of the first number is greater than or equal to 2;
[0033] The first configuration module is used to configure the target resource and obtain the configuration result of the target resource;
[0034] The first processing module is used to process the data processing request through the target resource according to the configuration result;
[0035] The resource devices in the target computing architecture are interconnected through the consensus interconnection protocol.
[0036] A resource management device, comprising:
[0037] Memory, used to store computer programs;
[0038] A processor for implementing the steps of any of the above-described resource management methods when executing the computer program.
[0039] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the resource management methods described above.
[0040] This application provides a resource management method applied to a target host. The method involves: acquiring data processing requests; acquiring target resources from a first number of resource devices in the target computing architecture based on a consistency interconnection protocol, where the first number is greater than or equal to 2; configuring the target resources to obtain configuration results; and processing the data processing requests using the target resources according to the configuration results. The resource devices in the target computing architecture are interconnected via the consistency interconnection protocol. In this application, the resource devices in the target computing architecture are interconnected via the consistency interconnection protocol. This allows the target host to uniformly configure the target resources after acquiring them from multiple resource devices, and to apply the multiple target resources to process data processing requests according to the configuration results, achieving efficient collaboration between resource devices in the computing architecture. This application also provides a resource management system, device, and computer-readable storage medium that solves the corresponding technical problems. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0042] Figure 1 A first flowchart of a resource management method provided in an embodiment of this application;
[0043] Figure 2 This is a schematic diagram showing how resource devices in a target compute node are connected to resource devices in other compute nodes via a virtual high-speed bus bridge.
[0044] Figure 3 A schematic diagram of the structure of interconnected devices;
[0045] Figure 4 This is a schematic diagram showing the connection between the first computing architecture and the second computing architecture.
[0046] Figure 5 A second flowchart of a resource management method provided in an embodiment of this application;
[0047] Figure 6 This is a schematic diagram illustrating the interconnection between memory resources;
[0048] Figure 7 A schematic diagram illustrating the addressing of pooled memory resources and the addressing of load / store operations;
[0049] Figure 8 This application provides a schematic diagram of the structure of a resource management system according to an embodiment of the present application.
[0050] Figure 9 This is a schematic diagram of the structure of a resource management device provided in an embodiment of this application;
[0051] Figure 10 This is another structural schematic diagram of a resource management device provided in an embodiment of this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] Please see Figure 1 , Figure 1 This is a first flowchart of a resource management method provided in an embodiment of this application.
[0054] This application provides a resource management method applied to a target host, which may include the following steps:
[0055] Step S101: Obtain data processing request.
[0056] In practical applications, the target host can first obtain the data processing request. The type of data processing request can be determined according to actual needs. For example, the data processing request can be an image processing request, a video processing request, a server security processing request, an AI model training processing request, etc.
[0057] Step S102: Based on the consensus interconnection protocol, obtain the target resources of a first number of resource devices in the target computing architecture, and the value of the first number is greater than or equal to 2. The resource devices in the target computing architecture are connected to each other through the consensus interconnection protocol.
[0058] In practical applications, after the target host obtains the data processing request, it can obtain the target resources of a first number of resource devices in the target computing architecture based on the Scalable Coherent Memory Interconnect Protocol (SCMP), and the value of the first number is greater than or equal to 2. That is, the target host can obtain the target resources of multiple resource devices in the target computing architecture based on the Scalable Coherent Memory Interconnect Protocol, so that the data processing request can be processed based on the target resources in the future.
[0059] It should be noted that the resource devices in the target computing architecture are interconnected through a consistency interconnect protocol. Considering that the function of the consistency interconnect protocol is to achieve resource synchronization between the resource devices and the target host, and this resource synchronization involves the target host transmitting physical signals to the resource devices, and the target host needs to maintain memory consistency with the resource devices, the consistency interconnect protocol of this application may include a first sub-protocol for transmitting physical signals to the resource devices, a second sub-protocol for maintaining memory consistency in the resource devices, and a third sub-protocol for the resource devices to access memory consistency from the target host. Furthermore, the type of resource device can be determined according to the specific application scenario. For example, resource devices may include edge computing nodes and / or hybrid memory nodes and / or heterogeneous computing nodes and / or ASIC computing nodes, etc.; wherein, hybrid memory nodes include persistent memory resource pools composed of CPUs and persistent memory; heterogeneous computing nodes include computing resource modules composed of CPUs and GPUs; and ASIC computing nodes include computing resource modules composed of heterogeneous computing resources of CPUs and ASICs. It should also be noted that the operating systems of each computing node in the target computing architecture can interact through Ethernet switches, and the accelerators and PCIe interface devices such as persistent memory can achieve cross-node sharing and pooled communication through PCIe switches. Of course, the target computing architecture may also include network card power supply modules, heat dissipation modules, management modules, etc., wherein the network card is interconnected with the Ethernet switch; the power supply module implements the power supply and emergency power failure protection measures for the entire computing system; the heat dissipation module implements the heat dissipation of the computing system; the management module implements the macro-control and management of the entire computing system, including power supply management, heat dissipation management, and maintenance, etc. This application does not make specific limitations on the target computing architecture.
[0060] In specific application scenarios, the target computing architecture generally includes the first computing architecture to which the target host belongs. In this case, when the target host acquires the target resources of the first number of resource devices in the target computing architecture, it can directly obtain the target resources from the first computing architecture, that is, it can acquire the target resources of the first number of resource devices in the target computing node to which the target host belongs. The target computing node includes the computing node to which the target host belongs in the first computing architecture. Alternatively, it can acquire the target resources of the first number of resource devices in other computing nodes, which include computing nodes in the first computing architecture other than the target computing node. Furthermore, the resource devices in the target computing node can be connected to resource devices in other computing nodes via a virtual high-speed bus bridge. A schematic diagram illustrating the connection between the resource devices in the target computing node and resource devices in other computing nodes via a virtual high-speed bus bridge is shown below. Figure 2As shown, Type-1 devices represent accelerators without local memory (such as smart network interface cards), using two sub-protocols: SCMP.io (first sub-protocol) and SCMP.cache (second sub-protocol) to achieve consistent reading of the CPU-side cache by the smart network interface card. Type-2 devices represent general-purpose accelerators with local memory (such as GPUs and ASICs), using three sub-protocols: SCMP.io, SCMP.cache, and SCMP.mem (third sub-protocol) to achieve both CPU reading of the cache in the accelerator and consistent reading of the CPU-side cache by the accelerator. Type-3 devices represent extended memory (such as DRAM and non-volatile memory), using two sub-protocols: SCMP.io and SCMP.mem to achieve consistent reading of the Type-3 device cache by the CPU. A Home Agent (HA) is placed on the CPU side, responsible for memory read and write operations; a Cache Agent (CA) is placed on the device side, responsible for managing cache contents. Both work together to maintain memory consistency. The root port is the root aggregation point, aggregating multiple Type-2 and Type-3 devices and mounting them below the CPU. The root port connects to a virtual high-speed bus bridge. Local Type-1, Type-2, and Type-3 devices connect and converge to the root port physical interface through virtual physical binding interfaces extended from the virtual high-speed bus bridge. (HSBB stands for High Speed Bus Bridge, VHS for Virtual HSBB Switch, and vHSBB for Virtual High Speed Bus Bridge). Cross-node memory expansion only supports extending Type-2 and Type-3 devices under VHS1 to VHS0. The implementation principle is as follows: when the memory of a Type-1 device or a Type-2 device in root port1 is logically allocated to root port0, the CPU in root port0 re-addresses the memory of the Type-1, Type-2, and Type-3 devices under root port0 and the Type-2 and Type-3 devices under root port1, allowing them to be uniformly managed and allocated by the CPU in root port0. Simultaneously, the CPU in root port1 and other computing devices lose access to the local Type-2 or Type-3 device memory.
[0061] In specific application scenarios, the target computing architecture may include a second computing architecture in addition to the first computing architecture. In this case, the target host, while acquiring the target resources of a first number of resource devices in the target computing architecture, can also acquire the target resources of a first number of resource devices in the second computing architecture. The first and second computing architectures are connected via switches and interconnect devices. The interconnect devices include a virtual variable-capacity memory connected to the first or second computing architecture, a first virtual cache pool and address translator connected to the virtual variable-capacity memory, a virtual memory-free accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memory-free accelerator. Specifically, the interconnect devices can be connected to the first or second computing architecture via an ultra-high-speed bus interface and an ultra-high-speed bus, or to a switch via a high-speed network interface and a high-speed network, etc. A schematic diagram can be shown below. Figure 3 and Figure 4 As shown, the high-speed bus can be a physical link using PCIe 5.0 or higher used within the rack / node, and the high-speed interconnect network can refer to an RDMA network based on IB / RoCE / iWARP. Its working principle is as follows: When the memory of a Type-2 or Type-3 device on the second computing architecture side is managed by the CPU on the first computing architecture side, the memory of the second computing architecture and the memory of the first computing architecture are uniformly addressed and managed. When the CPU of the first computing architecture uses the memory of the second computing architecture, the virtual memoryless accelerator module caches the memory data of the second computing architecture into the MOB (Memory Operation Buffer) in the virtual cache pool, achieving remote data remote coherence. The virtual variable-capacity memory of the first computing architecture reads the data from the virtual cache pool on the second computing architecture side into the local virtual cache pool on the first computing architecture side, achieving remote data local coherence. Furthermore, the virtual cache pool on the first computing architecture side and the virtual cache pool on the second computing architecture side can be composed of DDR5 memory, supporting a maximum capacity of 512GB. Address translation converts the memory addresses in the second computing architecture into memory addresses that the CPU in the first computing architecture can recognize. The above workflow is equivalent to a virtual variable capacity memory interfacing with processor memory access requests from an ultra-high-speed bus. The second computing architecture is equivalent to a memory-free accelerator accepting memory access requests from a high-speed network, initiating remote memory access through the remote node's ultra-high-speed bus and completing data backhaul, and reducing the average latency of cross-rack access through a data cache prefetching algorithm to achieve cross-rack distributed shared memory expansion.
[0062] Step S103: Configure the target resource to obtain the configuration result of the target resource.
[0063] In practical applications, after the target host obtains the target resources of the first number of resource devices in the target computing architecture based on the consensus interconnection protocol, it can configure the target resources and obtain the configuration results of the target resources, so as to uniformly manage and control the target resources with the help of the configuration results.
[0064] Step S104: Process the data processing request through the target resource according to the configuration result.
[0065] In practical applications, after configuring the target resource and obtaining the configuration result, the target host can process data processing requests through the target resource according to the configuration result. For example, when the target resource is a computing resource, the target host can configure the computing resources with unified numbers to obtain the numbering configuration results for each computing resource. Then, when a certain computing resource needs to be used, the data processing request only needs to be sent to the computing resource with the corresponding number, and the corresponding working relationship between the two can be clearly defined by simply recording the data processing request and the number of the computing resource.
[0066] This application provides a resource management method applied to a target host. The method involves: acquiring data processing requests; acquiring target resources from a first number of resource devices in the target computing architecture based on a consistency interconnection protocol, where the first number is greater than or equal to 2; configuring the target resources to obtain configuration results; and processing the data processing requests using the target resources according to the configuration results. The resource devices in the target computing architecture are interconnected via the consistency interconnection protocol. In this application, the resource devices in the target computing architecture are interconnected via the consistency interconnection protocol. This allows the target host to uniformly configure the target resources after acquiring them from multiple resource devices, and to apply the multiple target resources to process data processing requests according to the configuration results, thereby achieving efficient collaboration between resource devices in the computing architecture.
[0067] Please see Figure 5 , Figure 5 This is a second flowchart of a resource management method provided in an embodiment of this application.
[0068] This application provides a resource management method applied to a target host, which may include the following steps:
[0069] Step S201: Obtain data processing request.
[0070] Step S202: Based on the consensus interconnection protocol, obtain the target resources of a first number of resource devices in the target computing architecture, and the target resources include target memory resources, and the value of the first number is greater than or equal to 2.
[0071] Step S203: Split the target memory resource to obtain target memory sub-resources.
[0072] Step S204: Perform unified addressing on the target memory sub-resources to obtain the configuration result.
[0073] In practical applications, when the target resource includes the target memory resource, the target host can split the target memory resource into fine-grained target memory sub-resources during the configuration process of the target resource. The target memory sub-resources can then be uniformly addressed to obtain the configuration results. This allows subsequent data processing requests to be processed based on the fine-grained target memory sub-resources, thereby improving the utilization rate of memory resources.
[0074] For ease of understanding, let's assume the interconnection diagram between memory resources is as follows: Figure 6 As shown, D1, D2, ... D7 belong to the same physical memory module, which, after partitioning, is equivalent to 6 independent memory modules. Alternatively, multiple memory entities can be integrated to achieve the performance of a larger memory module. For example, the integration of the two independent memory modules D2 and D1-D7 using the SCMP protocol is equivalent to a single large memory module. The workflow of memory resources can then be as follows:
[0075] 1. Memory Resource Sharing: CPUs can act as hosts, or servers, while memory modules and accelerators are SCMP devices. Different colors in the diagram represent the device owners. For example, D1 and accelerator B belong to CPU1. A single device, D1, can only be allocated to one CPU, i.e., a server. A complete resource release and reallocation process is as follows: Assuming that after a period of time, CPU1 no longer needs device D1, after completing the computation task, D1 can be HOTRemove (without actually physically removing it), and D1 is released into the memory resource pool.
[0076] 2. Sharing of computing devices: Suppose that CPU2 needs more acceleration devices. After accessing the resource pool, it finds that accelerator A is idle. Accelerator A is then added to the CPU2 server via the SCMP.io, SCMP.mem, and SCMP.cacahe protocols. In other words, accelerator A is online on CPU2. This completes a full pooled resource allocation.
[0077] 3. Fine-grained memory sharing: All memory in D1 is allocated to CPU1; the entire persistent memory resource can also be divided into multiple sub-memory modules, and the CPU can be notified to request and use the resources. For example, if CPU2 needs a small amount of additional memory resources during computation, this can be achieved through fine-grained memory resource sharing. Sub-memory modules D2, D3, ... D7 can be assigned to CPU2 via HOT Add. This allocation granularity is finer and more flexible, providing significant flexibility for typical edge computing applications.
[0078] It should be noted that in specific application scenarios, when memory needs to be partitioned and consolidated, the physical principle can be as follows: The memory of the node where the processor resides is called local memory, including: ① DRAM and persistent memory connected to the processor via the memory bus, and ② accelerator device memory mapped to the system memory space through the compute architecture coherence interconnect protocol. A tree structure is used to hierarchically and uniformly address pooled memory resources, with the memory management subsystem as the virtual root node and compute nodes or memory nodes as secondary nodes, establishing a tree structure. For actual physical addressing, the rack unique ID is used as the rack base address (2 bits RBase), and the unique ID of the node within the rack is used as the node base address (2 bits NBase). The physical address consists of "RBase + NBase + node local hardware address LPB". Each node deploys a Pooling Memory Address Translation Logic (PMATL) (implemented in the DPU) to maintain a global address mapping table and for cross-node access address translation. When memory expansion and contraction occur, tree structures are added / deleted in the global address space; when node-level expansion and contraction occur, node structures are added / deleted in the rack address space; and the local address space within a node is managed by the operating system; when memory resource expansion and contraction occur, the memory management subsystem updates the global address mapping tables in all PMATLs. For example, a diagram illustrating pooled memory resource addressing and load / store operation addressing can be shown as follows: Figure 7 As shown.
[0079] Step S205: According to the configuration results, process the data processing request through the target resource; wherein, the resource devices in the target computing architecture are connected through a consistency interconnection protocol.
[0080] It should be noted that while using a massive page mechanism in the memory pool can reduce TLB misses and address translation overhead, massive page migration can sometimes cause significant CPU wait overhead and memory bandwidth overhead. Fine-grained page migration, on the other hand, can disrupt address continuity and affect the massive page mechanism. To balance fine-grained page migration and the massive page mechanism, this application proposes a hybrid-granularity page management mechanism consisting of page monitoring technology, hot data-aware placement technology, and a hybrid-granularity page migration mechanism. The three technologies are described below:
[0081] Page monitoring: The CPUs' Precise Event-Based Sampling (PEBS) mechanism is used to sample LLC and TLB misses, obtain recent load / store operation information in the virtual address space, generate page usage heatmaps, and realize real-time monitoring of pages across memory levels.
[0082] Hot data-aware placement: Based on page monitoring heatmaps and task runtime characteristics, the best page placement strategy is intelligently and dynamically selected. The lightweight allocation mechanism moves hot pages to the fast memory layer and keeps cold pages in the slow memory layer, reducing the cost of accessing high-frequency hot data.
[0083] Hybrid-granularity page migration mechanism: Based on the SCMP protocol, reinforcement learning methods are used to analyze features such as CPU utilization, memory utilization, memory bandwidth, inter-node communication latency, and page migration overhead of different granularities, to implement a hybrid-granularity page allocation and migration strategy, reducing page monitoring and address translation overhead.
[0084] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a resource management system provided in an embodiment of this application.
[0085] This application provides a resource management system applied to a target host, which may include:
[0086] The first acquisition module 101 is used to acquire data processing requests;
[0087] The second acquisition module 102 is used to acquire the target resources of a first number of resource devices in the target computing architecture based on the consensus interconnection protocol, and the value of the first number is greater than or equal to 2.
[0088] The first configuration module 103 is used to configure the target resource and obtain the configuration result of the target resource;
[0089] The first processing module 104 is used to process data processing requests through the target resources according to the configuration results; wherein, the resource devices in the target computing architecture are connected through a consistency interconnection protocol.
[0090] This application provides a resource management system applied to a target host. The consistency interconnection protocol includes a first sub-protocol for transmitting physical signals to resource devices, a second sub-protocol for maintaining memory consistency in resource devices, and a third sub-protocol for resource devices to access memory consistency from the target host.
[0091] This application provides a resource management system applied to a target host, wherein the target computing architecture includes a first computing architecture to which the target host belongs;
[0092] The second acquisition module may include:
[0093] The first acquisition unit is used to acquire the target resources of a first number of resource devices in the target computing node to which the target host belongs, and the target computing node includes the computing node to which the target host belongs in the first computing architecture.
[0094] This application provides a resource management system applied to a target host, wherein the second acquisition module may include:
[0095] The second acquisition unit is used to acquire the target resources of a first number of resource devices in other computing nodes, wherein the other computing nodes include computing nodes in the first computing architecture other than the target computing node.
[0096] In this system, the resource devices in the target computing node are connected to the resource devices in other computing nodes through a virtual high-speed bus bridge.
[0097] This application provides a resource management system applied to a target host, wherein the target computing architecture includes a second computing architecture in addition to a first computing architecture.
[0098] The second acquisition module may include:
[0099] The third acquisition unit is used to acquire the target resources of a first number of resource devices in the second computing architecture;
[0100] The first computing architecture and the second computing architecture are connected through a switch and interconnection devices. The interconnection devices include a virtual variable capacity memory connected to the first computing architecture or the second computing architecture, a first virtual cache pool and an address translator connected to the virtual variable capacity memory, a virtual memoryless accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memoryless accelerator.
[0101] This application provides a resource management system applied to a target host, where the target resources include target memory resources;
[0102] The first configuration module may include:
[0103] The first splitting unit is used to split the target memory resource to obtain the target memory sub-resources;
[0104] The first addressing unit is used to uniformly address the target memory sub-resources to obtain the configuration result.
[0105] This application provides a resource management system applied to a target host, where the resource devices include edge computing nodes and / or hybrid memory nodes and / or heterogeneous computer nodes and / or ASIC computing nodes;
[0106] Among them, the hybrid memory node includes a persistent memory resource pool consisting of CPUs and persistent memory; the heterogeneous computing node includes a computing resource module consisting of CPUs and GPUs; and the ASIC computing node includes a computing resource module consisting of heterogeneous computing resources of CPUs and ASICs.
[0107] This application also provides a resource management device and a computer-readable storage medium, both of which have the corresponding effects of a resource management method provided in the embodiments of this application. Please refer to... Figure 9 , Figure 9 This is a schematic diagram of the structure of a resource management device provided in an embodiment of this application.
[0108] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program, and the processor 202 executes the computer program to perform the following steps:
[0109] Get data processing request;
[0110] Based on the consensus interconnection protocol, obtain the target resources of the first number of resource devices in the target computing architecture, and the value of the first number is greater than or equal to 2;
[0111] Configure the target resource to obtain the configuration result of the target resource;
[0112] Based on the configuration results, the data processing request is processed through the target resource;
[0113] In this target computing architecture, resource devices are interconnected through a consistent interconnection protocol.
[0114] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program. When the processor 202 executes the computer program, it implements the following steps: a consistency interconnection protocol includes a first sub-protocol for transmitting physical signals to the resource device, a second sub-protocol for maintaining memory consistency in the resource device, and a third sub-protocol for the resource device to access memory consistency from the target host.
[0115] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program. When the processor 202 executes the computer program, it performs the following steps: the target computing architecture includes a first computing architecture to which the target host belongs; the target resources of a first number of resource devices in the target computing node to which the target host belongs are obtained, and the target computing node includes the computing node to which the target host belongs in the first computing architecture.
[0116] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program. When the processor 202 executes the computer program, it performs the following steps: acquiring target resources of a first number of resource devices in other computing nodes. The other computing nodes include computing nodes in a first computing architecture other than the target computing node. The resource devices in the target computing node are connected to resource devices in other computing nodes through a virtual high-speed bus bridge.
[0117] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program. When the processor 202 executes the computer program, it performs the following steps: the target computing architecture includes a second computing architecture other than a first computing architecture; the target resources of a first number of resource devices in the second computing architecture are acquired; wherein the first computing architecture and the second computing architecture are connected through a switch and interconnection devices; the interconnection devices include a virtual variable capacity memory connected to the first computing architecture or the second computing architecture, a first virtual cache pool and an address translator connected to the virtual variable capacity memory, a virtual memory-free accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memory-free accelerator.
[0118] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program, and when the processor 202 executes the computer program, it performs the following steps: the target resource includes a target memory resource; the target memory resource is split to obtain target memory sub-resources; the target memory sub-resources are uniformly addressed to obtain a configuration result.
[0119] This application provides a resource management device, including a memory 201 and a processor 202. The memory 201 stores a computer program, and the processor 202 executes the computer program to implement the following steps: the resource device includes edge computing nodes and / or hybrid memory nodes and / or heterogeneous computer nodes and / or ASIC computing nodes; wherein, the hybrid memory node includes a persistent memory resource pool composed of CPU and persistent memory; the heterogeneous computing node includes a computing resource module composed of CPU and GPU computing; and the ASIC computing node includes a computing resource module composed of CPU and ASIC heterogeneous computing resources.
[0120] Please see Figure 10 Another resource management device provided in this application embodiment may further include: an input port 203 connected to the processor 202 for transmitting commands input from the outside to the processor 202; a display unit 204 connected to the processor 202 for displaying the processing results of the processor 202 to the outside; and a communication module 205 connected to the processor 202 for enabling communication between the resource management device and the outside. The display unit 204 may be a display panel, a laser scanner, or the like; the communication method used by the communication module 205 includes, but is not limited to, Mobile High Definition Link (HML), Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), wireless connection: Wi-Fi, Bluetooth communication, Bluetooth Low Energy communication, and IEEE 802.11s-based communication technology.
[0121] This application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the following steps:
[0122] Get data processing request;
[0123] Based on the consensus interconnection protocol, obtain the target resources of the first number of resource devices in the target computing architecture, and the value of the first number is greater than or equal to 2;
[0124] Configure the target resource to obtain the configuration result of the target resource;
[0125] Based on the configuration results, the data processing request is processed through the target resource;
[0126] In this target computing architecture, resource devices are interconnected through a consistent interconnection protocol.
[0127] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the following steps: a consistency interconnection protocol includes a first sub-protocol for transmitting physical signals to resource devices, a second sub-protocol for maintaining memory consistency in resource devices, and a third sub-protocol for resource devices to access memory consistency from a target host.
[0128] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the following steps: the target computing architecture includes a first computing architecture to which the target host belongs; the target resources of a first number of resource devices in the target computing node to which the target host belongs are obtained, and the target computing node includes the computing node to which the target host belongs in the first computing architecture.
[0129] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the following steps: acquiring target resources of a first number of resource devices in other computing nodes, wherein the other computing nodes include computing nodes in a first computing architecture other than the target computing node; wherein the resource devices in the target computing node are connected to resource devices in other computing nodes through a virtual high-speed bus bridge.
[0130] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the following steps: a target computing architecture includes a second computing architecture other than a first computing architecture; acquiring target resources of a first number of resource devices in the second computing architecture; wherein the first computing architecture and the second computing architecture are connected through a switch and interconnection devices; the interconnection devices include a virtual variable-capacity memory connected to the first computing architecture or the second computing architecture, a first virtual cache pool and an address translator connected to the virtual variable-capacity memory, a virtual memory-free accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memory-free accelerator.
[0131] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the following steps: the target resource includes a target memory resource; the target memory resource is split to obtain target memory sub-resources; the target memory sub-resources are uniformly addressed to obtain a configuration result.
[0132] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the following steps: the resource device includes an edge computing node and / or a hybrid memory node and / or a heterogeneous computing node and / or an ASIC computing node; wherein, the hybrid memory node includes a persistent memory resource pool composed of a CPU and persistent memory; the heterogeneous computing node includes a computing resource module composed of a CPU and a GPU; and the ASIC computing node includes a computing resource module composed of heterogeneous computing resources of CPU and ASIC.
[0133] The computer-readable storage media involved in this application include random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage media known in the art.
[0134] For descriptions of relevant parts of the resource management system, device, and computer-readable storage medium provided in the embodiments of this application, please refer to the detailed descriptions of the corresponding parts in the resource management method provided in the embodiments of this application, which will not be repeated here. Furthermore, parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.
[0135] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0136] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A resource management method, characterized in that, Applied to the target host, including: Get data processing request; Based on the Consistent Interconnect Protocol, target resources of a first number of resource devices in the target computing architecture are obtained, and the value of the first number is greater than or equal to 2; the Consistent Interconnect Protocol includes a first sub-protocol for transmitting physical signals to the resource devices, a second sub-protocol for maintaining memory consistency in the resource devices, and a third sub-protocol for the resource devices to access memory consistency from the target host; the target computing architecture includes a first computing architecture to which the target host belongs, and a second computing architecture other than the first computing architecture; The step of acquiring the target resources of a first number of resource devices in the target computing architecture includes: acquiring the target resources of the first number of resource devices in the target computing node to which the target host belongs, wherein the target computing node includes the computing node to which the target host belongs in the first computing architecture; acquiring the target resources of the first number of resource devices in other computing nodes, wherein the other computing nodes include computing nodes in the first computing architecture other than the target computing node; the resource devices in the target computing node are connected to the resource devices in the other computing nodes through a virtual high-speed bus bridge; acquiring the target resources of the first number of resource devices in the second computing architecture; the first computing architecture and the second computing architecture are connected through a switch and interconnection devices; the interconnection devices include a virtual variable capacity memory connected to the first computing architecture or the second computing architecture, and connected to the virtual variable capacity memory. The system comprises a first virtual cache pool and an address translator, a virtual memory-free accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memory-free accelerator. The interconnect device is connected to the first or second computing architecture via an ultra-high-speed bus interface and an ultra-high-speed bus, and to the switch via a high-speed network interface and a high-speed network. When the memory of a device on the second computing architecture side is managed by the central processing unit on the first computing architecture side, the memory of the second computing architecture and the memory of the first computing architecture are uniformly addressed and managed. When the CPU of the first computing architecture uses the memory of the second computing architecture, the virtual memory-free accelerator module caches the memory data of the second computing architecture into the virtual cache pool (MOB), realizing local caching of remote data. The virtual variable-capacity memory of the first computing architecture reads the data from the virtual cache pool on the second computing architecture side into the local virtual cache pool on the first computer architecture side. Configure the target resource to obtain the configuration result of the target resource; According to the configuration results, the data processing request is processed through the target resource; The resource devices in the target computing architecture are interconnected via the consensus interconnection protocol; the resource devices include edge computing nodes and / or hybrid memory nodes and / or heterogeneous computer nodes and / or ASIC computing nodes; the hybrid memory node includes a persistent memory resource pool composed of CPUs and persistent memory; the heterogeneous computer node includes a computing resource module composed of CPUs and GPUs; the ASIC computing node includes a computing resource module composed of heterogeneous computing resources of CPUs and ASICs.
2. The method according to claim 1, characterized in that, The target resource includes target memory resources; The configuration of the target resource to obtain the configuration result of the target resource includes: The target memory resource is split into target memory sub-resources; The target memory sub-resources are uniformly addressed to obtain the configuration result.
3. A resource management system, characterized in that, Applied to the target host, including: The first acquisition module is used to acquire data processing requests; The second acquisition module is configured to acquire target resources of a first number of resource devices in a target computing architecture based on a consistency interconnection protocol, wherein the first number is greater than or equal to 2; the consistency interconnection protocol includes a first sub-protocol for transmitting physical signals to the resource devices, a second sub-protocol for maintaining memory consistency in the resource devices, and a third sub-protocol for the resource devices to access memory consistency from the target host; the target computing architecture includes a first computing architecture to which the target host belongs, and a second computing architecture other than the first computing architecture; wherein, acquiring the target resources of the first number of resource devices in the target computing architecture includes: acquiring the target resources of the first number of resource devices in the target computing node to which the target host belongs, wherein the target computing node includes the computing node to which the target host belongs in the first computing architecture. The system acquires the target resources of a first number of resource devices in other computing nodes, including computing nodes in the first computing architecture other than the target computing node; the resource devices in the target computing node are connected to the resource devices in the other computing nodes via a virtual high-speed bus bridge; acquires the target resources of the first number of resource devices in the second computing architecture; the first computing architecture and the second computing architecture are connected via a switch and interconnection devices; the interconnection devices include a virtual variable-capacity memory connected to the first computing architecture or the second computing architecture, a first virtual cache pool and an address translator connected to the virtual variable-capacity memory, a virtual memoryless accelerator connected to the address translator and the switch, and a second virtual cache pool connected to the virtual memoryless accelerator; The first configuration module is used to configure the target resource and obtain the configuration result of the target resource; The first processing module is used to process the data processing request through the target resource according to the configuration result; The resource devices in the target computing architecture are interconnected via the consensus interconnection protocol; the resource devices include edge computing nodes and / or hybrid memory nodes and / or heterogeneous computer nodes and / or ASIC computing nodes; the hybrid memory node includes a persistent memory resource pool composed of CPUs and persistent memory; the heterogeneous computer node includes a computing resource module composed of CPUs and GPUs; the ASIC computing node includes a computing resource module composed of heterogeneous computing resources of CPUs and ASICs.
4. A resource management device, characterized in that, include: Memory, used to store computer programs; A processor for implementing the steps of the resource management method as described in claim 1 or 2 when executing the computer program.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the resource management method as described in claim 1 or 2.
Citation Information
Patent Citations
Resource sharing device, resource management device, and resource management method
CN115586964A