Computing resource sharing system, method, device and equipment, medium and product

By generating and integrating global computing resource topology files and optimizing data transmission paths, the problem of low efficiency in multi-host computing resource sharing is solved, efficient and reliable computing resource sharing and communication is achieved, and larger-scale system expansion is supported.

CN120386636AActive Publication Date: 2025-07-29LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202510866080.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-29
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the prior art, the computing resource sharing efficiency between multiple hosts is low, mainly due to the high Ethernet transmission delay, which affects the efficiency of deep learning model training and other tasks.

Method used

By introducing the collaborative work of the resource management processor and controller, local computing resource topology files are generated and integrated, global computing resource topology files are built, data transmission paths are optimized, routing forwarding parameters and address mapping parameters in the switching module are configured to achieve efficient resource sharing and communication between multiple hosts.

Benefits of technology

It improves the efficiency and reliability of multi-host computing resource sharing, supports larger-scale computing resource sharing, improves the scalability and adaptability of the system, and ensures the accuracy and reliability of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386636A_ABST
    Figure CN120386636A_ABST
Patent Text Reader

Abstract

The invention discloses a computing resource sharing system, method, device and equipment, a medium and a product, and relates to the technical field of computers, the system comprises a plurality of hosts and a global communication module, each host comprises a plurality of computing units and a first switching module, the computing units in the same host are interconnected through the first switching module in the same host, and the global communication module is connected with the computing units in the same host. The global communication module comprises a second switching module and a resource management processor, and the plurality of hosts are interconnected through the second switching module; the host generates a local computing resource topology file; the resource management processor integrates the local computing resource topology files generated by the plurality of hosts to obtain a global computing resource topology file, and configures routing forwarding parameters in the second switching module based on the global computing resource topology file; and the host configures a routing forwarding parameter and an address mapping parameter in the first switching module based on the global computing resource topology file. According to the invention, the efficiency and reliability of multi-host computing resource sharing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a computing resource sharing system, method, device, equipment, medium, and product. Background Art

[0002] In the related art, computing units between multiple hosts are interconnected through Ethernet. However, the transmission delay of Ethernet is relatively high, which affects efficiency. For example, in the model training of deep learning, data and parameters need to be frequently exchanged between multiple computing units, and the high delay of Ethernet will result in low training efficiency. It can be seen that the resource sharing efficiency between multiple hosts in the related art is relatively low.

[0003] Therefore, how to achieve efficient sharing of computing resources between multiple hosts is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] The present invention provides a computing resource sharing system, method, device, equipment, medium, and product, which improves the efficiency and reliability of computing resource sharing among multiple hosts.

[0005] The present invention provides a computing resource sharing system, including multiple hosts and a global communication module. Each host includes multiple computing units and a first switching module. The computing units in the same host are interconnected through the first switching module in the same host. The global communication module includes a second switching module and a resource management processor. The multiple hosts are interconnected through the second switching module. The host is configured to: generate a local computing resource topology file and send the local computing resource topology file to the resource management processor; wherein, the local computing resource topology file is used to describe the identification information of the computing units in the host. The resource management processor is configured to: integrate the local computing resource topology files generated by the multiple hosts to obtain a global computing resource topology file, configure routing forwarding parameters in the second switching module based on the global computing resource topology file, and return the global computing resource topology file to the host. The host is further configured to: configure routing forwarding parameters and address mapping parameters in the first switching module based on the global computing resource topology file. The host is further configured to: access the computing units of itself or other hosts according to the address information in the memory address mapping information.

[0006] The present invention provides a computing resource sharing method, which is applied to a host in a computing resource sharing system. The method includes: generating a local computing resource topology file, and sending the local computing resource topology file to a resource management processor in a global communication module, so that the resource management processor integrates the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file, configures routing forwarding parameters in a second switching module in the global communication module based on the global computing resource topology file, and returns the global computing resource topology file to the host; wherein, the local computing resource topology file is used to describe the identification information of computing units in the host; configuring routing forwarding parameters and address mapping parameters in a first switching module in the host based on the global computing resource topology file; obtaining memory address mapping information, and accessing the computing units of itself or other hosts according to the address information of the memory address mapping information.

[0007] The present invention provides a computing resource sharing method, which is applied to a resource management processor in a global communication module in a computing resource sharing system. The method includes: receiving the local computing resource topology files sent by each host in the computing resource sharing system, integrating the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file; configuring routing forwarding parameters in a second switching module in the global communication module based on the global computing resource topology file; returning the global computing resource topology file to each host, so that each host configures routing forwarding parameters and address mapping parameters in a first switching module in each host based on the global computing resource topology file.

[0008] The present invention further provides a computing resource sharing device, which is applied to a host in a computing resource sharing system. The device includes: a generating module, configured to generate a local computing resource topology file, and send the local computing resource topology file to a resource management processor in a global communication module, so that the resource management processor integrates the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file, configures routing forwarding parameters in a second switching module in the global communication module based on the global computing resource topology file, and returns the global computing resource topology file to the host; wherein, the local computing resource topology file is used to describe the identification information of computing units in the host; a first configuration module, configured to configure routing forwarding parameters and address mapping parameters in a first switching module in the host based on the global computing resource topology file; an obtaining module, configured to obtain memory address mapping information, and access the computing units of itself or other hosts according to the address information of the memory address mapping information.

[0009] The present invention also provides a computing resource sharing device, which is applied to a resource management processor in a global communication module of a computing resource sharing system. The device includes: an integration module, configured to receive local computing resource topology files sent by each host in the computing resource sharing system, and integrate the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file; a second configuration module, configured to configure routing forwarding parameters in a second switching module in the global communication module based on the global computing resource topology file; a sending module, configured to return the global computing resource topology file to each host, so that each host configures routing forwarding parameters and address mapping parameters in a first switching module in each host based on the global computing resource topology file.

[0010] The present invention also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above computing resource sharing methods when executing the computer program.

[0011] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above computing resource sharing methods are implemented.

[0012] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above computing resource sharing methods are implemented.

[0013] The beneficial effects of the present invention are as follows: The resource sharing system provided by the present invention realizes the topological interconnection of multiple hosts in an out-of-band manner, and through the collaborative work of the resource management processor and the controller, realizes the generation and integration of the local computing resource topology file, thereby constructing a global computing resource topology file, providing a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the routing forwarding parameters and address mapping parameters in the first switching module and the routing forwarding parameters in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In addition, through multiple second switching modules in the global communication module, the present invention can be flexibly extended to the case of multiple switching modules, supporting a larger-scale computing resource sharing and enhancing the scalability of the system. In summary, the present invention improves the efficiency and reliability of multi-host computing resource sharing. The present invention also discloses a computing resource sharing method, device, an electronic device, a non-volatile storage medium, and a computer program product, which can also achieve the above technical effects.

[0014] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit the present invention. Brief Description of the Drawings

[0015] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is an architecture diagram of a computing resource sharing system in the related art.

[0017] Figure 2 It is a structural diagram of a computing resource sharing system shown according to an exemplary embodiment.

[0018] Figure 3 It is a structural diagram of another computing resource sharing system shown according to an exemplary embodiment.

[0019] Figure 4 It is an architecture diagram of a hardware topology management shown according to an exemplary embodiment.

[0020] Figure 5 It is a schematic diagram of memory address mapping information shown according to an exemplary embodiment.

[0021] Figure 6 It is a schematic diagram of a software topology management solution shown according to an exemplary embodiment.

[0022] Figure 7 It is a schematic diagram of the link connection involved when host 0 accesses the computing unit in host 1.

[0023] Figure 8 It is a flowchart of a computing resource sharing method shown according to an exemplary embodiment.

[0024] Figure 9 It is a flowchart of a topology routing configuration management shown according to an exemplary embodiment.

[0025] Figure 10 It is a flowchart of another computing resource sharing method shown according to an exemplary embodiment.

[0026] Figure 11 It is a structural diagram of a computing resource sharing device shown according to an exemplary embodiment.

[0027] Figure 12 It is a structural diagram of another computing resource sharing device shown according to an exemplary embodiment.

[0028] Figure 13 Structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0032] In the related art, NTB (Non-Transparent Bridge) is used to implement multi-host sharing of computing resources. As Figure 1 shown, the upstream port 501 of the first HOST (host) domain sends a first packet to the routing module. The routing module 502, in response to receiving the first packet sent by the upstream port of the first HOST domain, confirms whether the destination port of the first packet is mounted with NTB. In response to a confirmation of yes, the first packet is sent to the NTB virtual side through the centralized switching module; the centralized switching module 503 forwards the first packet from the routing module to the NTB virtual side; the NTB virtual side 504 forwards the first packet to the conversion module, and the NTB virtual side is a virtual endpoint; the conversion module 505 converts the address and ID information of the first packet from the first HOST domain to the second HOST domain, generates a second packet, and forwards the second packet through the routing module and the centralized switching module to the NTB link side. The NTB link side is an endpoint device physically connected to the second HOST domain in the first HOST domain; the routing module 502 and the centralized switching module 503 forward the second packet from the conversion module to the NTB link side; the NTB link side 506 receives the second packet.

[0033] In the above solution, multi-host communication is implemented through a VS (Virtual Switch) in the same switching module, making it difficult to scale to the case of multiple switching modules. In addition, in the above solution, communication between different computing units needs to be forwarded through upstream ports, resulting in low communication efficiency.

[0034] To solve the above problems, the present invention realizes the topological interconnection of multiple hosts in an out-of-band manner. By introducing the collaborative work of a resource management processor and a controller, the generation and integration of a local computing resource topology file are realized, thereby constructing a global computing resource topology file, providing a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the address trap register and port routing forwarding register in the first switching module and the port routing forwarding register in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In addition, through multiple second switching modules in the global communication module, the present invention can be flexibly extended to the case of multiple switching modules, supporting a larger-scale computing resource sharing and enhancing the scalability of the system.

[0035] This embodiment provides a computing resource sharing system, as Figure 2 shown, which includes multiple hosts 10 and a global communication module 20. Each host 10 includes multiple computing units 101 and a first switching module 102. The computing units 101 in the same host are interconnected through the first switching module 102 in the same host. The global communication module 20 includes a resource management processor 201 and a second switching module 202. The multiple hosts 10 are interconnected through the second switching module 202.

[0036] In this embodiment, the computing resource sharing system includes multiple hosts and a global communication module. In addition to multiple computing units 101 and a first switching module 102, host 10 may further include a processor and a controller. The computing unit may be a GPU (Graphics Processing Unit), and the controller may be a BMC (Baseboard Management Controller). The global communication module 20 includes a resource management processor 201 and multiple second switching modules 202. The first switching module 102 in host 10 is connected to the second switching module 202 in the global communication module 20. The first switching module and the second switching module may adopt a PCIe (Peripheral Component Interconnect Express) Switch, that is, the first switching module is connected to the computing unit through a PCIe bus, and the first switching module is connected to the second switching module through a PCIe bus.

[0037] In terms of the component connection relationship, the processor in the host is connected to the controller, and the controller is connected to multiple computing units through the first switching module, thereby constructing a computing resource network inside the host. The first switching module in the host can be connected to the second switching module in the global communication module through a Retimer device, enabling the processor inside the host to access the computing units inside other hosts through the first switching module and the second switching module, and also enabling the computing units inside different hosts to communicate through the first switching module and the second switching module. Inside the global communication module, the resource management processor can be connected to multiple second switching modules through a Uart (Universal Asynchronous Receiver / Transmitter) or an SDB (System Debug Bus). In addition, the resource management processor can be connected to the controllers in each host through a LAN (Local Area Network), making the resource management processor the management node of the entire system to collaboratively manage the controllers and other components in each host.

[0038] As a preferred embodiment, the host includes a first preset number of first switching modules. The first switching modules are connected to the controllers in the host, and each first switching module is connected to a second preset number of computing units. The global communication module includes a third preset number of second switching modules. The first switching modules are connected to the ports of the second switching modules, and the ports of the second switching modules connected by any two first switching modules are different.

[0039] In a specific implementation, each host includes one or more first switching modules, each first switching module is respectively connected to a plurality of different computing units, the global communication module includes one or more second switching modules, and the first switching module is connected to the port of the second switching module. The number of second switching modules included in the global communication module can be determined by the number of ports supported by the second switching module, the number of hosts included in the system, and the number of first switching modules included in each host. The sum of the number of ports included in all second switching modules in the global communication module needs to be greater than or equal to the sum of the number of all first switching modules in all hosts to ensure that the ports of the second switching modules connected by any two first switching modules are different.

[0040] For example, as Figure 3 shown, the computing resource sharing system includes 8 hosts, each host includes 4 first switching modules, each first switching module is connected to 2 computing units, that is, each host includes 8 computing units, the global communication module includes 4 second switching modules, each second switching module includes 8 ports, the first switching module 0 in each host is connected to a port of the second switching module 0, the first switching module 1 in each host is connected to a port of the second switching module 1, the first switching module 2 in each host is connected to a port of the second switching module 2, and the first switching module 3 in each host is connected to a port of the second switching module 3.

[0041] The hardware topology management architecture is as Figure 4 shown. The controller can be connected to the first switching module through I2C, and can perform out-of-band configuration management on the first switching module through this link. The resource management processor is located in the global communication module and serves as the management node of the entire system, coordinating the management of BMC, CPLD (Complex Programmable Logic Device), and switching module FW (Firmware). Multiple hosts are interconnected through the second switching module. It is necessary to decouple the computing unit, switching module, and processor logic in terms of timing, and control the coordinated power-on and power-off through the management software in the resource management processor.

[0042] Host 10 is used to: generate a local computing resource topology file and send the local computing resource topology file to the resource management processor 201; wherein, the local computing resource topology file is used to describe the identification information of the computing units in the host.

[0043] The resource management processor 201 is used to: integrate the local computing resource topology files generated by multiple hosts 10 to obtain a global computing resource topology file, configure the routing forwarding parameters in the second switching module based on the global computing resource topology file, and return the global computing resource topology file to the host 10.

[0044] The host 10 is further configured to: configure the routing forwarding parameters and address mapping parameters in the first switching module based on the global computing resource topology file.

[0045] The host 10 is further configured to: access its own or other hosts' computing units according to the address information in the memory address mapping information.

[0046] As a feasible implementation manner, the host is further configured to: allocate memory addresses for the computing units in the local host, and reserve placeholder memory addresses for the computing units in other hosts; perform device enumeration and resource allocation for the computing units in the local host, so as to allocate switching module domain addresses for the computing units in the local host, and map the memory addresses of the computing units to the switching module domain addresses.

[0047] In a specific implementation, when each host is started, the first switching module inside the host performs device enumeration and resource allocation for the computing units, creates switching module domain addresses for each computing unit, and the switching module domain addresses between each host are independent. In addition to allocating memory addresses for the local computing units, the BIOS (Basic Input / Output System) in each host also reserves a section of addresses for the computing units of other hosts, that is, placeholder memory addresses. Subsequently, the first switching module performs address mapping to establish a mapping between the memory addresses and the switching module domain addresses allocated by the first switching module. After that, the operating system (OS) of each host can access the computing units of all hosts through the memory addresses. Taking Figure 3 the shown computing resource sharing system as an example, the created memory address mapping information is as Figure 5 shown. By allocating memory addresses for the computing units by the BIOS and reserving placeholder memory addresses, and by the first switching module performing device enumeration and resource allocation for the computing units when the host is started, the dynamic allocation and management of computing resources are realized. This mechanism enables the system to flexibly respond to the dynamic changes of the computing units, improves the adaptability and flexibility of the system, and at the same time provides more efficient support for cross-host resource sharing and communication.

[0048] When implementing resource sharing among multiple hosts, the controller in the host is responsible for generating a local computing resource topology file, which details the bus identifiers of the computing units within the host, as well as the domain identifier and bus identifier from the perspective of the second switching module. The resource management processor generates a global computing resource topology file by collecting and integrating the local topology files of multiple hosts, and configures the port routing and forwarding registers in the second switching module accordingly. The controller further configures the address trap register and port routing and forwarding registers in the first switching module based on the global topology file, thus achieving seamless docking of intra-host and cross-host communications. Finally, the processor obtains the memory address mapping information through the controller and accesses local or remote computing units based on this information, realizing dynamic sharing and efficient utilization of computing resources.

[0049] As a feasible implementation, the local computing resource topology file is used to describe the bus identifier of the computing unit from the perspective of the local host, as well as the domain identifier and bus identifier from the perspective of the second switching module. The process by which the host generates the local computing resource topology file of the host where it is located includes: obtaining the asset information of the computing resources and populating the local bus identifier field in the local computing resource topology file based on the asset information; where the asset information includes the resource information of the computing units in the local host and the placeholder memory address information of the computing units in other hosts, and the local bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the local host; accessing the first switching module in the local host to obtain the memory address mapping information and global domain information of the computing resources, and populating the global domain identifier field and global bus identifier field in the local computing resource topology file based on the memory address mapping information and global domain information; where the global domain identifier field is used to describe the domain identifier of the computing units in the local host from the perspective of the second switching module, and the global bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the second switching module.

[0050] In a specific implementation, the controller obtains the asset information of computing resources, including the resource information of computing units in the local host and the placeholder memory address information of computing units in other hosts. As a feasible implementation, the basic input / output system in the host collects the asset information and sends the asset information to the controller through the H2B (Host to BIOS) shared memory. As the core component for system startup and initialization, the BIOS can collect the hardware resource information in the host more comprehensively, ensuring the accuracy and integrity of the asset information, thereby improving the stability and reliability of the entire resource sharing system. The controller fills the local bus identification field in the local computing resource topology file based on these asset information, and this field is used to describe the bus identification of the computing unit in the local host from the local perspective. Next, the controller accesses the first switching module in the local host to obtain the memory address mapping information and global domain information of the computing resources. As a feasible implementation, the controller accesses the first switching module in the host where it is located through the Management Component Transport Protocol over System Management Bus (MCTP over SMBus). The memory address mapping information is a detailed description of the location of the computing resources in the memory and how to access them, while the global domain information provides the location information of the computing resources in the entire system. The controller fills the global domain identification field and the global bus identification field in the local computing resource topology file based on these information, and these fields respectively describe the domain identification and bus identification of the computing unit in the local host from the perspective of the second switching module. It can be seen that this implementation improves the identifiability and manageability of computing resources, enabling the resource management processor to more effectively integrate and configure the global computing resource topology file, optimizing the cross-host communication path, and enhancing the system communication efficiency and performance.

[0051] On this basis, the global computing resource topology file includes the bus identification of each computing unit from the perspective of the host where it is located, the domain identification and bus identification from the perspective of the second switching module. As a preferred implementation, the global computing resource topology file further includes any one or several combinations of the address mapping field of each computing unit, the number of hosts included in the computing resource sharing system, the number of computing units, and the file length of the global computing resource topology file. The address mapping field is used to describe whether the address mapping of the computing unit is set up.

[0052] Taking Figure 3 the computing resource sharing system shown as an example, the global computing resource topology file is shown in Table 1: Table 1 Among them, Checksum is the checksum of the file, used to check the file integrity; Version is the major version number of the file; Host Number is the number of hosts included in the system; GPU Number is the number of computing units included in the system; Topo Length is the file length of the file; HxGx-Local-Bus is the bus identifier (Bus Number) of the x-th computing unit in the x-th host from the perspective of the host; HxGx-Global-Domain is the domain identifier (Domain ID) of the x-th computing unit in the x-th host from the perspective of the second switching module. The first 4 bits represent the host identifier (Host ID), and the last 4 bits represent the first switching module identifier (Switch ID); HxGx-Global-Bus is the bus identifier of the x-th computing unit in the x-th host from the perspective of the second switching module; HxGx-Mapping-Flag is the address mapping field of the x-th computing unit in the x-th host, used to describe whether the address mapping has been set. If it is 0, the setting has not been completed yet. If it is 1, the setting has been completed and it can be used normally; UUID is the host identifier.

[0053] As a feasible implementation manner, the process that the resource management processor configures the routing and forwarding parameters in the second switching module based on the global computing resource topology file includes: the resource management processor configures the port routing and forwarding register in the second switching module based on the global computing resource topology file. The process that the controller configures the routing and forwarding parameters and address mapping parameters in the first switching module based on the global computing resource topology file includes: the controller configures the port routing and forwarding register and address trap register in the first switching module based on the global computing resource topology file.

[0054] In specific implementation, the first switching module includes a port routing and forwarding register and an address trap register. The port routing and forwarding register and address trap register in the first switching module are configured based on the global computing resource topology file to implement the configuration of the routing and forwarding parameters and address mapping parameters in the first switching module. The second switching module includes a port routing and forwarding register. The port routing and forwarding register in the second switching module is configured based on the global computing resource topology file to implement the configuration of the routing and forwarding parameters in the second switching module.

[0055] As a feasible implementation manner, the process that the host configures the port routing and forwarding register and address trap register in the first switching module based on the global computing resource topology file includes: the controller in the host accesses the first switching module in the host to configure the port routing and forwarding register and address trap register in the first switching module based on the physical slot corresponding to each computing unit and the global computing resource topology file.

[0056] In a specific implementation, the controller configures the address trap register and the port routing and forwarding register in the first switching module by accessing the first switching module in the host where it is located, using the information in the global computing resource topology file and the physical slot corresponding to the computing unit. The address trap register is used to capture and process access requests for specific addresses, while the port routing and forwarding register is responsible for correctly routing data packets to the target computing unit according to the address information. This configuration method ensures the accuracy and efficiency of data transmission, optimizes network traffic, and reduces errors and delays.

[0057] As a feasible implementation, the processor obtains memory address mapping information from the baseboard management controller through in-band commands to access the computing units of its own host or other hosts according to the address information in the memory address mapping information.

[0058] In a specific implementation, the processor communicates with the baseboard management controller by executing in-band commands to obtain memory address mapping information. In-band commands refer to commands issued through the operating system to obtain hardware information when the operating system is running, such as KCS (Keyboard Controller Style) physical channel in-band IPMI (Intelligent Platform Management Interface) commands. The memory address mapping information includes the address identifiers of the computing units in the system, enabling the operating system client running on the processor to access each computing unit in each host. The use of in-band commands enables the processor to directly communicate with the BMC, quickly obtain the required memory address mapping information, thereby improving the system's response speed and operating efficiency, and ensuring that the processor can access each computing unit in a timely and accurate manner.

[0059] The software topology management solution is as Figure 6 shown. The basic input / output system is responsible for: obtaining the resource information of the local computing units and collecting the placeholder memory address information of the computing units in other hosts, and pushing them to the BMC in the form of asset information through the H2B shared memory.

[0060] The baseboard management controller is responsible for: parsing the resource information of the local computing unit in the asset information and the placeholder memory address information of the computing units in other hosts; accessing multiple first switching modules via MCTP over SMBus to obtain the memory address mapping information of each computing unit in the first switching modules; providing a Redfish (an open standard protocol for data center management) interface for the management software in the resource management processor to obtain the local computing resource topology file; providing a Redfish interface for the management software in the resource management processor to set the global computing resource topology file; accessing multiple first switching modules via MCTP over SMBus to configure the address trap register according to the Slot silk screen position information of different computing units corresponding to each first switching module, that is, the slot information; accessing multiple first switching modules via MCTP over SMBus to configure the port routing and forwarding register of the first switching module according to the Slot silk screen position information of different computing units corresponding to each first switching module; providing IPMI commands for the in-band operating system client (OS Client) to obtain the memory address mapping information in the first switching module.

[0061] The management software in the resource management processor is responsible for: setting the port routing and forwarding register in the second switching module in the global communication module via UART; collecting the local computing resource topology files of the baseboard management controllers of each host via Redfish commands; integrating the local computing resource topology files of all hosts to form a global computing resource topology file; sending the global computing resource topology file to the baseboard management controllers of each host via Redfish commands to achieve topology configuration.

[0062] The operating system client is responsible for: synchronously obtaining the memory address mapping information in the first switching module via in-band IPMI commands.

[0063] It should be noted that in order to implement Peer-to-Peer (P2P) communication between different computing units, it is necessary to configure the Address Trap register and the Port Routing Forwarding (DLUT) register of the first switching module. In the Address Trap register information configured in this process, it is necessary to include the global domain identifier, global bus identifier, and address information of all computing units of other hosts. This information needs to be integrated and distributed to each host by the management software in the resource management processor. If it is implemented in the OS Client, since the management software is an out-of-band network and does not participate in in-band services, it will cause pollution of the out-of-band and in-band networks, and the register configuration process has nothing to do with the host operating system services. Therefore, selecting the baseboard management controller as the main body for the management of the first switching module can achieve a more pure and seamless configuration of the first switching module, thereby ensuring the efficient operation of computing resource sharing and the P2P access function between computing units.

[0064] As a feasible implementation manner, the computing resource sharing system includes N hosts. The process by which the i-th host accesses the target computing unit in the j-th host according to the memory address mapping information includes: the processor in the i-th host generates an access request based on the address information of the target computing unit in the memory address mapping information. The access request is captured by the Address Trap register in the first switching module of the i-th host. The Address Trap register converts the address information in the access request into the switching module domain address of the target computing unit. The Port Routing Forwarding register in the first switching module of the i-th host routes the access request to the target computing unit in the j-th host through the second switching module in the global communication module based on the switching module domain address; where 1 ≤ i ≤ N, 1 ≤ j ≤ N, and i ≠ j.

[0065] In a specific implementation, the processor in the i-th host first determines the address of the target computing unit according to the memory address mapping information, and then generates an access request. This access request will be captured by the Address Trap register of the first switching module in the i-th host. The register converts the address information in the request into the address of the target computing unit in the switching module domain. Then, the Port Routing Forwarding register in the first switching module uses this switching module domain address to accurately route the access request to the target computing unit in the j-th host through the second switching module in the global communication module. This process realizes cross-host computing resource sharing and improves the resource utilization rate and system flexibility. For example, when host 0 accesses the computing unit in host 1, the schematic diagram of the link connection involved is as Figure 7 shown.

[0066] The resource sharing system provided by the embodiments of the present invention realizes the topological interconnection of multiple hosts through an out-of-band method. By introducing the collaborative work of a resource management processor and a controller, the generation and integration of local computing resource topology files are realized, thereby constructing a global computing resource topology file, which provides a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the address trap register and port routing and forwarding register in the first switching module and the port routing and forwarding register in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In addition, through multiple second switching modules in the global communication module, the embodiments of the present invention can be flexibly extended to the case of multiple switching modules, supporting a larger-scale computing resource sharing and enhancing the scalability of the system. In summary, the embodiments of the present invention improve the efficiency and reliability of multi-host computing resource sharing.

[0067] Embodiments of the present invention provide a method for sharing computing resources. Combining with the execution process of the method for sharing computing resources, the method will be described in detail. Refer to Figure 8 , a flowchart of a method for sharing computing resources shown according to an exemplary embodiment.

[0068] S101: Generate a local computing resource topology file, and send the local computing resource topology file to a resource management processor in a global communication module, so that the resource management processor integrates the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file, configures routing and forwarding parameters in a second switching module in the global communication module based on the global computing resource topology file, and returns the global computing resource topology file to the host; wherein, the local computing resource topology file is used to describe the identification information of computing units in the host.

[0069] The execution entity of this embodiment is the host. In this step, the controller in the host generates a local computing resource topology file that describes the layout of computing resources within its host. This file contains information such as the bus identifier and domain identifier of computing units from the perspective of the local host, and is crucial for understanding the internal resource distribution of the host. The controller then sends this file to the resource management processor, which is responsible for integrating such files from multiple hosts to create a global computing resource topology file. This global file enables the resource management processor to understand the resource distribution across the entire system and configure the routing and forwarding parameters in the second switching module accordingly to optimize the data transmission path. As a feasible implementation, the port routing and forwarding registers in the second switching module are configured based on the global computing resource topology file. After the configuration is completed, the global computing resource topology file is returned to each controller, enabling them to understand the resource layout of the entire system. The beneficial effect of this step is to achieve global management and optimization of the entire computing resources, improving resource utilization and system performance.

[0070] As a feasible implementation, generating the local computing resource topology file includes: obtaining the asset information of the computing resources and populating the local bus identifier field in the local computing resource topology file based on the asset information; wherein, the asset information includes the resource information of computing units in the local host and the placeholder memory address information of computing units in other hosts, and the local bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the local host; accessing the first switching module in the local host to obtain the memory address mapping information and global domain information of the computing resources, and populating the global domain identifier field and global bus identifier field in the local computing resource topology file based on the memory address mapping information and global domain information; wherein, the global domain identifier field is used to describe the domain identifier of the computing units in the local host from the perspective of the second switching module, and the global bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the second switching module.

[0071] In a specific implementation, the controller obtains the asset information of computing resources, including the resource information of computing units in the local host and the placeholder memory address information of computing units in other hosts. As a feasible implementation, the basic input / output system in the host collects the asset information and sends the asset information to the controller through the H2B shared memory. The controller fills the local bus identification field in the local computing resource topology file based on this asset information, and this field is used to describe the bus identification of the computing units in the local host from the local perspective. Next, the controller accesses the first switching module in the local host to obtain the memory address mapping information and global domain information of the computing resources. As a feasible implementation, the controller accesses the first switching module in the host where it is located through the management and control transport protocol of the system management bus. The memory address mapping information is a detailed description of the location of the computing resources in the memory and how they are accessed, while the global domain information provides the location information of the computing resources in the entire system. The controller fills the global domain identification field and the global bus identification field in the local computing resource topology file based on this information, and these fields respectively describe the domain identification and bus identification of the computing units in the local host from the perspective of the second switching module. It can be seen that this implementation improves the identifiability and manageability of computing resources, enables the resource management processor to more effectively integrate and configure the global computing resource topology file, optimizes the path of cross-host communication, and improves the system communication efficiency and performance.

[0072] S102: Configure the routing forwarding parameters and address mapping parameters in the first switching module in the host based on the global computing resource topology file.

[0073] In this step, the controller configures the routing forwarding parameters and address mapping parameters of the first switching module in the host where it is located based on the global computing resource topology file. As a feasible implementation, configure the address trap register and port routing forwarding register in the first switching module in the host where the controller is located based on the global computing resource topology file. The address trap register and the port routing forwarding register. The address trap register is used to capture and process access requests for specific addresses, while the port routing forwarding register is responsible for correctly routing data packets to the corresponding computing units according to the destination address. Through this configuration, the controller ensures that data can be efficiently and accurately transmitted within the host and between hosts.

[0074] As a feasible implementation, configuring the address trap register and port routing forwarding register in the first switching module in the host based on the global computing resource topology file includes: accessing the first switching module in the host to configure the address trap register and port routing forwarding register in the first switching module in the host based on the physical slots corresponding to each computing unit and the global computing resource topology file.

[0075] In a specific implementation, the controller configures the address trap register and the port routing and forwarding register in the first switching module by accessing the first switching module in the host where it is located and using the information in the global computing resource topology file and the physical slot corresponding to the computing unit.

[0076] S103: Obtain memory address mapping information, and access the computing units of itself or other hosts according to the address information in the memory address mapping information.

[0077] In this step, the controller sends the memory address mapping information to the processor in the host where it is located. This information includes the addresses of the computing units in the memory, enabling the processor to directly access each computing unit according to these address information. The memory address mapping information is the key for the processor to understand and access the system memory, and it allows the processor to efficiently manage and utilize the memory resources.

[0078] As a feasible implementation manner, obtaining the memory address mapping information includes: obtaining the memory address mapping information through an in-band manner.

[0079] In a specific implementation, the processor communicates with the baseboard management controller by executing in-band commands to obtain the memory address mapping information. In-band commands refer to commands issued through the operating system to obtain hardware information when the operating system is running, such as the in-band IPMI command of the KCS physical channel. The memory address mapping information includes the address identifiers of the computing units in the system, enabling the operating system client running on the processor to access each computing unit in each host.

[0080] In this embodiment, the topology routing configuration management process is as Figure 9 shown. First, the host basic input / output system starts and pushes asset information to the host baseboard management controller. The management controller parses this asset information, generates a local computing resource topology file, and accesses the first switching module of the host to collect the global information of the local computing units. Then, the management controller supplements the global information of the local computing units to the local computing resource topology file and obtains this file. Subsequently, the management controller integrates all the local computing resource topology files to generate a global computing resource topology file. This file is sent to the management software, and the management software configures the address trap register. After the configuration is successful, the port routing and forwarding register is configured. Finally, the host operating system client obtains the memory address mapping information in-band to complete the entire process.

[0081] In the embodiment of the present invention, through the collaborative work of the resource management processor and the controller, the generation and integration of the local computing resource topology file are realized, thereby constructing the global computing resource topology file, which provides a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the routing forwarding parameters and address mapping parameters in the first switching module and the routing forwarding parameters in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In summary, the embodiment of the present invention improves the efficiency and reliability of multi-host computing resource sharing.

[0082] An embodiment of the present invention provides a method for sharing computing resources. In combination with the execution process of the method for sharing computing resources, the method will be described in detail. Refer to Figure 10 , a flowchart of another method for sharing computing resources shown according to an exemplary embodiment.

[0083] S201: Receive the local computing resource topology files sent by each host in the computing resource sharing system, and integrate the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file.

[0084] S202: Configure the routing forwarding parameters in the second switching module of the global communication module based on the global computing resource topology file.

[0085] S203: Return the global computing resource topology file to each host so that each host can configure the routing forwarding parameters and address mapping parameters in the first switching module in each host based on the global computing resource topology file.

[0086] The execution subject of this embodiment is the resource management processor in the global communication module. The resource management processor is responsible for integrating the local computing resource topology files from multiple hosts and creating a global computing resource topology file. This global file enables the resource management processor to understand the resource distribution within the entire system and configure the routing forwarding parameters in the second switching module accordingly to optimize the data transmission path. As a feasible implementation, configure the port routing forwarding register in the second switching module based on the global computing resource topology file. After the configuration is completed, the global computing resource topology file is returned to each controller so that they can understand the resource layout of the entire system. The beneficial effect of this step is to realize the global management and optimization of the entire computing resources, improving the resource utilization rate and system performance.

[0087] In the embodiments of the present invention, through the collaborative work of the resource management processor and the controller, the generation and integration of the local computing resource topology file are realized, thereby constructing the global computing resource topology file, which provides a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the routing forwarding parameters and address mapping parameters in the first switching module and the routing forwarding parameters in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In summary, the embodiments of the present invention improve the efficiency and reliability of multi-host computing resource sharing.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0089] Next, a computing resource sharing device provided by the embodiments of the present invention will be introduced. The computing resource sharing device described below can be referred to with the computing resource sharing method on the controller side described above. Refer to Figure 11 , a structural diagram of a computing resource sharing device shown according to an exemplary embodiment.

[0090] The generation module 801 is configured to generate a local computing resource topology file and send the local computing resource topology file to the resource management processor in the global communication module, so that the resource management processor integrates the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file, configures the routing forwarding parameters in the second switching module in the global communication module based on the global computing resource topology file, and returns the global computing resource topology file to the host; wherein, the local computing resource topology file is used to describe the identification information of the computing units in the host.

[0091] The first configuration module 802 is configured to configure the routing forwarding parameters and address mapping parameters in the first switching module in the host based on the global computing resource topology file.

[0092] The acquisition module 803 is configured to acquire memory address mapping information and access the computing units of itself or other hosts according to the address information in the memory address mapping information.

[0093] In the embodiments of the present invention, through the collaborative work of the resource management processor and the controller, the generation and integration of the local computing resource topology file are realized, thereby constructing the global computing resource topology file, which provides a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the routing forwarding parameters and address mapping parameters in the first switching module and the routing forwarding parameters in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In summary, the embodiments of the present invention improve the efficiency and reliability of multi-host computing resource sharing.

[0094] Based on the above embodiments, as a preferred implementation manner, the local computing resource topology file is used to describe the bus identifier of the computing unit from the perspective of the local host, the domain identifier and bus identifier from the perspective of the second switching module; the generation module 801 is specifically configured to: obtain the asset information of the computing resources, and fill the local bus identifier field in the local computing resource topology file based on the asset information; wherein, the asset information includes the resource information of the computing units in the controller local host and the placeholder memory address information of the computing units in other hosts, and the local bus identifier field is used to describe the bus identifier of the computing units in the controller local host from the perspective of the local host; access the first switching module in the local host to obtain the memory address mapping information and global domain information of the computing resources, and fill the global domain identifier field and global bus identifier field in the local computing resource topology file based on the memory address mapping information and global domain information; wherein, the global domain identifier field is used to describe the domain identifier of the computing units in the local host from the perspective of the second switching module, and the global bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the second switching module.

[0095] Based on the above embodiments, as a preferred implementation manner, the first configuration module 802 is specifically configured to: configure the address trap register and port routing forwarding register in the first switching module in the host based on the global computing resource topology file.

[0096] Based on the above embodiments, as a preferred implementation manner, the first configuration module 802 is specifically configured to: access the first switching module in the host to configure the address trap register and port routing forwarding register in the first switching module in the host based on the physical slot corresponding to each computing unit and the global computing resource topology file.

[0097] Based on the above embodiments, as a preferred implementation manner, the acquisition module 803 is specifically configured to: obtain the memory address mapping information through the in-band method.

[0098] The following introduces another computing resource sharing device provided by the embodiments of the present invention. The another computing resource sharing device described below can be referred to in conjunction with the computing resource sharing method on the resource management processor side described above. Refer to Figure 12 , which is a structural diagram of another computing resource sharing device shown according to an exemplary embodiment.

[0099] The integration module 901 is configured to receive the local computing resource topology files sent by each host in the computing resource sharing system, and integrate the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file.

[0100] The second configuration module 902 is configured to configure the routing forwarding parameters in the second switching module in the global communication module based on the global computing resource topology file.

[0101] The sending module 903 is configured to return the global computing resource topology file to each host, so that each host configures the routing forwarding parameters and address mapping parameters in the first switching module in each host based on the global computing resource topology file.

[0102] In the embodiments of the present invention, through the collaborative work of the resource management processor and the controller, the generation and integration of the local computing resource topology file are realized, thereby constructing a global computing resource topology file, which provides a basis for efficient resource sharing and communication. Further, based on the global computing resource topology file, the routing forwarding parameters and address mapping parameters in the first switching module, and the routing forwarding parameters in the second switching module are configured, optimizing the data transmission path, ensuring the accuracy and reliability of data transmission, and improving the data transmission efficiency. The processor can flexibly access each computing unit in each host according to the memory address mapping information, realizing the dynamic sharing and efficient utilization of computing resources. In summary, the embodiments of the present invention improve the efficiency and reliability of multi-host computing resource sharing.

[0103] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0104] The embodiments of the present invention also provide an electronic device, Figure 13 which is a structural diagram of an electronic device shown according to an exemplary embodiment. As Figure 13 shown, the electronic device includes: a communication interface 1, capable of interacting with other devices such as network devices; a processor 2, connected to the communication interface 1 to implement information interaction with other devices, and when running a computer program, executing the computing resource sharing method provided by one or more of the above technical solutions. The computer program is stored on the memory 3.

[0105] Of course, in practical applications, each component in the electronic device is coupled together through the bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 13 all kinds of buses are labeled as the bus system 4.

[0106] The memory 3 in the embodiment of the present invention is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.

[0107] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.

[0108] The method disclosed in the above embodiments of the present invention can be applied to or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 2 or by instructions in the form of software. The above-mentioned processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 3. The processor 2 reads the program in the memory 3 and combines its hardware to complete the steps of the foregoing method.

[0109] When the processor 2 executes the program, it implements the corresponding processes in each method of the embodiments of the present invention. For the sake of brevity, it will not be elaborated here.

[0110] The embodiments of the present invention also provide a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any of the above embodiments of the computing resource sharing method when running.

[0111] In an exemplary embodiment, the above computer-readable storage medium may include but not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks and other various media that can store computer programs.

[0112] The embodiments of the present invention also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above embodiments of the computing resource sharing method.

[0113] The embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above embodiments of the computing resource sharing method.

[0114] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0115] The above has introduced in detail a computing resource sharing system, method, device, equipment, medium, and product provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A computing resource sharing system, characterized in that, It includes multiple hosts and a global communication module. Each host includes multiple computing units and a first switching module. The computing units within the same host are interconnected through the first switching module in the same host. The global communication module includes a second switching module and a resource management processor, and multiple hosts are interconnected through the second switching module; The host is used for: generating a local computing resource topology file and sending the local computing resource topology file to the resource management processor; wherein, the local computing resource topology file is used to describe the identification information of the computing units in the host; The resource management processor is used for: integrating the local computing resource topology files generated by multiple hosts to obtain a global computing resource topology file, configuring the routing and forwarding parameters in the second switching module based on the global computing resource topology file, and returning the global computing resource topology file to the host; The host is also used for: configuring the routing and forwarding parameters and address mapping parameters in the first switching module based on the global computing resource topology file; The host is also used for: accessing the computing units of its own or other hosts according to the address information in the memory address mapping information.

2. The computing resource sharing system according to claim 1, wherein The local computing resource topology file is used to describe the bus identifier of the computing unit from the perspective of the local host, the domain identifier and bus identifier from the perspective of the second switching module; The process by which the host generates the local computing resource topology file of the host where it is located includes: Obtaining the asset information of the computing resources and filling the local bus identifier field in the local computing resource topology file based on the asset information; wherein, the asset information includes the resource information of the computing units in the local host and the placeholder memory address information of the computing units in other hosts, and the local bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the local host; Accessing the first switching module in the local host to obtain the memory address mapping information and global domain information of the computing resources, and filling the global domain identifier field and global bus identifier field in the local computing resource topology file based on the memory address mapping information and the global domain information; wherein, the global domain identifier field is used to describe the domain identifier of the computing units in the local host from the perspective of the second switching module, and the global bus identifier field is used to describe the bus identifier of the computing units in the local host from the perspective of the second switching module.

3. The computing resource sharing system according to claim 2, wherein The global computing resource topology file includes the bus identifier of each computing unit from the perspective of the host where it is located, the domain identifier and bus identifier from the perspective of the second switching module; The global computing resource topology file further includes any one or any combination of the address mapping fields of each computing unit, the number of hosts included in the computing resource sharing system, the number of computing units, and the file length of the global computing resource topology file. The address mapping field is used to describe whether the address mapping of the computing unit is set up.

4. The computing resource sharing system according to claim 1, wherein The host is also used for: allocating memory addresses for the computing units in the local host and reserving placeholder memory addresses for the computing units in other hosts; Perform device enumeration and resource allocation for the computing units in the local host, so as to allocate a switching module domain address for the computing units in the local host, and map the memory addresses of the computing units to the switching module domain addresses.

5. The computing resource sharing system according to claim 1, wherein The host includes a processor and a baseboard management controller. The processor obtains memory address mapping information from the baseboard management controller through in-band commands, so as to access the computing units of its own host or other hosts according to the address information in the memory address mapping information.

6. The computing resource sharing system according to claim 1, wherein The process in which the resource management processor configures the routing forwarding parameters in the second switching module based on the global computing resource topology file includes: The resource management processor configures the port routing forwarding register in the second switching module based on the global computing resource topology file; The process in which the host configures the routing forwarding parameters and address mapping parameters in the first switching module based on the global computing resource topology file includes: The host configures the port routing forwarding register and the address trap register in the first switching module based on the global computing resource topology file.

7. The computing resource sharing system according to claim 6, wherein The process in which the host configures the port routing forwarding register and the address trap register in the first switching module based on the global computing resource topology file includes: The controller in the host accesses the first switching module in the host, so as to configure the port routing forwarding register and the address trap register in the first switching module based on the physical slots corresponding to the computing units and the global computing resource topology file.

8. The computing resource sharing system according to claim 7, wherein The computing resource sharing system includes N hosts. The process in which the i-th host accesses the target computing unit in the j-th host according to the memory address mapping information includes: The processor in the i-th host generates an access request based on the address information of the target computing unit in the memory address mapping information. The access request is captured by the address trap register in the first switching module of the i-th host. The address trap register converts the address information in the access request into the switching module domain address of the target computing unit. The port routing forwarding register in the first switching module of the i-th host routes the access request to the target computing unit in the j-th host through the second switching module in the global communication module based on the switching module domain address; where 1≤i≤N, 1≤j≤N, and i≠j.

9. The computing resource sharing system according to claim 1, wherein The host includes a first preset number of first switching modules, each of the first switching modules is connected to a second preset number of computing units. The global communication module includes a third preset number of second switching modules. The first switching modules are connected to the ports of the second switching modules, and the ports of the second switching modules connected by any two first switching modules are different.

10. A computing resource sharing method, characterized in that, Applied to the host in the computing resource sharing system, the method includes: Generate a local computing resource topology file and send the local computing resource topology file to a resource management processor in a global communication module, so that the resource management processor integrates the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file, configures routing forwarding parameters in a second switching module in the global communication module based on the global computing resource topology file, and returns the global computing resource topology file to the host; wherein, the local computing resource topology file is used to describe identification information of computing units in the host. Configure routing forwarding parameters and address mapping parameters in a first switching module in the host based on the global computing resource topology file. Obtain memory address mapping information and access computing units of itself or other hosts according to the address information of the memory address mapping information.

11. The computing resource sharing method according to claim 10, wherein, The local computing resource topology file is used to describe a bus identifier of a computing unit from the perspective of a local host, a domain identifier and a bus identifier from the perspective of the second switching module. The generating of the local computing resource topology file includes: Obtain asset information of computing resources and fill a local bus identifier field in the local computing resource topology file based on the asset information; wherein, the asset information includes resource information of computing units in a local host and placeholder memory address information of computing units in other hosts, and the local bus identifier field is used to describe a bus identifier of a computing unit in the local host from the perspective of the local host. Access a first switching module in the local host to obtain memory address mapping information and global domain information of computing resources, and fill a global domain identifier field and a global bus identifier field in the local computing resource topology file based on the memory address mapping information and the global domain information; wherein, the global domain identifier field is used to describe a domain identifier of a computing unit in the local host from the perspective of the second switching module, and the global bus identifier field is used to describe a bus identifier of a computing unit in the local host from the perspective of the second switching module.

12. The computing resource sharing method according to claim 10, wherein Configuring routing forwarding parameters and address mapping parameters in a first switching module in the host based on the global computing resource topology file includes: Configure an address trap register and a port routing forwarding register in a first switching module in the host based on the global computing resource topology file.

13. The computing resource sharing method according to claim 12, wherein Configuring an address trap register and a port routing forwarding register in a first switching module in the host based on the global computing resource topology file includes: Access the first switching module in the host to configure the address trap register and the port routing forwarding register in the first switching module in the host based on the physical slot corresponding to each computing unit and the global computing resource topology file.

14. The computing resource sharing method according to claim 10, wherein Obtaining memory address mapping information includes: Obtain memory address mapping information in a in-band manner.

15. A computing resource sharing method, characterized in that, Applied to a resource management processor in a global communication module in a computing resource sharing system, the method includes: Receive the local computing resource topology files sent by each host in the computing resource sharing system, and integrate the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file; Configure the routing forwarding parameters in the second switching module of the global communication module based on the global computing resource topology file; Return the global computing resource topology file to each host so that each host can configure the routing forwarding parameters and address mapping parameters in the first switching module of each host based on the global computing resource topology file.

16. A computing resource sharing device, characterized in that, Applied to the host in the computing resource sharing system, the device includes: A generation module, configured to generate a local computing resource topology file, and send the local computing resource topology file to the resource management processor in the global communication module, so that the resource management processor integrates the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file, configures the routing forwarding parameters in the second switching module of the global communication module based on the global computing resource topology file, and returns the global computing resource topology file to the host; wherein, the local computing resource topology file is used to describe the identification information of the computing units in the host; A first configuration module, configured to configure the routing forwarding parameters and address mapping parameters in the first switching module of the host based on the global computing resource topology file; An acquisition module, configured to acquire memory address mapping information, and access the computing units of itself or other hosts according to the address information of the memory address mapping information.

17. A computing resource sharing device, characterized in that, Applied to the resource management processor in the global communication module of the computing resource sharing system, the device includes: An integration module, configured to receive the local computing resource topology files sent by each host in the computing resource sharing system, and integrate the local computing resource topology files corresponding to multiple hosts to obtain a global computing resource topology file; A second configuration module, configured to configure the routing forwarding parameters in the second switching module of the global communication module based on the global computing resource topology file; A sending module, configured to return the global computing resource topology file to each host so that each host can configure the routing forwarding parameters and address mapping parameters in the first switching module of each host based on the global computing resource topology file.

18. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to implement the steps executed by the computing resource sharing method according to any one of claims 10 to 15 when executing the computer program.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the steps executed by the computing resource sharing method according to any one of claims 10 to 15.

20. A computer program product, characterized in that, Includes a computer program, and when the computer program is executed, it implements the steps executed by the computing resource sharing method according to any one of claims 10 to 15.

Citation Information

Patent Citations

  • SDN routing method and routing system based on virtual network mapping

    CN116938811A

  • External device topology configuration method, data processor, device and program product

    CN118069568A

  • Equipment resource management method, related system and storage medium

    CN118860624A

  • Optical communication network routing optimization method and system

    CN119946469A

  • Device resource management method, related system, and storage medium

    WO2024222777A1

Cited By

  • Computing system, computing management method, and electronic device

    CN121349379A

  • Interconnect system, method, device, medium, and program product

    WO2026145042A1