Computer system, resource sharing method, device and equipment, medium and product

By introducing a resource management processor in the computer system, integrating the configuration information of the switch and host, realizing in-band topological interconnection between multiple hosts, the problem of low efficiency in sharing computing resources by multi-hosts is solved and the communication efficiency and stability of the system is improved.

CN120371767AActive Publication Date: 2025-07-25INSPUR SUZHOU INTELLIGENT TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510866431.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the prior art, multi-host computing resource sharing efficiency is low and difficult to scale, especially the communication efficiency between different accelerator cards is low, and there is a lack of monitoring management and flexible scheduling of global resources.

Method used

By introducing a resource management processor in the computer system, integrating the configuration information of the switch and host, generating global configuration information, realizing in-band topological interconnection and resource management between multiple hosts, allowing the host to directly access the acceleration cards of other hosts to avoid upstream port forwarding.

Benefits of technology

It improves the efficiency and reliability of multi-host computing resource sharing, realizes more efficient, lower latency and flexible and scalable resource management, and improves the communication efficiency and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371767A_ABST
    Figure CN120371767A_ABST
Patent Text Reader

Abstract

The invention discloses a computer system, a resource sharing method, device and equipment, a medium and a product, and relates to the technical field of computers, the system comprises a plurality of hosts and a resource management processor, the hosts are connected with a plurality of acceleration cards through switches, the resource management processor is connected with a plurality of switches, and the switches are mutually connected; the switch generates first local configuration information, identification information describing a local acceleration card and connection port information; the host generates second local configuration information, identification information describing each accelerator card and local resource allocation information; the resource management processor generates global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host; and each host accesses the acceleration card of the host or other hosts based on the global configuration information, and the local acceleration card of each host accesses the acceleration cards of other hosts based on the global configuration information. According to the invention, the efficiency and reliability of multi-host computing resource sharing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a computer system, a resource sharing method, apparatus, device, medium, and product. Background Art

[0002] In the related art, when implementing multi-host sharing of computing resources through a switch, it usually relies on the NTB (Non-Transparent Bridge) scheme. This scheme has complex logic, low efficiency, and is difficult to scale to the case of multiple switches. In addition, in the above scheme, the communication between different acceleration cards needs to be forwarded through the upstream port, further reducing the communication efficiency.

[0003] Therefore, how to achieve efficient sharing of computing resources between multiple hosts is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a computer system, a resource sharing method, apparatus, device, medium, and product, which improves the efficiency and reliability of multi-host computing resource sharing.

[0005] This application provides a computer system, including multiple hosts and a resource management processor. The hosts are connected to multiple acceleration cards through switches, the resource management processor is connected to multiple switches, and the multiple switches are interconnected with each other;

[0006] The switch generates first local configuration information; wherein, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card;

[0007] The host generates second local configuration information; wherein, the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information;

[0008] The resource management processor generates global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch;

[0009] Each host accesses its own or other hosts' acceleration cards based on the global configuration information, or, the local acceleration cards of each host access the acceleration cards of other hosts based on the global configuration information.

[0010] This application provides a computing resource sharing method, which is applied to a switch in a computer system. The method includes:

[0011] Generating first local configuration information; wherein, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card;

[0012] Receive the global configuration information sent by the resource management processor; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host.

[0013] Write the global configuration information into the configuration space so that the host can access its own or other hosts' acceleration cards based on the global configuration information.

[0014] This application provides a computing resource sharing method, which is applied to a host in a global communication module of a computer system. The method includes:

[0015] Generate the second local configuration information; wherein, the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information.

[0016] Access its own or other hosts' acceleration cards based on the global configuration information; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch.

[0017] This application provides a computing resource sharing method, which is applied to a resource management processor in a global communication module of a computer system. The method includes:

[0018] Obtain the first local configuration information generated by each switch and the second local configuration information generated by each host; wherein, the first local configuration information is used to describe the identification information of the local acceleration card and the connection port information, and the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information.

[0019] Generate the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and return the global configuration information to each switch.

[0020] This application also provides a computing resource sharing device, which is applied to a switch in a computer system. The device includes:

[0021] The first generation module is used to generate the first local configuration information; wherein, the first local configuration information is used to describe the identification information of the local acceleration card and the connection port information.

[0022] The receiving module is used to receive the global configuration information sent by the resource management processor; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host.

[0023] The writing module is used to write the global configuration information into the configuration space so that the host can access its own or other hosts' acceleration cards based on the global configuration information.

[0024] The present application also provides a computing resource sharing device, which is applied to a host in a global communication module of a computer system. The device includes:

[0025] A second generation module, configured to generate second local configuration information; wherein the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information;

[0026] An access module, configured to access its own or other hosts' acceleration cards based on the global configuration information; wherein, the resource management processor generates global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch.

[0027] The present application also provides a computing resource sharing device, which is applied to a resource management processor in a global communication module of a computer system. The device includes:

[0028] An acquisition module, configured to acquire the first local configuration information generated by each switch and the second local configuration information generated by each host; wherein the first local configuration information is used to describe the identification information of the local acceleration card and the connection port information, and the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information;

[0029] A third generation module, configured to generate global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and return the global configuration information to each switch.

[0030] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above computing resource sharing methods when executing the computer program.

[0031] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above computing resource sharing methods are implemented.

[0032] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above computing resource sharing methods are implemented.

[0033] The computer system provided by this application realizes the topological interconnection of multiple hosts in a band - in - band manner, integrates the configuration information of each switch and host through a resource management processor, and realizes the efficient management and scheduling of global resources. Each host can directly access its own or other hosts' acceleration cards based on the global configuration information, avoiding the cumbersome process of communication between different acceleration cards that needs to pass through upstream port forwarding, thus significantly improving the communication efficiency and realizing the efficient sharing of computing resources among multiple hosts. This application also discloses a computing resource sharing method, device, an electronic device, a computer - readable storage medium, and a computer program product, which can also achieve the above - mentioned technical effects.

[0034] It should be understood that the above general description and the following detailed description are only exemplary and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0036] Figure 1 It is a structural diagram of a computer system shown according to an exemplary embodiment;

[0037] Figure 2 It is a schematic diagram of the relationship between hardware and software levels in a computer system shown according to an exemplary embodiment;

[0038] Figure 3 It is a structural diagram of a computer system with 4 hosts expanding 16 acceleration cards shown according to an exemplary embodiment;

[0039] Figure 4 It is a flowchart of an application embodiment provided by this application;

[0040] Figure 5 It is a flowchart of a computing resource sharing method shown according to an exemplary embodiment;

[0041] Figure 6 It is a flowchart of another computing resource sharing method shown according to an exemplary embodiment;

[0042] Figure 7 It is a flowchart of yet another computing resource sharing method shown according to an exemplary embodiment;

[0043] Figure 8 It is a structural diagram of a computing resource sharing device shown according to an exemplary embodiment;

[0044] Figure 9 Structural diagram of another computing resource sharing device shown according to an exemplary embodiment;

[0045] Figure 10 Structural diagram of yet another computing resource sharing device shown according to an exemplary embodiment;

[0046] Figure 11 Structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0048] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0049] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0050] In a PCIe (Peripheral Component Interconnect Express) interconnection system, a PCIe NTB technology is usually used to implement a multi-HOST (host) communication solution, which is specifically as follows:

[0051] 1. Two hosts are respectively connected to PCIe Switch chips, and the two PCIe Switch chips are connected by a Crosslink line. The connection ports of the PCIe Switch chips are set to the NTB mode, and two GPUs (Graphics Processing Unit) are interconnected and communicate through NTB.

[0052] 2. GPU heterogeneous communication software is respectively integrated in the two hosts.

[0053] 3. The hosts are interconnected through PCIe NT Crosslink (PCIe non-transparent cross), and the NTB mechanism is used to achieve cross-host memory access. Then, the DMA (Direct Memory Access) inside the PCIe Switch chip is used to move data.

[0054] 4. By opening a window in NTB, the remote node address space is mapped to the local address space through NTB, so as to open up the access path between the two GPUs.

[0055] 5. The two GPUs are the local GPU and the remote GPU, and the two hosts are the local host and the remote host. The remote GPU BAR (Base Address Register) space is windowed in the local NTB, and the local GPU initiates access to the remote GPU BAR space through DMA, thereby realizing P2P (Peer-to-Peer) transmission.

[0056] In this technical solution, PCIe NTB technology is actually implemented through Virtual Switch. The address domains at both ends are completely isolated. The address mapping depends on the complex logic of the NTB controller. Cross-domain address mapping and interrupt mechanisms will introduce additional delays and hardware overhead. NTB technology only supports 2-level switch interconnection, which is greatly limited when it involves the interaction of more than two host nodes. The scalability of the system topology of multiple host nodes is also poor. In the cross-host domain expansion GPU card system formed by PCIeNTB technology, the management solution cannot be embedded in the control logic, so there is a lack of monitoring and management of global resources, which also leads to poor flexibility in resource scheduling within the system, weak fault handling capabilities, and poor stability during operation.

[0057] Therefore, in this application, the resource management processor and the host use the configuration space of the northbound interface of the switch to transmit and maintain global configuration information to make full use of hardware resources. Without the need for an additional in-band management network, the monitoring, management and scheduling updates of interface device resources can be realized at the software level, reducing the cost of system construction and maintenance. At the same time, combined with the fully interconnected topology of the switch, it is possible to access the accelerator card across host domains, and then realize the redeployment of interface device resources at the software level, and the flexible scheduling of global computing resources, thereby improving resource utilization. In addition, it provides a more efficient, lower latency, flexible and scalable management and monitoring solution for multi-host computer systems, thereby improving the communication efficiency and stability of the entire system.

[0058] This embodiment provides a computer system, such as Figure 1As shown in the figure, it includes multiple hosts 10 and a resource management processor 20. The hosts 10 are connected to multiple acceleration cards 40 through a switch 30. The resource management processor 20 is connected to multiple switches 30, and the multiple switches 30 are interconnected with each other.

[0059] In a specific implementation, each host 10 is directly connected to a switch 30. The switch 30 is connected to multiple acceleration cards 40. The acceleration cards 40 connected to a host 10 through a switch 30 are called the local acceleration cards of the host 10. That is, a host 10, a switch 30, and multiple acceleration cards 40 form a computing node. All the switches 30 are interconnected with each other. The resource management processor 20 is connected to each switch 30. Each switch 30 can be connected to the resource management processor 20 through UART (Universal Asynchronous Receiver / Transmitter). The switch 30 can adopt a PCIe Switch, that is, the switch 30 is connected to the host 10 and the acceleration cards 40 through the PCIe bus. The acceleration card can be a GPU, and the resource management processor 20 can be an mCPU (Micro CPU, microprocessor / management unit).

[0060] The schematic diagram of the relationship between the hardware and software levels in the computer system is as Figure 2 shown. At the firmware / software level, the computing node pool includes the operating system (OS) client in the host. The switching pool includes the switch firmware and the resource management processor. The resource pool includes the acceleration card northbound interface configuration space. The operating system client is connected to the resource management processor through the switch firmware. The switch firmware is connected to the acceleration card northbound interface configuration space. The operating system client is connected to the acceleration card northbound interface configuration space. At the hardware level, the computing node pool includes multiple hosts. The switching pool includes multiple switches and the resource management processor. The resource pool includes multiple risers, and each riser is connected to multiple acceleration cards. The switch is connected to the corresponding host through the northbound interface. Multiple switches are all connected to the resource management processor, and multiple switches are interconnected with each other.

[0061] A computer system with 4 hosts expanding 16 acceleration cards is as Figure 3 shown. The figure contains four hosts, each host is respectively connected to the corresponding switch. Each switch is connected to 4 acceleration cards and a network interface card (NIC). Multiple switches are interconnected through ports. The resource management processor is connected to multiple switches respectively.

[0062] The switch 30 generates first local configuration information; wherein, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card.

[0063] The host 10 generates second local configuration information; wherein, the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information;

[0064] Based on the first local configuration information generated by each switch and the second local configuration information generated by each host, the resource management processor 20 generates global configuration information and returns the global configuration information to each switch;

[0065] Each host 10 accesses the acceleration card 40 of itself or other hosts 10 based on the global configuration information.

[0066] In this embodiment, after each switch enumerates the downstream acceleration cards, it generates first configuration information (# GPUSwitch Info). After the host is powered on, the operating system (OS) client accesses the downstream acceleration cards, generates second configuration information (# GPU HOST Info), and writes it into the configuration space of the northbound interface of the switch. The northbound interface is defined as the interface for receiving and processing requests from upstream devices. The resource management processor assigns a switch identifier (Switch ID) to each switch and a GPU identifier (GPU ID) to each acceleration card. After accessing all switches and obtaining the first configuration information and the second configuration information, it forms global configuration information covering global resources and updates the available resources in the current system. At the same time, the global configuration information can be updated to the configuration space of the northbound interface of each switch, and the host OS Client obtains the update by accessing the configuration space of the northbound interface of the switch.

[0067] As a feasible implementation manner, the resource management processor assigns a switch identifier to each switch and assigns a global bus identifier and a global access address to each acceleration card.

[0068] In a specific implementation, the resource management processor assigns a switch identifier (Switch ID) to each switch and assigns a global bus identifier (GPU Global BUS) and a global access address (GPU Global Address) to each acceleration card. This assignment method ensures that each component in the system can be uniquely identified, thus facilitating management and scheduling. In this way, the resource management processor can construct a global resource view and effectively manage and schedule the resources in the system.

[0069] As a feasible implementation manner, the first local configuration information includes the switch identifier, the accelerator card identifier, and the connection port identifier of the local accelerator card. The first local configuration information further includes any one or any combination of the presence status, the global bus identifier, and the global access address. The switch identifier of the local accelerator card is the identifier of the switch directly connected to the local accelerator card. The connection port identifier of the local accelerator card is the identifier of the switch port directly connected to the local accelerator card. The presence status of the local accelerator card is used to describe whether the local accelerator card is present. The global bus identifier of the local accelerator card is the bus identifier assigned by the resource management processor to the local accelerator card. The global access address of the local accelerator card is the access address assigned to the local accelerator card.

[0070] In specific implementation, each switch generates the first local configuration information for describing various statuses and identifiers of the local accelerator card. Specifically, it includes: the switch identifier, which is the identifier of the switch directly connected to the local accelerator card; the accelerator card identifier, which is the unique identifier of the local accelerator card; the connection port identifier, which is the identifier of the switch port directly connected to the local accelerator card; the presence status, which describes whether the local accelerator card is working properly; the global bus identifier, which is the bus identifier assigned by the resource management processor to the local accelerator card; and the global access address, which is the access address assigned by the resource management processor to the local accelerator card.

[0071] As a feasible implementation manner, the second local configuration information includes the switch identifier, the accelerator card identifier, and the local resource allocation information of each accelerator card. The local resource allocation information includes the local bus identifier and / or the local access address. The second local configuration information further includes the presence status of each accelerator card. The switch identifier of the accelerator card is the identifier of the switch directly connected to the accelerator card. The presence status of the accelerator card is used to describe whether the accelerator card is present. The local bus identifier is the bus identifier assigned by the host to the accelerator card. The local access address is the access address assigned by the host to the accelerator card.

[0072] In specific implementation, each host generates the second local configuration information for describing the local status and identifier of the accelerator card. Specifically, it includes: the switch identifier, which is the identifier of the switch directly connected to the accelerator card; the accelerator card identifier, which is the unique identifier of the accelerator card; the presence status, which describes whether the accelerator card is working properly; the local bus identifier, which is the bus identifier assigned by the host to the accelerator card; and the local access address, which is the access address assigned by the host to the accelerator card.

[0073] As a feasible implementation, after the switch generates the first local configuration information, it writes the first local configuration space into the configuration space; the process for the host to generate the second local configuration information includes: the host accessing the configuration spaces of the local acceleration card and the local switch to generate the second local configuration information, and writing the second local configuration information into the configuration space of the local switch; the resource management processor accesses the configuration spaces of each switch to obtain the first local configuration information generated by each switch and the second local configuration information generated by each host.

[0074] In a specific implementation, after the switch generates the first local configuration information, it writes it into the configuration space of the northbound interface. The process for the host to generate the second local configuration information includes accessing the northbound interface configuration spaces of the local acceleration card and the switch, collecting necessary information, and writing the generated second local configuration information into the northbound interface configuration space of the local switch. The resource management processor obtains the first local configuration information and the second local configuration information by accessing the northbound interface configuration spaces of each switch, and thus integrates and generates the global configuration information.

[0075] As a feasible implementation, the resource management processor monitors the presence status of each switch at a preset period to update the global configuration information, and returns the updated global configuration information to each switch.

[0076] In a specific implementation, the resource management processor monitors the presence status of each switch at a preset period to ensure the accuracy and timeliness of the global configuration information. When any change is detected, the resource management processor updates the global configuration information and returns the updated information to each switch. This dynamic update mechanism ensures that the system can respond promptly to resource changes, such as the addition or removal of acceleration cards, so as to maintain the efficient operation of the system.

[0077] As a feasible implementation, when the switch detects that the target local acceleration card goes offline, it notifies the resource management processor; the resource management processor updates the global configuration information according to the switch identifier and the acceleration card identifier of the target local acceleration card, and returns the updated global configuration information to each switch.

[0078] In a specific implementation, when the switch detects that a certain local acceleration card goes offline, it notifies the resource management processor. The resource management processor updates the global configuration information according to the received switch identifier and acceleration card identifier, and returns the updated information to each switch. This mechanism ensures that the system can quickly adapt when the acceleration card changes, maintaining the accuracy of resource allocation.

[0079] As a feasible implementation, multiple hosts in a computer system include a first host and a second host. The process for the first host to access a target acceleration card in the second host includes: the first host sends an access request to the corresponding first switch; the first switch sends the access request to the second switch corresponding to the second host according to the target switch identifier of the target acceleration card in the access request; the second switch determines the connection port identifier corresponding to the target acceleration card identifier of the target acceleration card in the access request, and sends the access request to the port corresponding to the connection port identifier to achieve access to the target acceleration card.

[0080] In a specific implementation, there are two hosts in the computer system, which are respectively called the first host and the second host. The first host sends an access request to the corresponding first switch. The first switch forwards the access request to the second switch corresponding to the second host according to the target switch identifier of the target acceleration card in the access request. The second switch determines the connection port identifier according to the target acceleration card identifier in the access request, and sends the access request to the corresponding port, thereby achieving access to the target acceleration card.

[0081] The computer system provided by the embodiments of the present application realizes the topological interconnection of multiple hosts in a band-in-band manner, and integrates the configuration information of each switch and host through a resource management processor to achieve efficient management and scheduling of global resources. Each host can directly access the acceleration cards of itself or other hosts based on the global configuration information, avoiding the cumbersome process of communication between different acceleration cards passing through the upstream port forwarding, thereby significantly improving the communication efficiency and realizing the efficient sharing of computing resources between multiple hosts.

[0082] The following introduces an application embodiment provided by the present application, as Figure 4 shown, including the following steps:

[0083] Step 1: Configure the port connecting the acceleration card (GPU) and the switch (Switch) as the PCIe EP (Endpoint) mode. During the startup process of the PCIe Switch Box, the Arm processor of each switch (Switch) obtains the switch identifier (Switch ID) from the resource management processor (mCPU), and enumerates the downstream acceleration cards (GPUs). According to whether the enumeration is successful, update the presence identifier Present for each acceleration card (GPU). Generate the first local configuration information (# GPU SwitchInfo) and write it into the switch (Switch) northbound interface configuration space.

[0084] The format of the first local configuration information (# GPU Switch Info) is as follows: # GPU Switch Info { GPU 00-00:{ “Switch ID” = “0x00”; “GPU ID” = “0x00”; “Switch Port” = “0x00”; “Present” = “0x01”; “GPU Global BUS” = “0x01”; “GPU Global Address” = “0x1_0000_0000”; }, … GPU 01-08 :{ “Switch ID” = “0x01”; “GPU ID” = “0x08”; “Switch Port” = “0x08”; “Present” = “0x01”; “GPU Global BUS” = “0x08”; “GPU Global Address” = “0x8_0000_0000”; }, … }

[0085] Among them, “GPU Global BUS” is the bus identifier enumerated by the resource management processor (mCPU) for the downstream acceleration card (GPU), and “GPU Global Address” is the access address assigned by the resource management processor (mCPU) for the downstream acceleration card (GPU).

[0086] Step 2: The resource management processor (mCPU) accesses the northbound interface configuration space of the switch (Switch) in sequence through communication methods such as UART to obtain the first local configuration information (# GPU Switch Info). At this time, each GPU corresponds to a unique management and scheduling address identifier {Switch_ID - Swich Port}.

[0087] Step 3: The host (HOST) needs to enumerate the local downstream acceleration cards (GPUs) and reserve resources for the acceleration cards (GPUs) under other switches (Switches) in the system for subsequent access. The specific process is as follows: The operating system client (OS Client) in the host (HOST) accesses the northbound interface configuration space of the acceleration card (GPU) and the northbound interface configuration space of the switch (Switch) to obtain the local {Switch_ID-GPU_ID} index information and the corresponding presence information, and maps the local acceleration card (GPU) to the predefined local access address (Local Address). At the same time, according to the predefined number of acceleration cards (GPUs) in the system, PCIe resources are reserved for other acceleration cards (GPUs) in the system and mapped to the local access address (LocalAddress). The format of the second local configuration information (# GPU HOST Info) is as follows: # GPU HOST Info { GPU 00-00 :{ “Switch ID”=“0x00”; “GPU ID”=“0x00”; “Present”=“0x01”; “GPU HOST BUS”=“0x50”; “GPU Local Address”=“0x1_0000_0000”; }, GPU 00-01 :{ “Switch ID”=“0x00”; “GPU ID”=“0x01”; “Present”=“0x01”; “GPU HOST BUS”=“0x51”; “GPU Local Address”=“0x2_0000_0000”; }, … GPU 01-00 :{ “Switch ID”=“0x01”; “GPU ID”=“0x00”; “Present”=“0x00”; “GPU HOST BUS”=“0x58”; "GPU Local Address" = "0x8_0000_0000"; }, GPU 01-01 : { "Switch ID" = "0x01"; "GPU ID" = "0x01"; "Present" = "0x00"; "GPU HOST BUS" = "0x59"; "GPU Local Address" = "0x9_0000_0000"; }, … }

[0088] Among them, "GPU HOST BUS" is the bus identifier enumerated by the host (HOST) for the accelerator card (GPU), and "GPU HOST Address" is the access address assigned by the host (HOST) for the accelerator card (GPU).

[0089] Step 4: The operating system client (OS Client) writes the second local configuration information (# GPU HOST Info) into the switch (Switch) northbound port configuration space through the PCIe link.

[0090] Step 5: The resource management processor (mCPU) accesses the switch (Switch) northbound port configuration space in sequence through communication methods such as UART to collect all the second local configuration information (# GPU HOST Info).

[0091] Step 6: At this time, the management and scheduling address identifier of each accelerator card (GPU) {Switch_ID - Swich Port}, that is, the physical port, and the address identifier {Switch_ID - GPU_ID} used by the host (HOST) to access the device have a one-to-one mapping relationship. Through the switch (Switch) interconnection network, each host (HOST) can access the accelerator cards (GPUs) attached to other hosts (HOSTs) across the host domain. For example Figure 3Taking the request from Host 0 to access the accelerator card 0 under Host 1 as an example, it will be first sent to the switch 0 node connected to Host 0. According to the address identifier {Switch_ID - GPU_ID}, it is forwarded in the switch (Switch) interconnection network to the target switch 1. Then, according to the correspondence between the internal accelerator card (GPU) of switch 1 and the {Switch_ID - Swich Port} field, it is forwarded to the physical port connected to the device accelerator card 0. The resource management processor (mCPU) integrates all the first local configuration information and the second local configuration information to form global configuration information (including the routing path and the in-place status of the accelerator card) and writes it back to the configuration space of the northbound interfaces of all switches (Switch).

[0092] Step 7: The operating system client (OS Client) obtains the global configuration information from the configuration space of the northbound port of the switch (Switch) to complete the configuration deployment. The user can view the current global configuration information in the operating system client (OS Client) and perform management scheduling by inputting scheduling information containing the address identifier of the accelerator card (GPU).

[0093] Step 8: After the initial configuration is completed, the resource management processor (mCPU) periodically accesses the switch (Switch), and updates the global configuration information according to the in-place status of the switch (Switch).

[0094] Step 9: The operating system client (OS Client) obtains the update of the global configuration information from the configuration space of the northbound port of the switch (Switch).

[0095] An example of the update of the global configuration information is as follows:

[0096] Step 10: Due to the characteristics of hot removal and hot addition of the switch (Switch), the uplink and downlink connected devices can still be unplugged and plugged in after the configuration is completed. That is, if a host (HOST) or an accelerator card (GPU) goes offline, it does not affect the normal operation of the switch (Switch), and the change is synchronized to the resource management processor (mCPU) in a timely manner through Uart communication. Similarly, if the resource management processor (mCPU) detects that the switch (Switch) goes offline, it can also access the configuration space in other in-place switches (Switch) and synchronize the change. Through this update feedback process, the impact on the normal operation of other nodes when a single node adjustment or failure occurs in the system can be reduced.

[0097] Step 11: The resource management processor (mCPU) can update the Present in-place status according to the index of the offline {Switch_ID - GPU_ID}, that is, it indicates that the switch (Switch) and the downstream accelerator card (GPU) are unavailable.

[0098] Step 12: The resource management processor (mCPU) synchronizes the updates to the configuration spaces of all the northbound interfaces of the switches (Switches).

[0099] Step 13: The operating system client (OS Client) monitors the global configuration information synchronized by the resource management processor (mCPU) within the northbound port configuration spaces of the switches (Switches), and detects the available situation of the current global resources.

[0100] Embodiments of the present application provide a computing resource sharing method. In combination with the execution flow of the computing resource sharing method, the method is described in detail.

[0101] See Figure 5 , a flowchart of a computing resource sharing method shown according to an exemplary embodiment, as Figure 5 shown, including:

[0102] S101: Generate first local configuration information; wherein, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card;

[0103] The execution subject of this embodiment is the switch in the computer system. Among them, the local acceleration card refers to the acceleration card directly connected to the current switch.

[0104] In this step, the switch generates the first local configuration information, which is used to describe the identification information and connection port information of the local acceleration card. The first local configuration information may include: switch identification, the identification of the switch to which the local acceleration card is directly connected; acceleration card identification, the unique identification of the local acceleration card; connection port identification, the identification of the switch port to which the local acceleration card is directly connected; in-place status, describing whether the local acceleration card is working properly; global bus identification, the bus identification assigned by the resource management processor for the local acceleration card; global access address, the access address assigned by the resource management processor for the local acceleration card.

[0105] S102: Receive the global configuration information sent by the resource management processor; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host;

[0106] In this step, the resource management processor sequentially accesses the upstream port configuration spaces of each switch through communication means such as UART to collect all the first local configuration information. At the same time, the resource management processor also collects the second local configuration information generated by each host. Then, the resource management processor integrates this information into the global configuration information and sends it to each switch through the northbound interface.

[0107] S103: Write the global configuration information into the configuration space so that the host can access the accelerator cards of itself or other hosts based on the global configuration information.

[0108] In this step, after the switch receives the global configuration information sent by the resource management processor, it writes the information into the configuration space of the northbound interface. The northbound interface refers to the interface that receives and processes requests from upstream devices and is used for communication between the host and the switch. The host can obtain the global configuration information by accessing the configuration space of the northbound interface, and thus can access the accelerator cards of the local or other hosts according to the configuration information.

[0109] It can be seen that by writing the global configuration information into the configuration space of the northbound interface, the host can flexibly access the accelerator cards of other hosts in the system based on the global configuration information, realizing cross-host domain resource access and improving the resource utilization rate and flexibility of the system.

[0110] Based on the above embodiments, as a preferred implementation, it further includes: when it is detected that the target accelerator card goes offline, notify the resource management processor so that the resource management processor updates the global configuration information according to the switch identifier and accelerator card identifier of the target accelerator card; receive the updated global configuration information sent by the resource management processor, and write the updated global configuration information into the configuration space.

[0111] In specific implementation, when the switch detects that the target accelerator card going downstream goes offline, the switch notifies the resource management processor, and the resource management processor updates the global configuration information according to the switch identifier and accelerator card identifier of the target accelerator card. For example, if a certain accelerator card goes offline, the resource management processor will mark the corresponding "Present" status as "0x00" (indicating unavailable). The updated global configuration information will be sent back to the switch and written into the configuration space of the northbound interface, so as to ensure that the host can access the accelerator cards in the system based on the latest global configuration information.

[0112] It can be seen that the above implementation can handle the offline situation of the accelerator card more intelligently and efficiently. This mechanism not only improves the stability and reliability of the system, but also ensures that other hosts in the system can obtain the latest resource status in a timely manner, thereby optimizing the resource management and scheduling efficiency of the entire system.

[0113] The embodiments of the present application provide a computing resource sharing method. Combining with the execution process of the computing resource sharing method, the method is described in detail.

[0114] See Figure 6 , according to the flowchart of another computing resource sharing method shown in an exemplary embodiment, as Figure 6 shown, it includes:

[0115] S201: Generate the second local configuration information, where the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information.

[0116] The execution entity of this embodiment is the host in the computer system. In this step, the host generates the second local configuration information, which is used to describe the local status and identification of the acceleration card. Specifically, it includes: switch identification, the identification of the switch directly connected to the acceleration card; acceleration card identification, the unique identification of the acceleration card; in-place status, which describes whether the acceleration card is working properly; local bus identification, the bus identification allocated by the host for the acceleration card; local access address, the access address allocated by the host for the acceleration card.

[0117] S202: Access the acceleration cards of itself or other hosts based on the global configuration information; where the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch.

[0118] In this step, the host obtains the global configuration information from the configuration space of the northbound interface. The global configuration information contains the detailed information of all acceleration cards in the system, including their identification information, connection port information, and global resource allocation information. The host accesses the target acceleration card through the interconnection network of the switch according to the address identification and routing path in the global configuration information. For example, if the host needs to access the acceleration card under another host, it will find the correct routing path according to the global configuration information and send an access request through the link. The switch forwards the access request correctly to the target acceleration card according to the routing information in the global configuration information.

[0119] It can be seen that accessing the acceleration card based on the global configuration information enables the host to flexibly access any acceleration card in the system, regardless of whether the acceleration card is directly connected to the host. This greatly improves the resource utilization rate and flexibility of the system, allowing the system to dynamically reallocate resources during operation to meet different computing requirements. At the same time, this mechanism also improves the scalability of the system because new acceleration cards can be more easily added to the system and accessed by other hosts.

[0120] The embodiments of this application provide a computing resource sharing method. Combining with the execution process of the computing resource sharing method, the method is described in detail.

[0121] See Figure 7 , according to the flowchart of another computing resource sharing method shown in an exemplary embodiment, as Figure 7 shown, includes:

[0122] S301: Obtain the first local configuration information generated by each switch and the second local configuration information generated by each host. Among them, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card, and the second local configuration information is used to describe the identification information of each acceleration card and the local resource allocation information.

[0123] The execution subject of this embodiment is the resource management processor in the computer system. In this step, the resource management processor communicates with each switch and host through a predefined communication link. It accesses each switch and host in turn to obtain the local configuration information they generate. For the switch, the resource management processor obtains the first local configuration information, which includes the identification information, connection port information, and global resource allocation information of each acceleration card. For the host, the resource management processor obtains the second local configuration information, which includes the identification information of each acceleration card and the local resource allocation information.

[0124] S302: Generate global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and return the global configuration information to each switch.

[0125] In this step, the resource management processor integrates the obtained first local configuration information and second local configuration information. It summarizes the identification information, connection port information, and global resource allocation information of each acceleration card. Ensure that the mapping relationship between the global address and the local address of each acceleration card is correct. Generate global configuration information covering all acceleration cards in the system. The resource management processor sends the generated global configuration information to each switch through the northbound interface. The switch receives and stores this information so that the host can access the latest global configuration information through the northbound interface.

[0126] It can be seen that by integrating all local configuration information, the global configuration information provides a complete view of all acceleration cards in the system, enabling the host to access any acceleration card in the system based on the latest global configuration information.

[0127] Based on the above embodiments, as a preferred implementation manner, it further includes: monitoring the presence status of each switch at a preset period to update the global configuration information, and returning the updated global configuration information to each switch.

[0128] In a specific implementation, the resource management processor sets a timer and sends status query requests to each switch via a communication link at a preset period. After receiving the request, the switch returns its current in-service status information, including whether it is working properly, the status of the connected acceleration card, etc. The resource management processor updates the relevant records in the global configuration information according to the in-service status information returned by the switch. For example, if a certain switch or its connected acceleration card goes offline, the resource management processor will mark the corresponding "Present" status as "0x00" (indicating unavailable) and update the global configuration information. The resource management processor sends the updated global configuration information to each switch via the northbound interface. The switch receives and stores this information so that the host can access the latest global configuration information via the northbound interface.

[0129] Through the above implementation, the system can monitor the status of the switch more intelligently and efficiently and update the global configuration information in a timely manner. This mechanism not only improves the stability and reliability of the system but also ensures that all hosts in the system can obtain the latest resource status information, thus optimizing the resource management and scheduling efficiency of the entire system.

[0130] Based on the above embodiments, as a preferred implementation, it further includes: when receiving the offline notice of the target acceleration card, updating the global configuration information according to the switch identifier and acceleration card identifier of the target acceleration card and returning the updated global configuration information to each switch.

[0131] In a specific implementation, when the switch detects that the target acceleration card goes offline, it sends an offline notice to the resource management processor through a predefined communication mechanism. The notice content includes the switch identifier and acceleration card identifier of the target acceleration card. After receiving the offline notice, the resource management processor finds the corresponding record of the target acceleration card in the global configuration information according to the switch identifier and acceleration card identifier in the notice. Update the status of this record, for example, mark the "Present" status as "0x00" (indicating unavailable), and update the global configuration information. The resource management processor sends the updated global configuration information to each switch via the northbound interface. The switch receives and stores this information so that the host can access the latest global configuration information via the northbound interface.

[0132] Through the above implementation, the system can handle the offline situation of the acceleration card more intelligently and efficiently. This mechanism not only improves the stability and reliability of the system but also ensures that all hosts in the system can obtain the latest resource status information, thus optimizing the resource management and scheduling efficiency of the entire system.

[0133] The following introduces a computing resource sharing device provided by an embodiment of the present application. The following-described computing resource sharing device and the computing resource sharing method on the controller side described above can be referred to each other.

[0134] See Figure 8 , a structural diagram of a computing resource sharing device shown according to an exemplary embodiment, as Figure 8 shown, includes:

[0135] A first generation module 101, configured to generate first local configuration information; wherein, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card.

[0136] A receiving module 102, configured to receive global configuration information sent by a resource management processor; wherein, the resource management processor generates global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host.

[0137] A writing module 103, configured to write the global configuration information into the configuration space, so that a host can access its own or other hosts' acceleration cards based on the global configuration information.

[0138] Based on the above embodiment, as a preferred implementation manner, it further includes:

[0139] A detection module, configured to notify the resource management processor when it detects that a target acceleration card goes offline, so that the resource management processor updates the global configuration information according to the switch identification and acceleration card identification of the target acceleration card.

[0140] A first update module, configured to receive the updated global configuration information sent by the resource management processor and write the updated global configuration information into the configuration space.

[0141] The following introduces another computing resource sharing device provided by an embodiment of the present application. The following-described another computing resource sharing device and the computing resource sharing method on the resource management processor side described above can be referred to each other.

[0142] See Figure 9 , a structural diagram of another computing resource sharing device shown according to an exemplary embodiment, as Figure 9 shown, includes:

[0143] A second generation module 201, configured to generate second local configuration information; wherein, the second local configuration information is used to describe the identification information and local resource allocation information of each acceleration card.

[0144] An access module 202 for accessing its own or other hosts' acceleration cards based on global configuration information; wherein, the resource management processor generates global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch.

[0145] Another computing resource sharing device provided by an embodiment of the present application is introduced below. The other computing resource sharing device described below can be referred to in mutual reference with the computing resource sharing method on the resource management processor side described above.

[0146] See Figure 10 , a structural diagram of another computing resource sharing device shown according to an exemplary embodiment, as Figure 10 shown, includes:

[0147] An acquisition module 301 for acquiring the first local configuration information generated by each switch and the second local configuration information generated by each host; wherein, the first local configuration information is used to describe the identification information and connection port information of the local acceleration card, and the second local configuration information is used to describe the identification information and local resource allocation information of each acceleration card.

[0148] A third generation module 302 for integrating the first local configuration information generated by each switch and the second local configuration information generated by each host, generating global configuration information, and returning the global configuration information to each switch.

[0149] Based on the above embodiments, as a preferred implementation manner, it further includes:

[0150] A monitoring module for monitoring the on-site status of each switch at a preset period to update the global configuration information and return the updated global configuration information to each switch.

[0151] Based on the above embodiments, as a preferred implementation manner, it further includes:

[0152] A second update module for updating the global configuration information according to the switch identification and acceleration card identification of the target acceleration card when receiving the offline notification of the target acceleration card, and returning the updated global configuration information to each switch.

[0153] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated in detail here.

[0154] An embodiment of the present application further provides an electronic device, Figure 11 as a structural diagram of an electronic device shown according to an exemplary embodiment, as Figure 11 shown, the electronic device includes:

[0155] A communication interface 1, capable of interacting with other devices such as network devices for information exchange;

[0156] A processor 2, connected to the communication interface 1 to achieve information interaction with other devices, and when used to run a computer program, execute the computing resource sharing method provided by one or more of the above technical solutions. And the computer program is stored on a memory 3.

[0157] Of course, in actual application, each component in the electronic device is coupled together through a bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 11 all kinds of buses are labeled as the bus system 4.

[0158] The memory 3 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.

[0159] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable types of memories.

[0160] The method disclosed in the embodiments of the present application above can be applied to the processor 2 or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 2. The above-mentioned processor 2 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, which is located in the memory 3. The processor 2 reads the program in the memory 3 and combines its hardware to complete the steps of the foregoing method.

[0161] When the processor 2 executes the program, it implements the corresponding processes in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.

[0162] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any of the above-described embodiments of the computing resource sharing method when running.

[0163] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0164] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above-described embodiments of the computing resource sharing method.

[0165] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above-described embodiments of the computing resource sharing method.

[0166] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0167] The above has introduced in detail a computer system, a computing resource sharing method, a device, a device, a medium, and a product provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A computer system, characterized in that, It includes multiple hosts and a resource management processor. The hosts are connected to multiple acceleration cards through switches, the resource management processor is connected to multiple switches, and the multiple switches are interconnected with each other; The switch generates first local configuration information; wherein, the first local configuration information is used to describe the identification information and connection port information of local acceleration cards; The host generates second local configuration information; wherein, the second local configuration information is used to describe the identification information of each acceleration card and local resource allocation information; The resource management processor generates global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch; Each host accesses the acceleration cards of other hosts based on the global configuration information, or the local acceleration cards of each host access the acceleration cards of other hosts based on the global configuration information.

2. The computer system according to claim 1, wherein The first local configuration information includes the switch identification, acceleration card identification, and connection port identification of the local acceleration card, and the first local configuration information further includes any one or any combination of the in-position state, global bus identification, and global access address; The switch identification of the local acceleration card is the identification of the switch directly connected to the local acceleration card, the connection port identification of the local acceleration card is the identification of the switch port directly connected to the local acceleration card, the in-position state of the local acceleration card is used to describe whether the local acceleration card is in position, the global bus identification of the local acceleration card is the bus identification allocated by the resource management processor for the local acceleration card, and the global access address of the local acceleration card is the access address allocated for the local acceleration card.

3. The computer system according to claim 2, characterized in that, The resource management processor allocates switch identifications for each switch, and allocates global bus identifications and global access addresses for each acceleration card.

4. The computer system according to claim 1, wherein The second local configuration information includes the switch identification, acceleration card identification, and local resource allocation information of each acceleration card. The local resource allocation information includes a local bus identification and / or a local access address, and the second local configuration information further includes the in-position state of each acceleration card; The switch identification of the acceleration card is the identification of the switch directly connected to the acceleration card, the in-position state of the acceleration card is used to describe whether the acceleration card is in position, the local bus identification is the bus identification allocated by the host for the acceleration card, and the local access address is the access address allocated by the host for the acceleration card.

5. The computer system according to claim 4, characterized in that, After the switch generates the first local configuration information, it writes the first local configuration information into the configuration space; The process of the host generating the second local configuration information includes: the host accesses the configuration space of the local acceleration card and the configuration space of the local switch to generate the second local configuration information, and writes the second local configuration information into the configuration space of the local switch; The resource management processor accesses the configuration spaces of each switch to obtain the first local configuration information generated by each switch and the second local configuration information generated by each host.

6. The computer system according to claim 1, wherein The resource management processor monitors the presence status of each switch according to a preset period to update the global configuration information, and returns the updated global configuration information to each switch.

7. The computer system according to claim 1, wherein When the switch detects that the target local acceleration card goes offline, it notifies the resource management processor; The resource management processor updates the global configuration information according to the switch identifier and acceleration card identifier of the target local acceleration card, and returns the updated global configuration information to each switch.

8. The computer system according to claim 1, wherein, The multiple hosts in the computer system include a first host and a second host. The process of the first host accessing the target acceleration card in the second host includes: The first host sends an access request to the corresponding first switch; The first switch sends the access request to the second switch corresponding to the second host according to the target switch identifier of the target acceleration card in the access request; The second switch determines the connection port identifier corresponding to the target acceleration card identifier of the target acceleration card in the access request, and sends the access request to the port corresponding to the connection port identifier to implement access to the target acceleration card.

9. A computing resource sharing method, characterized in that, Applied to the switch in the computer system according to any one of claims 1 to 8, the method includes: Generating first local configuration information; wherein, the first local configuration information is used to describe the identifier information and connection port information of the local acceleration card; Receiving the global configuration information sent by the resource management processor; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host; Writing the global configuration information into the configuration space so that the host can access the acceleration card of itself or other hosts based on the global configuration information.

10. The computing resource sharing method according to claim 9, wherein Further includes: When it detects that the target acceleration card goes offline, it notifies the resource management processor so that the resource management processor updates the global configuration information according to the switch identifier and acceleration card identifier of the target acceleration card; Receiving the updated global configuration information sent by the resource management processor, and writing the updated global configuration information into the configuration space.

11. A method for sharing computing resources, characterized in that, Applied to the host in the computer system according to any one of claims 1 to 8, the method includes: Generating second local configuration information; wherein, the second local configuration information is used to describe the identifier information of each acceleration card and the local resource allocation information; Accessing the acceleration card of itself or other hosts based on the global configuration information; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch.

12. A computing resource sharing method, characterized in that, Applied to the resource management processor in the computer system according to any one of claims 1 to 8, the method includes: Obtaining the first local configuration information generated by each switch and the second local configuration information generated by each host; wherein, the first local configuration information is used to describe the identifier information and connection port information of the local acceleration card, and the second local configuration information is used to describe the identifier information of each acceleration card and the local resource allocation information; Generate global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and return the global configuration information to each switch.

13. The computing resource sharing method according to claim 12, wherein Further included: Monitor the online status of each switch at a preset period to update the global configuration information, and return the updated global configuration information to each switch.

14. The computing resource sharing method according to claim 12, wherein Further included: When receiving a notification of the offline of a target acceleration card, update the global configuration information according to the switch identifier and acceleration card identifier of the target acceleration card, and return the updated global configuration information to each switch.

15. A computing resource sharing device, characterized in that, Applied to a switch in the computer system according to any one of claims 1 to 8, the device includes: A first generation module, configured to generate first local configuration information; wherein, the first local configuration information is used to describe the identifier information and connection port information of local acceleration cards. A receiving module, configured to receive the global configuration information sent by the resource management processor; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host. A writing module, configured to write the global configuration information into the configuration space, so that the host can access its own or other hosts' acceleration cards based on the global configuration information.

16. A computing resource sharing device, characterized in that, Applied to a host in the computer system according to any one of claims 1 to 8, the device includes: A second generation module, configured to generate second local configuration information; wherein, the second local configuration information is used to describe the identifier information of each acceleration card and local resource allocation information. An access module, configured to access its own or other hosts' acceleration cards based on the global configuration information; wherein, the resource management processor generates the global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and returns the global configuration information to each switch.

17. A computing resource sharing device, characterized in that, Applied to a resource management processor in the computer system according to any one of claims 1 to 8, the device includes: An acquisition module, configured to acquire the first local configuration information generated by each switch and the second local configuration information generated by each host; wherein, the first local configuration information is used to describe the identifier information and connection port information of local acceleration cards, and the second local configuration information is used to describe the identifier information of each acceleration card and local resource allocation information. A third generation module, configured to generate global configuration information based on the first local configuration information generated by each switch and the second local configuration information generated by each host, and return the global configuration information to each switch.

18. An electronic device, characterized in that, Included: A memory, configured to store a computer program. A processor, configured to implement the steps performed by the computing resource sharing method according to any one of claims 9 to 14 when executing the computer program.

19. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed, it implements the steps performed by the computing resource sharing method according to any one of claims 9 to 14.

20. A computer program product, characterized in that, Including a computer program, which when executed implements the steps performed by the computing resource sharing method according to any one of claims 9 to 14.

Citation Information

Patent Citations

  • PCIe-based vehicle resource sharing method and device, equipment, medium and vehicle

    CN117851298A

  • GPU cross-host communication interconnection system based on PCIe NTB

    CN119149480A

  • Server and equipment monitoring system and method thereof

    CN119906687A

  • Data processing system, method, device, medium and program product

    CN120144326A

  • Methods and apparatus for high-speed data bus connection and fabric management

    US20200081858A1