Server system, and job task processing method based on cloud service

By interconnecting CPUs, memory, network cards, and computing chips within a physical server, and constructing a second network using direct output channels and switching chips, the problem of limited communication in physical server clusters in existing technologies is solved. This enables an increase in the number of computing chips and an expansion of the cluster size, thereby improving computing performance.

WO2026066375A1PCT designated stage Publication Date: 2026-04-02HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing cloud vendor clusters, the network connection methods between physical servers limit the number of processors that can be expanded, which cannot meet the needs of large-scale clusters and results in the inability to efficiently complete tenant jobs.

Method used

By interconnecting CPUs, memory, network cards, and computing chips within a physical server, and building a second network using direct output channels and switching chips, computing chips can communicate directly with each other, reducing reliance on CPUs and network cards, thereby increasing the number of computing chips and the size of the cluster.

Benefits of technology

It enables efficient communication between computing chips, expands the cluster size, improves computing performance, and provides tenants with better cloud services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025104848_02042026_PF_FP_ABST
    Figure CN2025104848_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a server system, and a job task processing method based on a cloud service, which can provide a tenant with a higher-quality cloud service. The server system in the present application may include a plurality of physical server groups, wherein each physical server in these physical server groups includes a CPU, a memory, a network interface card, a disk and a computing chip. For any physical server, the CPU, memory, network interface card, disk and computing chip that are included in the physical server are interconnected inside the physical server, and network interface cards of all of the plurality of physical servers are all connected to a first network, wherein the first network enables the network interface cards of all of the plurality of physical servers to be interconnected within a group and between groups, and computing chips of all of the plurality of physical servers can all be connected to a second network by means of direct-output channels provided by the computing chips, wherein the second network enables the computing chips of all of the plurality of physical servers to be interconnected within a group and between groups.
Need to check novelty before this filing date? Find Prior Art

Description

Server system and cloud service-based job task processing method

[0001] The present application claims priority to the Chinese patent application No. 202411345904.8, filed on September 25, 2024, and entitled "Server system and cloud service-based job task processing method", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of cloud technology, and in particular to a server system and a cloud service-based job task processing method. BACKGROUND

[0003] With the rapid development of cloud technology, more and more tenants choose to use clusters provided by cloud vendors to complete their own job tasks. These clusters often contain multiple extension processors. Since each extension processor has a certain specification and performance, these clusters often have large enough specifications and strong enough performance to efficiently complete the tenants' job tasks, thereby meeting the tenants' job requirements.

[0004] In the related art, the cluster provided by the cloud vendor for the tenant can include multiple physical servers, and each physical server can include a certain number of extension processors, which can be used to process the tenants' job tasks. Among the multiple physical servers of the cluster, the physical distance between some physical servers can be far. In these physical servers, when the extension processor of a certain physical server needs to send a message of a job task to the extension processor of another physical server, it can successfully send the message to the extension processor of the other physical server through the network between them.

[0005] In the above process, the extension processor of the physical server needs to pass through the central processing unit (CPU) and the network card of the physical server to access the network. Such a network connection method often limits the size of the cluster, i.e., the number of extension processors included in the cluster is limited. In some large demand scenarios, it is often impossible to meet the tenants' demand for building large-scale clusters. SUMMARY

[0006] Embodiments of the present application provide a server system and a cloud service-based job task processing method, which can provide tenants with better cloud services.

[0007] The first aspect of the embodiments of the present application provides a server system, which is infrastructure for providing cloud services. The infrastructure can include a plurality of physical server groups, each of which includes a plurality of physical servers, each of which includes a CPU, a memory, a network card, a disk, and a computing chip.

[0008] For any one of the plurality of physical server groups, the CPU, the memory, the network card, the disk, and the computing chip included in the physical server are interconnected within the physical server, and the network cards of all the physical servers in the plurality of physical server groups are connected to a first network, which can interconnect the network cards of all the physical servers in the plurality of physical server groups within and between groups, and the computing chips of all the physical servers in the plurality of physical servers can be connected to a second network through the direct channels they have, and the second network can interconnect the computing chips of all the physical servers in the plurality of physical servers within and between groups.

[0009] As can be seen from the above system, when the computing chip of a certain physical server needs to communicate with the computing chip of another physical server, the computing chip of the certain physical server can directly communicate with the computing chip of the another physical server through the second network, without the need to access the first network through the CPU and the network card of the certain physical server to communicate with the computing chip of the another physical server. As can be seen, in the physical server cluster, the communication between the computing chips is no longer limited by the CPU and the network card, and the cloud provider can increase the number of computing chips in the physical server to increase the number of computing chips in the physical server cluster, which is equivalent to expanding the size of the cluster, thereby making the cluster have stronger computing performance and providing better cloud services for tenants.

[0010] In one possible implementation, a plurality of physical server groups are used to create a logical node for a tenant, the logical node comprising at least two physical servers in the plurality of physical server groups, a first sub-network of a first network, and a second sub-network of a second network, network cards of the at least two physical servers being configured to receive data of a job task of the tenant via the first sub-network, memories of the at least two physical servers and disks of the at least two physical servers being configured to store the data, and CPUs of the at least two physical servers being configured to instruct computing chips of the at least two physical servers to jointly process the data via the second sub-network to complete the job task. In the foregoing implementation, the infrastructure can be managed by a cloud management platform, and when the tenant has a job task to be processed, the tenant can input a job task processing request set by the tenant to a task processing interface provided by the cloud management platform, so that the cloud management platform can receive the job task processing request sent by the tenant via the task processing interface. Since the job task processing request comprises data of the job task and performance requirements of the job task, the cloud management platform can select a plurality of idle physical servers that meet the performance requirements of the job task from the plurality of physical server groups, and create a logical node for the tenant on the plurality of physical servers. The logical node comprises the plurality of physical servers, a network (a first sub-network of a first network) between network cards of the plurality of physical servers, and a network (a second sub-network of a second network) between computing chips of the plurality of physical servers. After the logical node is created for the tenant, the cloud management platform can install an operating system image specified by the tenant in the logical node, and since the logical node installed with the operating system image comprises the plurality of physical servers selected by the cloud management platform, the CPUs of the plurality of physical servers are equivalent to being installed with the operating system, so that the cloud management platform can send the job task and the data of the job task to the network cards of the plurality of physical servers via the first sub-network, so that the CPUs of the plurality of physical servers store the data in the memories and the disks of the plurality of physical servers after receiving the job task and the data via the network cards. Then, the CPUs of the plurality of physical servers can call the computing chips of the plurality of physical servers to obtain the data from the memories and the disks of the plurality of physical servers, and jointly process the data via the second sub-network to complete the job task. Thus, the cloud management platform can provide a certain scale of computing chips and a network (a sub-network of a second network) between the computing chips for the tenant according to performance requirements of a job task of the tenant, so as to meet customized requirements of the tenant for computing services and network services.

[0011] In one possible implementation, the second network includes a first switch chip corresponding to each physical server group and a second switch chip corresponding to the multiple physical server groups; the compute chips in each physical server group access the first switch chip through direct-out channels, and the first switch chip is configured to support the compute chips in each physical server group to be interconnected within the group; the compute chips in the multiple physical server groups access the second switch chip through direct-out channels, and the second switch chip is configured to support the compute chips in the multiple physical server groups to be interconnected between groups. In the foregoing implementation, for any one of the multiple physical server groups, the compute chips of any two physical servers in the physical server group can access the first switch chip through the direct-out channels possessed by the compute chips to achieve communication connection. Similarly, the compute chips of the physical servers in any two of the multiple physical server groups can also access the second switch chip through the direct-out channels possessed by the compute chips to achieve communication connection. As can be seen, the second network between the compute chips can be implemented through the first switch chip and the second switch chip.

[0012] In one possible implementation, the second network includes a second switch chip corresponding to the multiple physical server groups; the compute chips in each physical server group are interconnected within the group through direct-out channels; the compute chips in the multiple physical server groups access the second switch chip through direct-out channels, and the second switch chip is configured to support the compute chips in the multiple physical server groups to be interconnected between groups. In the foregoing implementation, for any one of the multiple physical server groups, the compute chips of any two physical servers in the physical server group can achieve communication connection through the direct-out channels possessed by the compute chips, and similarly, the compute chips of the physical servers in any two of the multiple physical server groups can also access the second switch chip through the direct-out channels possessed by the compute chips to achieve communication connection. As can be seen, the second network between the compute chips can be implemented through the direct-out channels and the second switch chip.

[0013] In a possible implementation, when the computing chips of the at least two physical servers transmit the data-associated message through the second sub-network, the sub-network is configured to complete the transmission of the message after determining that the message carries the identifier of the second sub-network. In the foregoing implementation, the cloud management platform can configure a globally unique identifier for the second sub-network included in the logical node of the tenant. Then, when the computing chips of the plurality of physical servers in the logical node process the job task, if the computing chips of the plurality of physical servers need to transmit a data-associated message through the second sub-network, the second sub-network can determine whether the message carries the unique identifier of the second sub-network. If the message carries the unique identifier of the second sub-network, the second sub-network forwards the message. If the message does not carry the unique identifier of the second sub-network, the second sub-network discards the message. In this way, the cloud management platform assigns a unique identifier to the network (that is, the second sub-network) between the computing chips in the logical node of each tenant, so that the network between the computing chips in the logical nodes of different tenants is isolated and does not affect the processing of the respective job tasks.

[0014] In a possible implementation, the direct-out channel includes a first direct-out channel and a second direct-out channel. When the first direct-out channel is in a working state, the computing chips of the at least two physical servers access the second sub-network through the first direct-out channel. When the first direct-out channel is in a fault state, the computing chips of the at least two physical servers access the second sub-network through the second direct-out channel. In the foregoing implementation, for the computing chip of any physical server included in the logical node, the cloud management platform can divide the direct-out channels of the computing chip into two parts, that is, the first direct-out channel and the second direct-out channel. When the first direct-out channel is in a working state, the computing chip of the physical server can access the second sub-network through the first direct-out channel. When the first direct-out channel is in a fault state, the cloud management platform can instruct the computing chip of the physical server to access the second sub-network through the second direct-out channel. In this way, the cloud management platform divides the direct-out channels of the computing chips in the logical node of the tenant into a direct-out channel implementing a master port (the first direct-out channel) and a direct-out channel implementing a backup port (the second direct-out channel), so that the computing chips can continue to access the network between the computing chips in the logical node through the direct-out channel implementing the backup port when the direct-out channel implementing the master port fails, thereby enabling the computing chips in the logical node to have a certain fault tolerance, that is, a fault isolation capability.

[0015] In a possible implementation, the bandwidth of the first direct-out channel is determined based on any one or any combination of the number of the first direct-out channels, the communication protocol used by the first direct-out channels, and the physical link constructed by the first direct-out channels.

[0016] In a possible implementation, the straight-out channel is a line board, a cable or an optical fiber.

[0017] In a possible implementation, the computing chip of each physical server is a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU) or a data processing unit (DPU).

[0018] In a possible implementation, the first network or the second network is implemented through a peripheral component interconnect express (PCIE) network, an infiniBand (IB) network or a compute express link (CXL) network.

[0019] A second aspect of the embodiments of the present application provides a cloud service-based job task processing method, the method is applied to a cloud management platform, the cloud management platform is used for managing an infrastructure providing cloud services for tenants, the infrastructure includes a plurality of physical server groups, each physical server group includes a plurality of physical servers, a CPU, a memory, a network card, a disk and a computing chip in each physical server are interconnected inside the physical server, the network card of each physical server accesses to a first network, the first network is used to realize the interconnection of the network cards in each physical server group within and between groups, the computing chip of each physical server includes a direct-out channel, and the computing chip of each physical server accesses to a second network through the direct-out channel, the second network is used to realize the interconnection of the computing chips in each physical server group within and between groups; the method comprises the following steps: the cloud management platform acquires a job task processing request input by a tenant, the job task processing request includes data of a job task of the tenant and performance requirements of the job task; the cloud management platform creates a logical node of the tenant based on the job task processing request, the logical node includes at least two idle physical servers in the plurality of physical server groups that meet the performance requirements, a first sub-network of the first network and a second sub-network of the second network; the cloud management platform determines an operating system image input or selected by the tenant; and the cloud management platform notifies the logical node to install the operating system image, in the logical node in which the operating system image is installed, the network card of the at least two physical servers is used to receive data sent by the cloud management platform through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers are used to store the data, and the CPU of the at least two physical servers is used to notify the computing chip of the at least two physical servers to jointly process the data through the second sub-network, so as to complete the job task.

[0020] In a possible implementation manner, the second network includes a first switch chip corresponding to each physical server group and a second switch chip corresponding to the plurality of physical server groups; the computing chip in each physical server group accesses the first switch chip through the direct-out channel, and the first switch chip is used to support the interconnection of the computing chips in each physical server group within the group; the computing chips in the plurality of physical server groups access the second switch chip through the direct-out channel, and the second switch chip is used to support the interconnection of the computing chips in the plurality of physical server groups between groups.

[0021] In a possible implementation manner, the second network includes a second switch chip corresponding to the plurality of physical server groups; the computing chips in each physical server group are interconnected within the group through the direct-out channel; the computing chips in the plurality of physical server groups access the second switch chip through the direct-out channel, and the second switch chip is used to support the interconnection of the computing chips in the plurality of physical server groups between groups.

[0022] In a possible implementation, when the message associated with the data is transmitted between the computing chips of the at least two physical servers through the second sub-network, the sub-network is configured to complete the transmission of the message after determining that the message carries the identifier of the second sub-network.

[0023] In a possible implementation, the direct access channel includes a first direct access channel and a second direct access channel, when the first direct access channel is in a working state, the computing chips of the at least two physical servers access the second sub-network through the first direct access channel, and when the first direct access channel is in a fault state, the computing chips of the at least two physical servers access the second sub-network through the second direct access channel.

[0024] In a possible implementation, the bandwidth of the first direct access channel is determined based on any one or any combination of the number of the first direct access channels, the communication protocol used by the first direct access channels, and the physical link constructed by the first direct access channels.

[0025] In a possible implementation, the direct access channel is a line board, a cable, or an optical fiber.

[0026] In a possible implementation, the computing chip of each physical server is a GPU, an NPU, a DPU, or a TPU.

[0027] In a possible implementation, the first network or the second network is implemented through a PCIE network, an IB network, or a CXL network.

[0028] A third aspect of the embodiments of the present application provides a cloud management platform, the cloud management platform being configured to manage an infrastructure for providing cloud services for tenants, the infrastructure comprising a plurality of physical server groups, each of the physical server groups comprising a plurality of physical servers, a CPU, a memory, a network card, a disk and a computing chip in each of the physical servers being interconnected within the physical server, the network card of each of the physical servers being connected to a first network, the first network being configured to realize interconnection of the network cards in each of the physical server groups within the group and between the groups, the computing chip of each of the physical servers comprising a direct-out channel, the computing chip of each of the physical servers being connected to a second network through the direct-out channel, the second network being configured to realize interconnection of the computing chips in each of the physical server groups within the group and between the groups; the cloud management platform comprising: an obtaining module configured to obtain a job task processing request input by a tenant, the job task processing request comprising data of a job task of the tenant and performance requirements of the job task; a creating module configured to create a logical node of the tenant based on the job task processing request, the logical node comprising at least two idle physical servers in the plurality of physical server groups that meet the performance requirements, a first sub-network of the first network and a second sub-network of the second network; a determining module configured to determine an operating system image input or selected by the tenant; and a notifying module configured to notify the logical node to install the operating system image, in the logical node in which the operating system image is installed, the network card of the at least two physical servers being configured to receive data sent by the cloud management platform through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers being configured to store the data, and the CPU of the at least two physical servers being configured to notify the computing chip of the at least two physical servers to jointly process the data through the second sub-network, so as to complete the job task.

[0029] In a possible implementation manner, the second network comprises a first switch chip corresponding to each of the physical server groups and a second switch chip corresponding to the plurality of physical server groups; the computing chip in each of the physical server groups is connected to the first switch chip through the direct-out channel, and the first switch chip is configured to support interconnection of the computing chips in each of the physical server groups within the group; the computing chips in the plurality of physical server groups are connected to the second switch chip through the direct-out channel, and the second switch chip is configured to support interconnection of the computing chips in the plurality of physical server groups between the groups.

[0030] In a possible implementation manner, the second network comprises a second switch chip corresponding to the plurality of physical server groups; the computing chips in each of the physical server groups are interconnected within the group through the direct-out channel; the computing chips in the plurality of physical server groups are connected to the second switch chip through the direct-out channel, and the second switch chip is configured to support interconnection of the computing chips in the plurality of physical server groups between the groups.

[0031] In a possible implementation, when the message associated with the data is transmitted between the computing chips of the at least two physical servers through the second sub-network, the sub-network is configured to complete the transmission of the message after determining that the message carries the identifier of the second sub-network.

[0032] In a possible implementation, the direct channels include a first direct channel and a second direct channel, when the first direct channel is in a working state, the computing chips of the at least two physical servers access the second sub-network through the first direct channel, and when the first direct channel is in a fault state, the computing chips of the at least two physical servers access the second sub-network through the second direct channel.

[0033] In a possible implementation, the bandwidth of the first direct channel is determined based on any one or any combination of the number of the first direct channels, the communication protocol used by the first direct channels, and the physical link constructed by the first direct channels.

[0034] In a possible implementation, the direct channel is a line board, a cable, or an optical fiber.

[0035] In a possible implementation, the computing chip of each physical server is a GPU, an NPU, a DPU, or a TPU.

[0036] In a possible implementation, the first network or the second network is implemented through a PCIE network, an IB network, or a CXL network.

[0037] A fourth aspect of the embodiments of the present application provides a cloud service system, which includes the infrastructure as described in the first aspect or any possible implementation manner of the first aspect, and the cloud management platform as described in the second aspect or any possible implementation manner of the second aspect.

[0038] A fifth aspect of the embodiments of the present application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory: the memory is configured to store instructions; and the processor is configured to execute the instructions to cause the computing device cluster to perform the method as described in the second aspect or any possible implementation manner of the second aspect.

[0039] A sixth aspect of the embodiments of the present application provides a computer storage medium, which stores one or more instructions, the instructions causing one or more computers to implement the method as described in the second aspect or any possible implementation manner of the second aspect when executed by the one or more computers.

[0040] A seventh aspect of the embodiments of the present application provides a computer program product, which stores instructions, the instructions causing a computer to implement the method as described in the second aspect or any possible implementation manner of the second aspect when executed by the computer.

[0041] In the embodiment, the physical server cluster in the cloud service system can include a plurality of physical server groups, each of which can include a plurality of physical servers, and each of which can include a CPU, a memory, a network card, a disk and a computing chip. For any one of the plurality of physical server groups, the CPU, the memory, the network card, the disk and the computing chip included in the physical server are interconnected inside the physical server, and the network cards of all the physical servers in the plurality of physical servers are connected to the first network, which can interconnect the network cards of all the physical servers in the plurality of physical servers within and between groups, and the computing chips of all the physical servers in the plurality of physical servers can be connected to the second network through the direct channel provided by the computing chips, and the second network can interconnect the computing chips of all the physical servers in the plurality of physical servers within and between groups. In the foregoing system, when the computing chip of a certain physical server needs to communicate with the computing chip of another physical server, the computing chip of the physical server can directly communicate with the computing chip of another physical server through the second network, without the need for the CPU and the network card of the physical server to access the first network to communicate with the computing chip of another physical server. As can be seen, in the physical server cluster, the communication between the computing chips is no longer limited by the CPU and the network card, and the cloud vendor can increase the number of computing chips in the physical server to increase the number of computing chips in the physical server cluster, which is equivalent to expanding the size of the cluster, thereby making the cluster have stronger computing performance and providing better cloud services for tenants. BRIEF DESCRIPTION OF DRAWINGS

[0042] FIG. 1 is a structural schematic diagram of a cloud service system provided by an embodiment of the present application;

[0043] FIG. 2 is a structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0044] FIG. 3 is another structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0045] FIG. 4 is another structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0046] FIG. 5 is another structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0047] FIG. 6 is another structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0048] FIG. 7 is another structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0049] FIG. 8 is another structural schematic diagram of an infrastructure provided by an embodiment of the present application;

[0050] FIG. 9 is a flow diagram of a cloud service-based job task processing method according to an embodiment of the present application;

[0051] FIG. 10 is another structural diagram of an infrastructure according to an embodiment of the present application;

[0052] FIG. 11 is a structural diagram of a cloud management platform according to an embodiment of the present application;

[0053] FIG. 12 is a structural diagram of a computing device according to an embodiment of the present application;

[0054] FIG. 13 is a structural diagram of a computing device cluster according to an embodiment of the present application;

[0055] FIG. 14 is a diagram of a network connection between computing devices in a computer cluster according to an embodiment of the present application. DETAILED DESCRIPTION

[0056] The server system and the cloud service-based job task processing method provided by the embodiments of the present application can provide better cloud services for tenants.

[0057] The terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms used in this way can be interchanged, and this is only a way of distinguishing the objects with the same attributes in the description of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or devices containing a series of units do not have to be limited to those units, but can include other units not clearly listed or inherent to these processes, methods, products or devices.

[0058] With the rapid development of cloud technology, more and more tenants choose to use clusters provided by cloud vendors to complete their own job tasks. These clusters often contain multiple extension processors. Since each extension processor has certain specifications and performance, these clusters often have large enough specifications and strong enough performance to efficiently complete tenants' job tasks and meet tenants' job requirements.

[0059] In the related art, a cluster provided by a cloud vendor for a tenant can include multiple physical servers, each of which can include a certain number of extension processors that can be used to process a job task of the tenant. Among the multiple physical servers of the cluster, the physical distance between some of the physical servers can be far. In these physical servers, when an extension processor of a certain physical server needs to send a message of a job task to an extension processor of another physical server, it can successfully send the message to the extension processor of the other physical server through a network between the two. For example, the cluster includes multiple physical server groups, physical server group 1 is arranged in data center 1, and physical server 2 is arranged in data center 2. When an extension processor on physical server 1 in physical server group 1 needs to communicate with an extension processor on physical server 2 in physical server group 2, the extension processor on physical server 1 can send a message to the extension processor on physical server 2 through a network between the server groups.

[0060] In the above process, the extension processor of the physical server needs to access the network through the CPU and the network card of the physical server (for example, the extension processor on physical server 1 can access the network between the server groups through the central processor and the network card of physical server 1). Such a network connection method often limits the size of the cluster, that is, the number of extension processors included in the cluster is limited. In some large demand scenarios, it is often impossible to meet the demand of the tenant to build a large-scale cluster.

[0061] To solve the above problems, an embodiment of the present application provides a job task processing method based on a cloud service system. The method can be implemented through a cloud service system. FIG. 1 is a structural schematic diagram of a cloud service system provided by an embodiment of the present application. As shown in FIG. 1, the cloud service system includes infrastructure that can provide cloud services and a cloud management platform that manages the infrastructure. The cloud management platform and the infrastructure are introduced respectively as follows:

[0062] The cloud management platform can manage the infrastructure in the whole cloud service system (for example, in the infrastructure, a logical node serving a tenant is created according to the instruction of the tenant, the logical node can be used to execute the business of the tenant, so as to meet the business demand of the tenant, and the like), and the cloud management platform can also be open to the tenant outside the cloud service system and respond to the request of the tenant. For example, the cloud management platform can provide various interfaces such as a login interface and a task processing interface for the client of the tenant (for example, a terminal device used by the tenant or a browser on the terminal device, and the like) to access. Among them, the cloud management platform can authenticate the client of the tenant through the login interface, and allow the client of the tenant to log in to the cloud management platform after the authentication is successful. For another example, the cloud management platform can also allow the client of the tenant to send a job task processing request of the tenant to the cloud management platform through the task processing interface. Since the job task processing request is used to indicate a job task to be processed, data required by the job task, performance requirement and the like, the cloud management platform can select a plurality of physical servers meeting the performance requirement required by the job task from a plurality of physical server groups of a physical server cluster based on the job task processing request, and the plurality of computing chips are usually located on different physical server groups (hyperpods) of the physical server cluster (of course, they can also be located in the same physical server group). Then, the cloud management platform can create a logical node serving the tenant on the plurality of physical servers selected, and the plurality of physical servers included in the logical node can process the data of the job task under the control of the cloud management platform, so as to complete the job task, thereby meeting the job demand of the tenant.

[0063] The infrastructure includes a physical server cluster, the physical server cluster includes a plurality of physical server groups, each physical server group includes a plurality of physical servers, and each physical server can include at least one CPU, at least one memory, at least one network card, at least one disk and at least one computing chip (also referred to as an extended processor (xPU)). The network connection mode between the devices of each physical server in the plurality of physical server groups is introduced as follows:

[0064] (1) As shown in FIG. 2 (which is a structural schematic diagram of the infrastructure provided by the embodiment of the present application), for any one of the plurality of physical server groups, the CPU, the memory, the network card, the disk and the computing chip included in the physical server can be interconnected in the internal of the physical server, that is, the CPU of the physical server can call the memory, the network card, the disk and the computing chip of the physical server, and the computing chip of the physical server can also call the memory, the network card and the disk of the physical server.

[0065] (2) For any one of the plurality of physical server groups, the network cards of the physical servers of the physical server group can access the first network. For the first network, the first network can include the first intra-group network of any one of the plurality of physical server groups and the first inter-group network between the plurality of physical servers, wherein for any one of the plurality of physical server groups, the first intra-group network of the physical server group can enable the network cards of all the physical servers in the physical server group to be interconnected within the physical server group, and for the plurality of physical server groups, the first inter-group network can enable the network cards of all the physical servers in the plurality of physical server groups to be interconnected between the plurality of physical server groups.

[0066] (3) For any one of the plurality of physical server groups, the computing chips of the physical servers of the physical server group can access the second network. For the second network, the second network can include the second intra-group network (hypernet) of any one of the plurality of physical server groups and the second inter-group network (hyperclusternet) between the plurality of physical servers, wherein for any one of the plurality of physical server groups, the second intra-group network of the physical server group can enable the computing chips of all the physical servers in the physical server group to be interconnected within the physical server group, and for the plurality of physical server groups, the second inter-group network can enable the computing chips of all the physical servers in the plurality of physical server groups to be interconnected between the plurality of physical server groups.

[0067] Specifically, for any one of the plurality of physical server groups, the computing chips of any two physical servers in the physical server group can be communicatively connected through the direct-out channel thereof, or the computing chips of the two physical servers can access the first switching chip (of the switch) through the direct-out channel thereof to be communicatively connected. Similarly, the computing chips of the physical servers of any two physical server groups in the plurality of physical server groups can also access the second switching chip through the direct-out channel thereof to be communicatively connected.

[0068] Therefore, the direct-out channel, the first switching chip and the second switching chip connect all the computing chips of the plurality of physical server groups, and form a second network (interconnection network) between all the computing chips. In the second network, the intra-group network of any one of the plurality of physical server groups can include the direct-out channel of all the computing chips of the physical servers in the physical server group or the first switching chip corresponding to the physical server group, and the inter-group network between the plurality of physical server groups can include the second switching chip.

[0069] It is worth noting that for any one computing chip in the multiple physical server groups, the number of direct channels of the computing chip is usually multiple, which can be divided into two parts, one part is x channel, and the other part is y channel, the x channel is used for the computing chip to access the second intra-group network of the physical server group where the computing chip is located, and the y channel is used for the computing chip to access the second inter-group network between the multiple physical server groups. As shown in FIG. 3 (FIG. 3 is another structure diagram of the infrastructure provided by the embodiment of the present application), the computing chip includes a transmitter and a receiver, the transmitter of the computing chip can be directly connected with the receiver at the opposite end through the direct channel of the computing chip, and the receiver of the computing chip can be directly connected with the transmitter at the opposite end through the direct channel of the computing chip, when the direct channels are x channels, the opposite end is the second switching chip corresponding to the physical server group where the computing chip is located or the remaining computing chips in the physical server group, so the computing chip can use the x channels it has to realize communication connection with the remaining computing chips in the physical server group (which can be directly connected with the remaining computing chips in the physical server group through the x channels, or indirectly connected with the remaining computing chips in the physical server group through the x channels by first accessing the first switching chip corresponding to the physical server group, etc.), so as to build the second intra-group network of the physical server group among all the computing chips in the physical server group. When the direct channels are y channels, the opposite end is the second switching chip or the computing chip in other physical server groups, so the computing chip can use the y channels it has to first access the second switching chip, and then indirectly connect with the computing chips in other physical server groups, so as to build the second inter-group network between the multiple physical server groups among all the computing chips in the multiple physical server groups.

[0070] It is also worth noting that for the multiple physical server groups, the physical networking mode between the multiple physical server groups can be multiple modes. For example, as shown in FIG. 4 (FIG. 4 is another structure diagram of the infrastructure provided by the embodiment of the present application), for any one physical server group in the multiple physical server groups, all the computing chips in the physical server group are connected with each other through the direct channels possessed by the computing chips, so as to build the second intra-group network of the physical server group. All the computing chips in the physical server group access the second switching chip through the direct channels possessed by the computing chips, and all the computing chips of other physical server groups also access the second switching chip through the direct channels possessed by the computing chips, so as to build the second inter-group network between the multiple physical server groups.

[0071] For example, as shown in FIG. 5 (FIG. 5 is another structure diagram of the infrastructure provided by the embodiment of the present application), for any one of the plurality of physical server groups, all the computing chips in the physical server group access the first switching chip corresponding to the physical server group through the direct access channel possessed by the computing chips, to build the second intra-group network of the physical server group. All the computing chips in the physical server group access the second switching chip through the direct access channel possessed by the computing chips, and all the computing chips in other physical server groups access the second switching chip through the direct access channel possessed by the computing chips, to build the second inter-group network between the plurality of physical server groups. It is worth noting that the first switching chip in the physical server group can include multiple layers (for example, two layers of first switching chips in FIG. 5), and the multiple layers of first switching chips need to be combined to complete the connection of all the computing chips in the physical server group, so the second intra-group network built by the multiple layers of first switching chips is a single network plane network. Similarly, the second switching chip between the plurality of physical server groups can also include multiple layers (for example, two layers of second switching chips in FIG. 5), and the multiple layers of second switching chips need to be combined to complete the connection between all the computing chips of the plurality of physical server groups, so the second inter-group network built by the multiple layers of second switching chips is also a single network plane network.

[0072] For example, as shown in FIG. 6 (FIG. 6 is another structure diagram of the infrastructure provided by the embodiment of the present application), the first switching chip in the physical server group can include one layer (for example, one layer of first switching chips in FIG. 6), and each first switching chip in the layer of first switching chips can separately complete the connection of all the computing chips in the physical server group, so the second intra-group network built by the layer of first switching chips is a multi-network plane network (each first switching chip in the layer of first switching chips forms one network plane of the second intra-group network of the physical server group). Similarly, the second switching chip between the plurality of physical server groups can also include one layer (for example, one layer of second switching chips in FIG. 6), and each second switching chip in the layer of second switching chips can also separately complete the connection between all the computing chips of the plurality of physical server groups, so the second inter-group network built by the layer of second switching chips is also a multi-network plane network (each second switching chip in the layer of second switching chips forms one network plane of the second inter-group network).

[0073] It is also worth noting that after the cloud management platform obtains a job task processing request for a job task of a tenant, the cloud management platform can select a number of computing chips (i.e., the aforementioned at least two physical servers) from a plurality of physical server groups based on the request, and create a logical node of the tenant on the number of computing chips. As shown in FIG. 7 (FIG. 7 is another structural diagram of an infrastructure provided by an embodiment of the present application, and FIG. 7 is drawn based on FIG. 4), the logical node includes the number of physical servers, a network between network cards of the number of physical servers (i.e., a first sub-network logically divided from the first network (not shown in FIG. 7)), and a network between computing chips of the number of physical servers (i.e., a second sub-network logically divided from the second network).

[0074] Since the number of physical servers are located in at least one of the plurality of physical server groups, the second sub-network can include a hypersubnet of a second intra-group network of the number of physical server groups and a hypercluster subnet of a second inter-group network between the plurality of physical server groups.

[0075] In order to further understand the aforementioned second network, second intra-group network, second inter-group network, and second sub-network, the following further introduces these concepts in combination with FIG. 8 and a specific application example. As shown in FIG. 8 (FIG. 8 is another structural diagram of an infrastructure provided by an embodiment of the present application), it is assumed that there are a physical server group 1 and a physical server group 2, the physical server group 1 includes physical servers 1 to 4, and the physical server group 2 includes physical servers 5 to 8. Since each physical server includes a computing chip, the physical server group 1 includes computing chips 1 to 4, and the physical server group 2 includes computing chips 5 to 8. It should be noted that each physical server also includes a CPU, a memory, a network card, and a disk, which are not expanded in this application example.

[0076] Since each of the computing chips 1-8 accesses the interconnection network (the aforementioned second network) through its own direct channel, the interconnection network includes the hypernet 1 accessed by the computing chips 1-4 through their x channels, the hypernet 2 accessed by the computing chips 5-8 through their x channels, and the hypercluster net accessed by the computing chips 1-8 through their y channels, the hypernet 1 being the intra-group network of the physical server group 1 (i.e., the aforementioned second intra-group network), the hypernet 2 being the intra-group network of the physical server group 2, and the hypercluster net being the inter-group network between the physical server 1 and the physical server 2 (the aforementioned second inter-group network), wherein the intra-group network is a large-bandwidth small-scale network, and the inter-group network is a small-bandwidth large-scale network.

[0077] Suppose the cloud management platform needs to create the logical node 1 for the tenant 1 and the logical node 2 for the tenant 2, the cloud management platform can select the computing chips 1, 2, 5, and 6 to create the logical node 1, and select the computing chips 3, 4, 7, and 8 to create the logical node 2.

[0078] Then, the hypernet 1 can be logically divided by the cloud management platform into the sub-network hypersubnet 1 and the sub-network hypersubnet 2, the hypersubnet 1 being the part of the hypernet 1 built between the computing chips 1 and 2, and the hypersubnet 2 being the part of the hypernet 1 built between the computing chips 3 and 4.

[0079] Similarly, the hypernet 2 can be logically divided by the cloud management platform into the sub-network hypersubnet 3 and the sub-network hypersubnet 4, the hypersubnet 3 being the part of the hypernet 2 built between the computing chips 5 and 6, and the hypersubnet 4 being the part of the hypernet 2 built between the computing chips 7 and 8.

[0080] Similarly, the hypercluster net can be logically divided into a subnet hypercluster subnet 1 and a subnet hypercluster subnet 2 by the cloud management platform, the hypercluster subnet 1 is the part of the hypercluster net constructed between the computing chip 1, the computing chip 2, the computing chip 5 and the computing chip 6, and the hypercluster subnet 2 is the part of the hypercluster net constructed between the computing chip 3, the computing chip 4, the computing chip 7 and the computing chip 8.

[0081] Therefore, the hypersubnet 1, the hypersubnet 3 and the hypercluster subnet 1 are the subnets of the interconnection network allocated to the logical node 1 by the cloud management platform, the hypersubnet 2, the hypersubnet 4 and the hypercluster subnet 2 are the subnets of the interconnection network allocated to the logical node 2 by the cloud management platform, and the subnets owned by the two logical nodes are isolated and do not interfere with the job tasks performed by each other.

[0082] It should be understood that, in the present application example, only one hyper net containing two hypersubnets is illustratively introduced, and in actual application, the hyper net can be logically divided into more hypersubnets by the cloud management platform, and the number of hypersubnets is not limited here. Similarly, in the present application example, only one hypercluster net containing two hypercluster subnets is illustratively introduced, and in actual application, the hypercluster net can be logically divided into more hypercluster subnets by the cloud management platform, and the number of hypercluster subnets is not limited here.

[0083] Further, after the cloud management platform creates the logical node for the tenant based on the job task processing request of the tenant, the cloud management platform can install the operating system image specified by the tenant in the logical node. Since the logical node installed with the operating system image contains a plurality of physical servers selected by the cloud management platform, the CPUs of the plurality of physical servers are equivalent to being installed with the operating system. Therefore, the cloud management platform can send the job task of the tenant and the data of the job task to the network cards of the plurality of physical servers, so that the CPUs of the plurality of physical servers store the data in the memories and disks of the plurality of physical servers based on the job task. Then, the CPUs of the plurality of physical servers call the computing chips of the plurality of physical servers, obtain the data from the memories and disks of the plurality of physical servers, and jointly process the data (through the second sub-network between the computing chips of the plurality of physical servers) to complete the job task.

[0084] Further, after the CPUs of the plurality of physical servers are installed with the operating system, the cloud management platform can further create virtual instances on the CPUs of the plurality of physical servers to call the computing chips of the plurality of physical servers to process the data through the virtual instances, so as to complete the job task. The virtual instances can be presented in various ways, for example, virtual machines (VMs) created by the cloud management platform on the physical servers selected by the cloud management platform through virtualization technology, for example, the virtual instances can also be containers (dockers) created by the cloud management platform on the physical servers selected by the cloud management platform through virtualization technology, for example, the virtual instances can also be micro virtual machines (microVMs) created by the cloud management platform on the physical servers selected by the cloud management platform through virtualization technology, and the like.

[0085] Further, in the embodiments of the present application, since the computing chips in the plurality of physical server groups in the physical server cluster are connected to the second network, the number of computing chips in each physical server cluster can be increased due to the existence of the second network. For example, the number of all computing chips in any one physical server group can reach 384, and the total number of all computing chips in the plurality of physical server groups can reach 128000. It can be seen that the scale of the physical server cluster is large and has sufficient computing performance and specifications.

[0086] Further, for any one of the computing chips in the plurality of physical server groups, the computing chip can be any one of a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), a data processing unit (DPU), and the like.

[0087] Further, for any one of the computing chips in the plurality of physical server groups, the direct-out channel of the computing chip can be presented in various ways, for example, the direct-out channel of the computing chip can be a printed circuit board (PCB), and for another example, the direct-out channel of the computing chip can also be a cable, and for another example, the direct-out channel of the computing chip can also be an optical fiber, and the like.

[0088] Further, for the plurality of physical server groups, the plurality of physical server groups can be deployed in the same site or different sites, and the site can be presented in various forms, for example, the site can be a region in the infrastructure, and for another example, the site can be an availability zone in the infrastructure, and for another example, the site can be a data center (DC) in the infrastructure, and for another example, the site can be a room in the infrastructure, and the like.

[0089] Further, the first network or the second network can have various presentation forms, for example, the first network or the second network can be a peripheral component interconnect express (PCIE) network, and for another example, the first network or the second network can also be an infiniBand (IB) network, and for another example, the first network or the second network can also be a compute express link (CXL) network, and for another example, the first network or the second network can also be an interconnection network between devices developed by a cloud vendor (the bandwidth of the network also needs to be not less than 40G / S), and the like. Accordingly, the direct-out channel possessed by each computing chip and the switching chip accessed thereby can be a communication device based on a PCIE protocol, a CXL protocol, an IB protocol, or a communication protocol developed by a cloud vendor. Similarly, the network card of each physical server and the interface of the computing chip for accessing the first network or the second network can be a PCIE interface, a CXL interface, an IB interface, or an interface developed by a cloud vendor, and the like.

[0090] Further, when the job task of the tenant is a model training task, in the second sub-network contained in the logical node of the tenant, the model training task can be completed between the computing chips of the sub-network accessing the second intra-group network in a manner of tensor parallelism (TP) and sequence parallelism (SP), and the model training task can be completed between the computing chips of the sub-network accessing the second inter-group network in a manner of pipeline parallelism (PP) and data parallelism (DP).

[0091] Based on the cloud service system, the physical server cluster in the cloud service system can include a plurality of physical server groups, each of which can include a plurality of physical servers, and each of which includes a CPU, a memory, a network card, a disk and a computing chip. For any one of the plurality of physical server groups, the CPU, memory, network card, disk and computing chip contained in the physical server are interconnected inside the physical server, and the network cards of all physical servers in the plurality of physical servers are connected to the first network. The first network can interconnect the network cards of all physical servers in the plurality of physical servers within and between groups, and the computing chips of all physical servers in the plurality of physical servers can be connected to the second network through the direct channel they have. The second network can interconnect the computing chips of all physical servers in the plurality of physical servers within and between groups. In the foregoing system, when the computing chip of a physical server needs to communicate with the computing chip of another physical server, the computing chip of the physical server can directly communicate with the computing chip of another physical server through the second network, without the need for the CPU and network card of the physical server to access the first network to communicate with the computing chip of another physical server. As can be seen, in the physical server cluster, the communication between the computing chips is no longer limited by the CPU and the network card. The cloud vendor can increase the number of computing chips in the physical server to increase the number of computing chips in the physical server cluster, which is equivalent to expanding the size of the cluster, thereby making the cluster have stronger computing performance and providing better cloud services for tenants. In order to further understand the working process of the cloud service system, the working process is further introduced below in combination with FIG. 9. FIG. 9 is a flowchart of a cloud service-based job task processing method according to an embodiment of the present application. The method can be implemented by the cloud service system shown in FIG. 1. The cloud management platform includes infrastructure for providing cloud services for tenants and a cloud management platform for managing the infrastructure. The infrastructure includes a plurality of physical server groups, each of which includes a plurality of physical servers. For any one of the plurality of physical server groups, the CPU, memory, network card, disk and computing chip in the physical server are interconnected inside the physical server, and the network cards of all physical servers in the plurality of physical servers are connected to the first network. The first network is used to interconnect the network cards of all physical servers in the plurality of physical server groups within and between groups, and the computing chips of all physical servers in the plurality of physical server groups each include a direct channel. The computing chips of these physical servers are connected to the second network through the direct channel they have. The second network is used to interconnect the computing chips of all physical servers in the plurality of physical server groups within and between groups. The method comprises:

[0092] 901、The cloud management platform obtains a job task processing request input by a tenant, the job task processing request including data of a job task of the tenant and performance requirements of the job task.

[0093] In this embodiment, when the tenant has a job task to be processed, the cloud management platform can provide a task processing interface (e.g., a job task input box of a tenant interface, etc.) to a client of the tenant. Then, the tenant can input a job task processing request set by the tenant to the task processing interface through the client used by the tenant, the job task processing request including the job task to be processed by the tenant, data of the job task, performance requirements of the job task, etc. In this way, the cloud management platform can receive the job task processing request sent by the client of the tenant through the task processing interface.

[0094] It should be noted that the data of the job task generally refers to data to be processed in the job task, and the performance requirements of the job task generally refer to information such as specifications of CPU, memory, network card, disk, and computing chip required for processing the job task.

[0095] 902、The cloud management platform creates a logical node of the tenant based on the job task processing request, the logical node including at least two physical servers that meet the performance requirements and are idle in a plurality of physical server groups, a first sub-network of a first network, and a second sub-network of a second network.

[0096] After obtaining the job task processing request, the cloud management platform can analyze the job task processing request to obtain the data of the job task and the performance requirements of the job task. Then, the cloud management platform can select a plurality of physical servers that meet the performance requirements of the job task and are idle in a plurality of physical server groups, and create a logical node of the tenant on the plurality of physical servers. The logical node includes the plurality of physical servers, a network between network cards of the plurality of physical servers, and a network between computing chips of the plurality of physical servers. The network between the network cards of the plurality of physical servers is a first sub-network logically divided from a first network, and the network between the computing chips of the plurality of physical servers is a second sub-network logically divided from a second network.

[0097] Specifically, for the second sub-network contained in the logical node, the cloud management platform can configure a global unique identifier for the second sub-network. Then, in the process of processing the job task, if the computing chip of a certain physical server needs to send a message associated with the data of the job task to the computing chip of another physical server (for example, the message carries the processing result of the data, etc.), the computing chip of the certain physical server can send the message to the second sub-network, and the second sub-network can determine whether the message carries the unique identifier of the second sub-network. If yes, the second sub-network forwards the message to the computing chip of the other physical server; if not, the second sub-network discards the message. As can be seen, the cloud management platform assigns a unique identifier to the network between the computing chips inside the logical node of the tenant (i.e., the second sub-network), so that the network between the computing chips inside the logical nodes of different tenants is isolated and does not affect the processing of the respective job tasks.

[0098] For example, as shown in FIG. 10 (FIG. 10 is another structure diagram of the infrastructure provided by the embodiment of the present application, and FIG. 10 is drawn on the basis of FIG. 8), after the cloud management platform receives the job task processing request of tenant 1, the cloud management platform can confirm the data of the job task to be processed by tenant 1 and the performance requirement, so the cloud management platform can select physical server 1, physical server 2, physical server 5 and physical server 6 that meet the performance requirement from physical server group 1 and physical server group 2 to construct the logical node 1 of tenant 1. The logical node 1 contains the four physical servers, the network between the network cards of the four physical servers (not shown in the figure) and the network between the computing chips (computing chip 1, computing chip 2, computing chip 5 and computing chip 6) of the four physical servers, that is, hypersubnet 1, hypersubnet 3 and hypercluster subnet 1.

[0099] Then, the cloud management platform can allocate a globally unique identification (ID) for hypersubnetl, hypersubnet3, and hyperclustersubnetl, which can be shared by hypersubnetl, hypersubnet3, and hyperclustersubnetl, so that hypersubnetl, hypersubnet3, and hyperclustersubnetl are isolated domains 1 of tenant 1, and the ID is the identification of the isolated domains 1. In the subsequent process of jointly processing data of the job task by the four computing chips, if computing chip 1 needs to send a message associated with the data to computing chip 2, hypersubnetl receives the message sent by computing chip 1, detects that the message carries the identification of isolated domains 1, and then forwards the message to computing chip 2. Similarly, the message transmission between computing chip 1, computing chip 2, computing chip 5, and computing chip 6 through any one of hypersubnetl, hypersubnet3, and hyperclustersubnetl needs to carry the identification of isolated domains 1 to complete the transmission.

[0100] In this way, isolated domains 1 of tenant 1 and isolated domains 2 of tenant 2 (hypersubnet2, hypersubnet4, and hyperclustersubnet2) do not interfere with each other, and connectivity isolation is formed between the two.

[0101] More specifically, for the computing chip of any one of the physical servers contained in the logical node, the computing chip of the physical server is provided with a plurality of direct-out channels, and the cloud management platform can divide the plurality of direct-out channels into two parts, respectively, a first direct-out channel (a direct-out channel for implementing a primary port) and a second direct-out channel (a direct-out channel for implementing a backup port). Then, when the first direct-out channel is in a working state, the computing chip of the physical server can access the second sub-network through the first direct-out channel, and when the first direct-out channel is in a fault state, the cloud management platform can clear the information of the first direct-out channel so that the first direct-out channel is no longer used by the computing chip of the physical server, and enable the second direct-out channel so that the computing chip of the physical server accesses the second sub-network through the second direct-out channel, thereby ensuring that the computing chip of the physical server successfully performs the job task. It can be seen that, by dividing the direct-out channels of the computing chips inside the logical node of the tenant into direct-out channels for implementing primary ports and direct-out channels for implementing backup ports, the cloud management platform can enable the computing chips to continue to use the direct-out channels for implementing backup ports to access the network between the computing chips in the logical node in the case that the direct-out channels for implementing primary ports fail, so that the computing chips inside the logical node have a certain fault tolerance, that is, have the ability of fault isolation.

[0102] More specifically, for the computing chip of any one of the physical servers contained in the logical node, the cloud management platform can set the bandwidth of the first direct-out channel of the computing chip of the physical server, and the bandwidth of the second direct-out channel of the computing chip of the physical server, wherein the bandwidth of the first direct-out channel can be determined based on any one or any combination of the number of the first direct-out channels, the communication protocol used by the first direct-out channels, and the physical links constructed by the first direct-out channels, and similarly, the bandwidth of the second direct-out channel can be determined based on any one or any combination of the number of the first direct-out channels, the communication protocol used by the first direct-out channels, and the physical links constructed by the first direct-out channels. It should be noted that there can be multiple physical links in the second network that are constructed by the first direct-out channels of the computing chip of the physical server (i.e., these physical links contain the first direct-out channels of the computing chip of the physical server, and can also contain intermediate switch chips, etc.), but the cloud management platform will allocate a certain number of physical links from these multiple physical links to the computing chip of the physical server according to the bandwidth requirements of the first direct-out channels of the computing chip of the physical server. Moreover, if a certain physical link allocated to the computing chip of the physical server is shared with the computing chips of the remaining physical servers (i.e., the physical link is also constructed based on the first direct-out channels of the computing chips of the remaining physical servers), the cloud management platform can logically divide the physical link (which can be divided in the manner of a virtual link, etc.) to divide the physical link into multiple sub-physical links, and then allocate the multiple sub-physical links to the computing chip of the physical server and the computing chips of the remaining physical servers, respectively. It can be seen that the cloud management platform can achieve performance isolation for different computing chips of different physical servers, i.e., performance isolation.

[0103] Still as in the above example, for the logical node 1, the computing chip 1 can contain the primary port hyperport 1.1 and hyperclusterport 1.1, and contain the backup port hyperport 1.2 and hyperclusterport 1.2. The hyperport 1.1 is implemented based on a part of the x channels of the computing chip 1, the hyperport 1.2 is implemented based on another part of the x channels of the computing chip 1, the hyperclusterport 1.1 is implemented based on a part of the y channels of the computing chip 1, and the hyperclusterport 1.2 is implemented based on another part of the y channels of the computing chip 1.

[0104] Generally, the computing chip 1 can access the hyper subnet 1 using the hyperport 1.1 and access the hyper cluster subnet 1 using the hyper cluster port 1.1. When the hyperport 1.1 fails, the cloud management platform can clear the information of the hyperport 1.1 and enable the hyperport 1.2, so that the computing chip 1 can access the hyper subnet 1 using the hyperport 1.2. Similarly, the hyper cluster port 1.1 and the hyper cluster port 1.2 are also the same, which will not be repeated here. Further, the computing chip 2, the computing chip 5 and the computing chip 6 are also the same (only part of the standby port is shown in FIG. 10), which will not be repeated here.

[0105] In addition, the cloud management platform can also set the bandwidth of each port, for example, the cloud management platform can set the number of x channels occupied by the hyperport 1.1, the communication protocol used by the hyperport 1.1 and the performance indicators such as the physical link constructed by the hyperport 1.1 according to the bandwidth requirement of the hyperport 1.1, so that the hyperport 1.1 has a certain bandwidth. Similarly, the hyperport 1.2, the hyper cluster port 1.1 and the hyper cluster port 1.2 are also the same, which will not be repeated here. Further, the computing chip 2, the computing chip 5 and the computing chip 6 are also the same, which will not be repeated here.

[0106] It should be understood that for the logical node 1 of the tenant 1, the cloud management platform sets the identification of the isolation domain 1 for the logical node 1, the division of the master and standby ports and the bandwidth of the master and standby ports. These information can be specified by the tenant 1 and can be included in the performance requirement of the job task processing request sent by the tenant 1 to the cloud management platform. Of course, these information can also be automatically generated by the cloud management platform, which is not limited here.

[0107] 903, the cloud management platform determines the operating system image input or selected by the tenant.

[0108] 904, the cloud management platform informs the logical node to install the operating system image. In the logical node where the operating system image is installed, the network card of the at least two physical servers is used to receive the data sent by the cloud management platform through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers are used to store data, and the CPU of the at least two physical servers is used to inform the computing chip of the at least two physical servers to jointly process the data through the second sub-network to complete the job task.

[0109] After the logical node is created for the tenant, the cloud management platform can also obtain an operating system image input by the tenant or selected on the cloud management platform. Then, the cloud management platform can install the operating system image specified by the tenant in the logical node, since the logical node in which the operating system image is installed contains a plurality of physical servers selected by the cloud management platform, the CPUs of the plurality of physical servers are equivalent to installing an operating system, so the cloud management platform can send a job task of the tenant and data of the job task to the network cards of the plurality of physical servers through the first sub-network, so that the CPUs of the plurality of physical servers store the data in the memories and disks of the plurality of physical servers after receiving the job task and the data. Then, the CPUs of the plurality of physical servers can call the computing chips of the plurality of physical servers, obtain the data from the memories and disks of the plurality of physical servers, and jointly process the data through the second sub-network to complete the job task.

[0110] In the embodiments of the present application, the physical server cluster in the cloud service system can include a plurality of physical server groups, each of which can include a plurality of physical servers, and each of which can include a CPU, a memory, a network card, a disk and a computing chip. For any one of the plurality of physical server groups, the CPU, the memory, the network card, the disk and the computing chip contained in the physical server are interconnected inside the physical server, and the network cards of all physical servers in the plurality of physical servers are connected to the first network, which can interconnect the network cards of all physical servers in the plurality of physical servers within and between groups, and the computing chips of all physical servers in the plurality of physical servers can be connected to the second network through the direct channels they have, and the second network can interconnect the computing chips of all physical servers in the plurality of physical servers within and between groups. In the foregoing system, when the computing chip of a physical server needs to communicate with the computing chip of another physical server, the computing chip of the physical server can directly communicate with the computing chip of another physical server through the second network, without the need for the CPU and the network card of the physical server to access the first network to communicate with the computing chip of another physical server. As can be seen, in the physical server cluster, the communication between the computing chips is no longer limited by the CPU and the network card, and the cloud vendor can increase the number of computing chips in the physical server to increase the number of computing chips in the physical server cluster, which is equivalent to expanding the size of the cluster, thereby making the cluster have stronger computing performance and providing better cloud services for tenants.

[0111] Further, in the embodiments of the present application, all the computing chips in the plurality of physical server groups and the second network built among these computing chips can be provided for use by the plurality of tenants, and a certain scale of computing chips and the network (a sub-network of the second network) among these computing chips can be provided for each tenant according to the performance requirement of the job task of each tenant, so as to meet the customized requirement of the tenant on the computing service and the network service.

[0112] Further, in the embodiments of the present application, the cloud management platform can also cause network isolation between different tenants, including the aforementioned connectivity isolation, fault isolation, and performance isolation, and the like.

[0113] The above is a detailed description of the cloud service-based job task processing method provided by the embodiments of the present application. The cloud management platform provided by the embodiments of the present application will be introduced below. FIG. 11 is a structural schematic diagram of the cloud management platform provided by the embodiments of the present application. As shown in FIG. 11, the cloud management platform is used to manage the infrastructure for providing cloud services for tenants. The infrastructure includes a plurality of physical server groups. Each physical server group includes a plurality of physical servers. The CPU, memory, network card, disk, and computing chip in each physical server are interconnected inside the physical server. The network card of each physical server accesses the first network. The first network is used to realize the interconnection of the network cards in each physical server group within the group and between groups. The computing chip of each physical server includes a direct channel. The computing chip of each physical server accesses the second network through the direct channel. The second network is used to realize the interconnection of the computing chips in each physical server group within the group and between groups. The cloud management platform includes:

[0114] The obtaining module 1101 is configured to obtain a job task processing request input by a tenant. The job task processing request includes data of a job task of the tenant and a performance requirement of the job task.

[0115] The creating module 1102 is configured to create a logical node of the tenant based on the job task processing request. The logical node includes at least two idle physical servers in the plurality of physical server groups that meet the performance requirement, a first sub-network of the first network, and a second sub-network of the second network.

[0116] The determining module 1103 is configured to determine an operating system image input or selected by the tenant.

[0117] The notification module 1104 is configured to notify the logical node to install an operating system image, and in the logical node installed with the operating system image, the network card of the at least two physical servers is configured to receive data sent by the cloud management platform through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers are configured to store the data, and the CPU of the at least two physical servers is configured to notify the computing chip of the at least two physical servers to jointly process the data through the second sub-network to complete a job task.

[0118] In a possible implementation, the second network includes a first switch chip corresponding to each physical server group and a second switch chip corresponding to the plurality of physical server groups; the computing chips in each physical server group access the first switch chip through direct-out channels, and the first switch chip is configured to support the computing chips in each physical server group to be interconnected within the group; the computing chips in the plurality of physical server groups access the second switch chip through direct-out channels, and the second switch chip is configured to support the computing chips in the plurality of physical server groups to be interconnected between groups.

[0119] In a possible implementation, the second network includes a second switch chip corresponding to the plurality of physical server groups; the computing chips in each physical server group are interconnected within the group through direct-out channels; and the computing chips in the plurality of physical server groups access the second switch chip through direct-out channels, and the second switch chip is configured to support the computing chips in the plurality of physical server groups to be interconnected between groups.

[0120] In a possible implementation, when the computing chips of the at least two physical servers transmit a packet associated with the data through the second sub-network, the sub-network is configured to complete transmission of the packet after determining that the packet carries an identifier of the second sub-network.

[0121] In a possible implementation, the direct-out channel includes a first direct-out channel and a second direct-out channel; when the first direct-out channel is in a working state, the computing chips of the at least two physical servers access the second sub-network through the first direct-out channel; and when the first direct-out channel is in a fault state, the computing chips of the at least two physical servers access the second sub-network through the second direct-out channel.

[0122] In a possible implementation, the bandwidth of the first direct-out channel is determined based on any one or any combination of the number of the first direct-out channels, a communication protocol used by the first direct-out channels, and a physical link constructed by the first direct-out channels.

[0123] In a possible implementation, the direct-out channel is a line board, a cable, or an optical fiber.

[0124] In a possible implementation, the computing chip of each physical server is a GPU, an NPU, a DPU, or a TPU.

[0125] In a possible implementation, the first network or the second network is implemented through a PCIE network, an IB network, or a CXL network.

[0126] It should be noted that the information interaction and implementation process between the modules / units of the apparatus are based on the same concept as the method embodiments of the present application, and the technical effects brought by the same are the same as those of the method embodiments of the present application. For details, refer to the foregoing description of the method embodiments of the present application, which will not be repeated here.

[0127] Please refer to FIG. 12, which is a structural schematic diagram of a computing device provided by an embodiment of the present application. As shown in FIG. 12, the computing device 1200 (which can be used to present the cloud management platform mentioned above) includes a processor 1201, a memory 1202, a communication interface 1203, and a bus 1204. The processor 1201, the memory 1202, and the communication interface 1203 are coupled through a bus (not labeled in the figure). The memory 1202 stores instructions, and when the instructions in the memory 1202 are executed, the computing device 1200 performs the method performed by the cloud management platform in the above method embodiments.

[0128] The computing device 1200 can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For another example, when the units in the apparatus can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For another example, these units can be integrated together to implement a system-on-a-chip (SOC).

[0129] The processor 1201 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.

[0130] The memory 1202 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).

[0131] The memory 1202 stores executable program code, and the processor 1201 executes the executable program code to respectively implement the functions of the aforementioned acquisition module, creation module, determination module and notification module and the like, thereby implementing the aforementioned cloud service-based job task processing method. That is, the memory 1202 stores instructions for executing the aforementioned cloud service-based job task processing method.

[0132] The communication interface 1203 uses a transceiving module such as, but not limited to, a network interface card, a transceiver, and the like to enable communication between the computing device 1200 and other devices or communication networks.

[0133] The bus 1204 can include, in addition to a data bus, a power bus, a control bus, and a state signal bus, and the like. The bus can be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), and the like. The bus can be divided into an address bus, a data bus, a control bus, and the like.

[0134] Referring to FIG. 13, FIG. 13 is a structural schematic diagram of a computing device cluster provided by an embodiment of the present application. As shown in FIG. 13, the computing device cluster 1300 includes at least one computing device 1200.

[0135] As shown in FIG. 13, the computing device cluster 1300 includes at least one computing device 1200. The memory 1202 in one or more computing devices 1200 in the computing device cluster 1300 can store the same instructions for performing the cloud service-based job task processing method described above.

[0136] In some possible implementations, the memory 1202 in one or more computing devices 1200 in the computing device cluster 1300 can also respectively store partial instructions for performing the cloud service-based job task processing method described above. In other words, the combination of one or more computing devices 1200 can collectively perform the cloud service-based job task processing method described above.

[0137] It should be noted that the memory 1202 in different computing devices 1200 in the computing device cluster 1300 can store different instructions, respectively, for performing part of the functions of the cloud management platform described above. That is, the instructions stored in the memory 1202 in different computing devices 1200 can implement the functions of one or more of the obtaining module, the creating module, the determining module, and the notifying module, and the like.

[0138] In some possible implementation manners, one or more of the computing devices 1200 in the computing device cluster 1300 can be connected through a network. The network can be a wide area network or a local area network, etc.

[0139] Referring to FIG. 14, FIG. 14 is a schematic diagram of connection of the computing devices in the computer cluster through a network according to an embodiment of the present application. As shown in FIG. 14, two computing devices 1200A and 1200B are connected through a network. Specifically, the computing devices are connected to the network through the communication interfaces in the computing devices.

[0140] In a possible implementation manner, the memory in the computing device 1200A stores instructions for performing the functions of the obtaining module and the like. Meanwhile, the memory in the computing device 1200B stores instructions for performing the functions of the creating module, the determining module, the notifying module and the like.

[0141] It should be understood that the functions of the computing device 1200A shown in FIG. 14 can also be completed by multiple computing devices. Similarly, the functions of the computing device 1200B can also be completed by multiple computing devices.

[0142] The embodiments of the present application also relate to a computer storage medium, which stores a program for performing signal processing, and when the program runs on a computer, the computer executes the steps performed by the cloud management platform in the embodiment shown in FIG. 9.

[0143] The embodiments of the present application also relate to a computer program product, which stores instructions, and when the instructions are executed by a computer, the computer executes the steps performed by the cloud management platform in the embodiment shown in FIG. 9.

[0144] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0145] In the several embodiments of the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner for actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0146] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0147] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0148] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, read-only memory), a random access memory (RAM, random access memory), a magnetic disk or an optical disk, and various program code storage media.

Claims

1. A server system, characterized by The server system serves as an infrastructure providing cloud services, and the infrastructure includes a plurality of physical server groups, each of which includes a plurality of physical servers, and a central processing unit (CPU), a memory, a network card, a disk, and a computing chip in each of the physical servers are interconnected within the physical server; The network card of each physical server accesses a first network, and the first network is used to realize the interconnection of the network cards in each physical server group within the group and between groups; The computing chip of each physical server includes a direct access channel, and the computing chip of each physical server accesses a second network through the direct access channel, and the second network is used to realize the interconnection of the computing chips in each physical server group within the group and between groups.

2. The system of claim 1, wherein, The plurality of physical server groups are used to create a logical node of a tenant, which includes at least two physical servers in the plurality of physical server groups, a first sub-network of the first network, and a second sub-network of the second network, the network card of the at least two physical servers is used to receive data of a job task of the tenant through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers are used to store the data, and the CPU of the at least two physical servers is used to inform the computing chips of the at least two physical servers to jointly process the data through the second sub-network to complete the job task.

3. The system of claim 1 or 2, wherein, The second network includes a first switch chip corresponding to each physical server group and a second switch chip corresponding to a plurality of physical server groups; The computing chip in each physical server group accesses the first switch chip through the direct access channel, and the first switch chip is used to support the interconnection of the computing chips in each physical server group within the group; The computing chips in the plurality of physical server groups access the second switch chip through the direct access channel, and the second switch chip is used to support the interconnection of the computing chips in the plurality of physical server groups between groups.

4. The system of claim 1 or 2, wherein, The second network includes a second switch chip corresponding to a plurality of physical server groups; The computing chips in each physical server group are interconnected within the group through the direct access channel; The computing chips in the plurality of physical server groups access the second switch chip through the direct access channel, and the second switch chip is used to support the interconnection of the computing chips in the plurality of physical server groups between groups.

5. The system of any one of claims 2 to 4, wherein, When the computing chips of the at least two physical servers transmit a packet associated with the data through the second sub-network, the second sub-network is used to complete the transmission of the packet after determining that the packet carries an identifier of the second sub-network.

6. The system of any one of claims 2 to 5, wherein, The direct access channel includes a first direct access channel and a second direct access channel, when the first direct access channel is in a working state, the computing chips of the at least two physical servers access the second sub-network through the first direct access channel, and when the first direct access channel is in a fault state, the computing chips of the at least two physical servers access the second sub-network through the second direct access channel.

7. The system of claim 6, wherein, The bandwidth of the first direct-out channel is determined based on any one or any combination of the number of the first direct-out channels, a communication protocol used by the first direct-out channels, and a physical link built by the first direct-out channels.

8. The system of any one of claims 1 to 7, wherein, The direct-out channel is a line board, a cable, or an optical fiber.

9. The system according to any one of claims 1 to 8, characterized in that, The computing chip of each physical server is a graphics processing unit (GPU), a neural processing unit (NPU), a data processing unit (DPU), or a tensor processing unit (TPU).

10. The system according to any one of claims 1 to 9, characterized in that, The first network or the second network is implemented through a PCIE network, an IB network, or a CXL network.

11. A cloud service-based job task processing method, characterized by, The method is applied to a cloud management platform for managing an infrastructure for providing cloud services for tenants, the infrastructure including a plurality of physical server groups, each physical server group including a plurality of physical servers, a CPU, a memory, a network card, a disk, and a computing chip in each physical server being interconnected within the physical server, the network card of each physical server being connected to a first network, the first network being used to implement interconnection of the network cards in each physical server group within the group and between groups, the computing chip of each physical server including a direct-out channel, the computing chip of each physical server being connected to a second network through the direct-out channel, the second network being used to implement interconnection of the computing chips in each physical server group within the group and between groups; the method includes: The cloud management platform obtains a job task processing request input by a tenant, the job task processing request including data of a job task of the tenant and performance requirements of the job task; The cloud management platform creates a logical node of the tenant based on the job task processing request, the logical node including at least two idle physical servers in the plurality of physical server groups that meet the performance requirements, a first sub-network of the first network, and a second sub-network of the second network; The cloud management platform determines an operating system image input or selected by the tenant; The cloud management platform notifies the logical node to install the operating system image, in the logical node in which the operating system image is installed, the network card of the at least two physical servers is used to receive the data sent by the cloud management platform through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers are used to store the data, and the CPU of the at least two physical servers is used to notify the computing chip of the at least two physical servers to jointly process the data through the second sub-network to complete the job task.

12. The method of claim 11, wherein, The second network includes a first switch chip corresponding to each physical server group and a second switch chip corresponding to the plurality of physical server groups. The computing chip in each physical server group is connected to the first switch chip through the direct-out channel, and the first switch chip is used to support interconnection of the computing chips in each physical server group within the group. The computing chips in the plurality of physical server groups are connected to the second switch chip through the direct-out channel, and the second switch chip is used to support interconnection of the computing chips in the plurality of physical server groups between groups.

13. The method of claim 11, wherein, The second network comprises a second switch chip corresponding to a plurality of physical server groups; The computing chips in each physical server group are interconnected within the group through the direct-out channel; The computing chips in the plurality of physical server groups access the second switch chip through the direct-out channel, and the second switch chip is used to support the computing chips in the plurality of physical server groups to be interconnected between groups.

14. The method according to any one of claims 11 to 13, characterized in that, When the computing chips of the at least two physical servers transmit a message associated with the data through the second sub-network, the second sub-network is used to complete the transmission of the message after determining that the message carries the identification of the second sub-network.

15. The method according to any one of claims 11 to 14, characterized in that, The direct-out channel comprises a first direct-out channel and a second direct-out channel, when the first direct-out channel is in a working state, the computing chips of the at least two physical servers access the second sub-network through the first direct-out channel, and when the first direct-out channel is in a fault state, the computing chips of the at least two physical servers access the second sub-network through the second direct-out channel.

16. The method of claim 15, wherein, The bandwidth of the first direct-out channel is determined based on any one or any combination of the number of the first direct-out channel, the communication protocol used by the first direct-out channel, and the physical link constructed by the first direct-out channel.

17. The method according to any one of claims 11 to 16, characterized in that, The direct-out channel is a circuit board, a cable or an optical fiber.

18. The method according to any one of claims 11 to 17, characterized in that, The computing chip of each physical server is a graphics processing unit (GPU), a neural network processing unit (NPU), a data processing unit (DPU) or a tensor processing unit (TPU).

19. The method according to any one of claims 11 to 18, characterized in that, The first network or the second network is implemented through a PCIE network, an IB network or a CXL network.

20. A cloud management platform, characterized in that, The cloud management platform is used to manage an infrastructure for providing cloud services for tenants, the infrastructure comprises a plurality of physical server groups, each physical server group comprises a plurality of physical servers, the CPU, memory, network card, disk and computing chip in each physical server are interconnected within the physical server, the network card of each physical server accesses the first network, the first network is used to implement the interconnection of the network cards in each physical server group within and between groups, the computing chip of each physical server comprises a direct-out channel, and the computing chip of each physical server accesses the second network through the direct-out channel, and the second network is used to implement the interconnection of the computing chips in each physical server group within and between groups. The cloud management platform comprises: An acquisition module is configured to acquire a job task processing request input by a tenant, the job task processing request comprising data of a job task of the tenant and performance requirements of the job task; A creation module is configured to create a logical node of the tenant based on the job task processing request, the logical node comprising at least two physical servers in the plurality of physical server groups that meet the performance requirements and are idle, a first sub-network of the first network, and a second sub-network of the second network; A determination module is configured to determine an operating system image input or selected by the tenant; The notification module is configured to notify the logical node to install the operating system image, and in the logical node installed with the operating system image, the network card of the at least two physical servers is configured to receive the data sent by the cloud management platform through the first sub-network, the memory of the at least two physical servers and the disk of the at least two physical servers are configured to store the data, and the CPU of the at least two physical servers is configured to notify the computing chip of the at least two physical servers to jointly process the data through the second sub-network to complete the job task.

21. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, and each computing device includes a processor and a memory: The memory is configured to store instructions; The processor is configured to execute the instructions to cause the computing device cluster to perform the method of any one of claims 11-19.

22. A computer storage medium, comprising, The computer storage medium stores one or more instructions, which, when executed by one or more computers, cause the one or more computers to implement the method of any one of claims 11-19.

23. A computer program product, characterised in that, The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 11-19.

Citation Information

Patent Citations

  • Data processing method and device

    CN111400238A

  • Computing system and communication method

    CN117319324A

  • Method for managing a multi-tenant server enclosure

    US20200344311A1