Data transmission method, device, equipment and storage medium

By using communication cables between the processor and the switch for parallel data transmission, the problem of inefficient data transmission in multiprocessor systems is solved, and more efficient data transmission is achieved.

CN118606254BActive Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410172983.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2025-06-06
Estimated Expiration
2044-02-05

AI Technical Summary

Technical Problem

In the prior art, when data transmission between multiple processors is transmitted, data transmission cannot be realized in time due to channel bandwidth limitations, resulting in inefficiency.

Method used

By using the communication cable connection between the processor and the switch, data transmission between the processor is realized, and the bandwidth resources of the bus connection and communication cable are used in parallel, expanding the data transmission method.

Benefits of technology

It improves the data transmission efficiency between multiple processors, makes full use of the bandwidth resources of the communication cable between the processor and the switch, and avoids the waste of bandwidth resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118606254B_ABST
    Figure CN118606254B_ABST
Patent Text Reader

Abstract

The present application discloses a data transmission method, device, equipment and storage medium, which belongs to the field of communication technology. The method is executed by a first processor, the first processor, the second processor and the third processor belong to the same server, the first processor and the second processor are connected through a bus in the server, and the first processor and the third processor are connected to a switch through a communication cable. The method includes: based on the bus connection in the server, transmitting the first data between the first processor and the second processor; based on the communication cable and the switch, transmitting the second data between the first processor and the third processor; wherein the first data and the second data are transmitted in parallel. By utilizing the communication cable connection between the processor and the switch, data transmission between two processors belonging to the same server is realized, and the data transmission efficiency between multiple processors belonging to the same server is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a data transmission method, device, equipment and storage medium. Background Art

[0002] With the development of Internet technology, the amount of data that needs to be calculated is increasing, and there is a need to use multiple processors to collaboratively calculate data.

[0003] In the related art, taking a graphics processing unit (GPU) as an example, GPUs belonging to the same server transmit data through a high-speed interconnection communication channel to achieve information transfer between GPUs, thereby achieving collaborative computing of data.

[0004] However, as the amount of data that needs to be transmitted increases, data transmission cannot be achieved in a timely manner due to channel bandwidth limitations. Summary of the invention

[0005] The present application provides a data transmission method, apparatus, device and storage medium, and the technical solution is as follows:

[0006] According to one aspect of the present application, a data transmission method is provided, the method being executed by a first processor, the first processor, a second processor, and a third processor belonging to the same server, the first processor and the second processor being connected via a bus in the server, and the first processor and the third processor being connected to a switch via a communication cable, the method comprising:

[0007] transmitting first data between the first processor and the second processor based on the bus connection in the server;

[0008] Based on the communication cable and the switch, transmitting second data between the first processor and the third processor;

[0009] The first data and the second data are transmitted in parallel.

[0010] According to another aspect of the present application, a data transmission method is provided, the method being executed by a control device in a server, the first processor, the second processor, and the third processor belonging to the server, the first processor and the second processor being connected via a bus in the server, the first processor and the third processor being connected to a switch via a communication cable, the method comprising:

[0011] controlling first data to be transmitted between the first processor and the second processor, wherein the first data is transmitted based on the bus connection in the server;

[0012] Controlling second data to be transmitted between the first processor and the third processor, wherein the second data is transmitted based on the communication cable and the switch;

[0013] The first data and the second data are transmitted in parallel.

[0014] According to another aspect of the present application, a data transmission device is provided, wherein a first processor, a second processor, and a third processor belong to the same server, the first processor and the second processor are connected via a bus in the server, and the first processor and the third processor are connected to a switch via a communication cable, and the device comprises:

[0015] a transmission module, configured to transmit first data between the first processor and the second processor based on the bus connection in the server;

[0016] The transmission module is further used to transmit second data between the first processor and the third processor based on the communication cable and the switch;

[0017] The first data and the second data are transmitted in parallel.

[0018] In an optional design of the present application, the first processor and the third processor are respectively connected to a first access layer switch, and the first access layer switch is used to directly connect multiple processors; the transmission module is also used to:

[0019] The first processor transmits the second data to the third processor via a first path, wherein the first path starts from one of the first processor and the third processor and ends at the other processor, passing through the first access layer switch.

[0020] In an optional design of the present application, the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, the first aggregation layer switch is connected to the second access layer switch and the third access layer switch respectively; the transmission module is further used to:

[0021] The first processor transmits the second data to the third processor via a second path, wherein the second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence.

[0022] In an optional design of the present application, the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch respectively; the transmission module is also used for:

[0023] The first processor transmits the second data to the third processor via a third path, wherein the third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence.

[0024] In an optional design of the present application, the transmission module is further used for:

[0025] The second data is transmitted between the first processor and the third processor based on the communication cable and the network channel provided by the switch; the network channel is used to provide a data transmission path between the two processors connected by the communication cable.

[0026] In an optional design of the present application, the network channel is a channel based on a ring structure;

[0027] Among them, a transmission path based on a network channel between any two processors among the first processor, the second processor and the third processor passes through at least one switch.

[0028] In an optional design of the present application, the transmission module is further used for:

[0029] The first data is transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection; the interconnection channel is used to provide a data transmission path between at least two processors in the server.

[0030] According to another aspect of the present application, a data transmission device is provided, wherein a first processor, a second processor, and a third processor belong to a server, the first processor and the second processor are connected via a bus in the server, and the first processor and the third processor are connected to a switch via a communication cable, and the device comprises:

[0031] a control module, configured to control transmission of first data between the first processor and the second processor, wherein the first data is transmitted based on the bus connection in the server;

[0032] The control module is further used to control the transmission of second data between the first processor and the third processor, wherein the second data is transmitted based on the communication cable and the switch;

[0033] The first data and the second data are transmitted in parallel.

[0034] In an optional design of the present application, the control module is also used for:

[0035] The first data is controlled to be transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection.

[0036] In an optional design of the present application, the control module is also used for:

[0037] The second data is controlled to be transmitted between the first processor and the third processor based on the communication cable and the network channel provided by the switch.

[0038] In an optional design of the present application, the device further includes:

[0039] a processing module, configured to allocate a first channel resource bandwidth corresponding to the bus connection to at least one candidate interconnection channel, and allocate a second channel resource bandwidth corresponding to the communication cable and the switch to at least one candidate network channel, in the case where the first processor, the second processor, and the third processor are directly connected to each other;

[0040] The first data is transmitted via an interconnection channel, and the at least one candidate interconnection channel includes the interconnection channel; the second data is transmitted via a network channel, and the at least one candidate network channel includes the network channel.

[0041] In an optional design of the present application, the processing module is further used for:

[0042] When any one of the first processor, the second processor, and the third processor obtains the bus identifier of the other processor in a manner without going through an intermediate device, the first channel resource bandwidth corresponding to the bus connection is allocated to the at least one candidate interconnection channel, and the second channel resource bandwidth corresponding to the communication cable and the switch is allocated to the at least one candidate network channel.

[0043] In an optional design of the present application, the processing module is further used for:

[0044] Determining a candidate channel resource ratio between the candidate interconnect channel and the candidate network channel according to a ratio between the first channel resource bandwidth and the second channel resource bandwidth;

[0045] According to the preset number of channels and the channel resource ratio, a first number of the candidate interconnection channels and a second number of the candidate network channels are determined.

[0046] In an optional design of the present application, the processing module is further used for:

[0047] According to the data volume of the first data and the second data, a ratio of the number of transmission channels of the interconnection channel and the network channel is determined in a load balancing manner.

[0048] According to another aspect of the present application, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the data transmission method as described above.

[0049] According to another aspect of the present application, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the computer-readable storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the data transmission method as described above.

[0050] According to another aspect of the present application, a computer program product is provided, which includes computer instructions stored in a computer-readable storage medium, and a processor reads and executes the computer instructions from the computer-readable storage medium to implement the data transmission method described above.

[0051] The beneficial effects of the technical solution provided by this application include at least:

[0052] By utilizing the communication cable connection between the processor and the switch, data transmission between two processors belonging to the same server is achieved, and the bandwidth resources of the communication cable between the processor and the switch are fully utilized; the bandwidth resources of the communication cable and the bus connection are used simultaneously in a parallel manner, and a new data transmission method is expanded beyond the bus connection in the server; the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in related technologies is avoided, and the data transmission efficiency between multiple processors belonging to the same server is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0054] Figure 1 is a schematic diagram of a computer system provided by an exemplary embodiment of the present application;

[0055] Figure 2 is a schematic diagram of a processor connection relationship provided by an exemplary embodiment of the present application;

[0056] Figure 3 is a schematic diagram of an information transmission method provided by an exemplary embodiment of the present application;

[0057] Figure 4 is a flow chart of an information transmission method provided by an exemplary embodiment of the present application;

[0058] Figure 5 is a flow chart of an information transmission method provided by an exemplary embodiment of the present application;

[0059] Figure 6 is a flow chart of an information transmission method provided by an exemplary embodiment of the present application;

[0060] Figure 7 is a flow chart of an information transmission method provided by an exemplary embodiment of the present application;

[0061] Figure 8 is a schematic diagram of a connection relationship between a switch and a server provided by an exemplary embodiment of the present application;

[0062] Fig. 9 is a schematic diagram of candidate network channels and candidate interconnection channels provided by an exemplary embodiment of the present application;

[0063] Fig.10 is a schematic diagram of the performance of a single processor provided by an exemplary embodiment of the present application;

[0064] Fig.11 is a schematic diagram of the performance of a single processor provided by an exemplary embodiment of the present application;

[0065] Fig.12 is a schematic diagram of the iteration time of model training execution by a processor provided by an exemplary embodiment of the present application;

[0066] Fig.13 is a structural block diagram of a data transmission device provided by an exemplary embodiment of the present application;

[0067] Fig.14 is a structural block diagram of a data transmission device provided by an exemplary embodiment of the present application;

[0068] Fig.15 It is a structural block diagram of a server provided by an exemplary embodiment of the present application.

[0069] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application. DETAILED DESCRIPTION

[0070] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0071] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0072] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms of "a", "said" and "the" used in this disclosure and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the bus identification and other information of the processor involved in this application are all obtained with full authorization.

[0074] It should be understood that although the terms first, second, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0075] Figure 1 A schematic diagram of a computer system provided by an embodiment of the present application is shown. The computer system can be implemented as a system architecture of a data transmission method.

[0076] The server 10 uses a processor as a computing core and has data processing capabilities. This embodiment is described by taking the server 10 as including four processors (a first processor 11, a second processor 12, a third processor 13, and a fourth processor 14), but it does not exclude that the processor includes a greater or lesser number of processors; optionally, the number of processors included in a server 10 is a positive integer power of 2. Exemplarily, the server 10 also includes a control device, which controls the processor in the server to perform at least one of data transmission, data operation, and data storage through signaling.

[0077] This embodiment is described by taking the example that the processor is a graphics processing unit (GPU), but does not exclude the case where in other examples, the processor is implemented as a main processor and / or other coprocessors.

[0078] Exemplarily, the main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor is implemented as a graphics processing unit (GPU) as described above; exemplarily, the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0079] Exemplarily, the multiple processors included in the server 10 are installed on the card slots, and the card slots provide hardware connections between the multiple processors in the server 10 to achieve interconnection and communication between the multiple processors in the server 10. In one example, the GPUs of the same server 10 are connected to each other through a bus to achieve interconnection and communication. Furthermore, data is transmitted based on a high-speed interconnection communication channel (such as an Nvlink channel).

[0080] Each processor included in the server 10 is connected to at least one switch 20; the switch 20 is used to connect multiple processors to achieve data exchange and forwarding for the multiple connected processors; illustratively, the switch 20 performs exchange and forwarding on data carried by Ethernet frames based on Ethernet.

[0081] In some embodiments, the switch includes an access layer switch (Access Layer, LA), a convergence layer switch (Layer Convergence, LC), and a core layer switch (Switched Gigabit Linecard, SGLC); wherein each LA can be connected to one or more processors, and each LA can be connected to one or more LCs; the core layer switch is used to connect multiple LCs; this application is not limited to this.

[0082] Exemplarily, each processor included in the server 10 is connected to a common switch. Exemplarily, the processor can be directly connected to the common switch, or can be indirectly connected to the common switch through other switches.

[0083] In some embodiments, processors communicate with each other in a ring-based manner, and communication specifically refers to data (traffic) transmission between processors. The data transmission paths between processors are connected end to end to form a ring, which is also called a communication traffic ring. Exemplarily, for any processor on the ring structure, there is a left-neighboring processor and a right-neighboring processor; the above processor receives data from the left-neighboring processor and sends data to the right-neighboring processor.

[0084] like Figure 2 As shown in the figure, taking the communication between four processors (first processor 11, second processor 12, third processor 13, fourth processor 14) based on the ring structure as an example, the above four processors communicate based on the Ring-based method, and the traffic interaction will be connected end to end to form a ring. For example, for processor 11, it receives data from its left neighbor processor, that is, processor 14; and sends data to its right neighbor processor, that is, processor 12. For example, in Figure 2The storage space of each processor is divided into four data blocks (chunks); dividing the processor into chunks can increase the parallelism of communication. Each processor processes the corresponding data block, and different processors can perform communication operations in parallel, thereby improving the overall throughput.

[0085] Figure 3 A schematic diagram of an information transmission method provided by an embodiment of the present application is shown.

[0086] In this embodiment, the server 10 includes four GPUs (GPU0, GPU1, GPU2, GPU3) for example. It can be understood that in different examples, the server 10 may include more or fewer GPUs.

[0087] The above four GPUs are all connected to an access layer switch; specifically, LA0 is connected to GPU0 and GPU1 via communication cables, LA1 is connected to GPU1 via communication cables, LA2 is connected to GPU2 via communication cables, and LA3 is connected to GPU3 via communication cables.

[0088] The above four access layer switches are all connected to the aggregation layer switch; specifically, LC0 is connected to LA0 via a communication cable, LC1 is connected to LA1 and LA2 via a communication cable, LC2 is connected to LA2 via a communication cable, and LC3 is connected to LA3 via a communication cable. The above four aggregation layer switches are all connected to the core layer switch via communication cables. The four GPUs included in the server 10 have a one-to-one corresponding network interface controller (Network Interface Card, NIC; also known as a network adapter or network card) for providing a connection interface between the GPU and the switch.

[0089] Based on the bus connection within the server 10, the first data is transmitted between GPU1 and GPU2; exemplarily, the bus connection provides an interconnection channel, which can be specifically implemented as a high-speed interconnection communication channel (such as an Nvlink channel) to realize data transmission between GPU1 and GPU2.

[0090] Based on the communication cable and the switch, the second data is transmitted between GPU1 and GPU0; illustratively, the communication cable and the switch provide a network channel. Further, the switch and the communication cable implement communication of the network channel based on Ethernet.

[0091] For example, transmitting the first data between GPU1 and GPU2 can be realized as GPU1 transmitting the first data to GPU2, or as GPU1 receiving the first data sent by GPU2 (i.e., GPU2 transmitting the first data to GPU1). Similarly, transmitting the second data between GPU1 and GPU0 can be realized as GPU1 sending or receiving the second data.

[0092] The first data and the second data are transmitted in parallel, so as to realize the simultaneous transmission of the first data and the second data at the same time. The bandwidth resources of the interconnection channel and the network channel are used in parallel, and a new data transmission channel is expanded in addition to the interconnection channel in the server; the waste of bandwidth resources caused by using the network channel only for data transmission across servers in the related technology is avoided, and the data transmission efficiency between multiple processors belonging to the same server is improved.

[0093] Figure 4 A flowchart of an information transmission method provided by an exemplary embodiment of the present application is shown. The method can be executed by a first processor. The method includes:

[0094] Step 510: Transmitting first data between the first processor and the second processor based on a bus connection in the server;

[0095] The first processor, the second processor, and the third processor belong to the same server. In one example, the server is provided with card slots for the processors belonging to the server, or are integrated on a mainboard in the server. The first processor and the second processor are connected via a bus in the server.

[0096] The transmission of the first data is implemented by the first processor, which may be implemented as the first processor sending the first data to the second processor, or may be implemented as the first processor receiving the first data sent by the second processor.

[0097] Exemplarily, the transmission of the first data may be exclusive to the bandwidth of the bus connection, or the data may be transmitted in parallel within the bus connection to share the bandwidth of the bus connection.

[0098] As introduced above, the processor in the present application can be implemented as different types of processors corresponding to different functions. In the following embodiments, the processor is implemented as a GPU as an example, but the situation of being implemented as other types of processors is not excluded.

[0099] Exemplarily, the first data is operation data that needs to be calculated in the process of the first processor and the second processor collaboratively performing data processing. For example, the server is used to perform training of an artificial intelligence neural network, and the first data is the calculation result data of the first GPU. The first GPU sends the first data to the second GPU, and the second GPU uses the first data as an input parameter in the above-mentioned training process to enable the first GPU and the second GPU in the server to collaboratively perform training of the artificial intelligence neural network.

[0100] Step 520: Transmitting second data between the first processor and the third processor based on the communication cable and the switch;

[0101] The first processor and the third processor are connected to the switch via a communication cable. The switch is used to provide data forwarding and forward messages or data packets carrying data to the connected processors.

[0102] The transmission of the second data by the first processor may be implemented as the first processor sending the second data to the third processor, or may be implemented as the first processor receiving the second data sent by the third processor. Similar to the transmission of the first data, the transmission of the second data may exclusively use the bandwidth of the communication cable, or may transmit data in parallel within the communication cable to share the bandwidth of the communication cable.

[0103] The first data and the second data are transmitted in parallel. In addition to the bus connection in the server, a new way of transmitting data is expanded, the bandwidth of data transmission in the server is increased, and the data transmission efficiency is improved; the bandwidth resources of the communication cable connection between the processor and the switch are fully utilized.

[0104] It should be noted that, in various embodiments of the present application, the third processor is usually a processor different from the second processor, but it does not exclude the situation that the second processor and the third processor are the same processor.

[0105] In summary, the method provided in this embodiment realizes data transmission between two processors belonging to the same server by utilizing the communication cable connection between the processor and the switch, thereby making full use of the bandwidth resources of the communication cable between the processor and the switch; simultaneously using the bandwidth resources of the communication cable and the bus connection in a parallel manner, expanding a new data transmission method beyond the bus connection within the server; avoiding the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, thereby improving the data transmission efficiency between multiple processors belonging to the same server.

[0106] Next, the manner in which the first processor and the third processor are connected to the switch through the communication cable is further introduced.

[0107] In one implementation, the first processor and the third processor are respectively connected to the first access layer switch, and the first access layer switch is used to directly connect multiple processors. It can be understood that the first access layer switch is usually connected to a larger number of processors. The processors connected to the first access layer switch may only include other processors in the server to which the first processor, the second processor, and the third processor belong, or may include other processors belonging to different servers.

[0108] Accordingly, the transmission mode of the second data can be specifically implemented as follows:

[0109] The first processor transmits second data to the third processor via a first path. The first path starts from one of the first processor and the third processor and ends at the other processor, passing through the first access layer switch.

[0110] Further, the first processor sends a first message, and the first message is forwarded to the second processor via the first access layer switch;

[0111] Exemplarily, the sender address of the first message is the address of the first processor, the destination address is the address of the third processor, and the first message carries the second data.

[0112] As described above, the second data may be sent by the first processor to the third processor, or may be sent by the third processor to the first processor, and accordingly:

[0113] And / or, the first processor receives a second message, where the second message is sent by the second processor and forwarded to the first processor via the first access layer switch;

[0114] Exemplarily, the sender address of the second message is the address of the third processor, the destination address is the address of the first processor, and the second message carries the second data.

[0115] refer to Figure 3 In one example, the second data is sent from GPU1 to GPU0. Specifically, GPU1 sends the second data to GPU0 via the following path: starting from GPU1, passing through LA0, and ending at GPU0.

[0116] In summary, the method provided in this embodiment realizes data transmission between two processors belonging to the same server by utilizing the communication cable connection between the processor and the switch, thereby making full use of the bandwidth resources of the communication cable between the processor and the first access layer switch; using the bandwidth resources of the communication cable and the bus connection in parallel at the same time, in addition to the bus connection in the server, a new data transmission method is expanded for the two processors that are commonly connected to the first access layer switch; avoiding the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, thereby improving the data transmission efficiency between multiple processors belonging to the same server.

[0117] In another implementation, the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, and the first aggregation layer switch is connected to the second access layer switch and the third access layer switch, respectively. It can be understood that the second access layer switch and the third access layer switch are similar to the first access layer switch mentioned above, and are usually connected to a larger number of processors. The first aggregation layer switch is usually connected to a larger number of access layer switches. Exemplarily, the access layer switch can be connected to one or more aggregation layer switches.

[0118] Accordingly, the transmission mode of the second data can be specifically implemented as follows:

[0119] The first processor transmits second data to the third processor via a second path. The second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence.

[0120] Further, the first processor sends a third message, and the third message passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence, and is forwarded to the second processor one by one;

[0121] Exemplarily, the sender address of the third message is the address of the first processor, the destination address is the address of the third processor, and the third message carries the second data.

[0122] As described above, the third processor may send the information to the first processor, and accordingly:

[0123] And / or, the first processor receives a fourth message, which is sent by the second processor and is sequentially forwarded to the first processor through the third access layer switch, the first aggregation layer switch, and the second access layer switch;

[0124] Exemplarily, the sender address of the fourth message is the address of the third processor, the destination address is the address of the first processor, and the fourth message carries the second data.

[0125] refer to Figure 3 In one example, the second data is sent from GPU2 to GPU1. Specifically, GPU2 sends the second data to GPU1 via the following path: starting from GPU2, passing through LA2, LC1, and LA1 in sequence, and ending at GPU1.

[0126] To summarize, the method provided in this embodiment realizes data transmission between two processors belonging to the same server by utilizing the communication cable connection between the processor and the switch, and fully utilizes the bandwidth resources of the communication cables between the processor and the second access layer switch, the first aggregation layer switch, and the third access layer switch; uses the bandwidth resources of the communication cables and bus connections in parallel, and expands a new data transmission method for the two processors that are commonly connected to the first aggregation layer switch in addition to the bus connection in the server; avoids the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, and improves the data transmission efficiency between multiple processors belonging to the same server.

[0127] In another implementation, the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch respectively; it can be understood that the fourth access layer switch and the fifth access layer switch are similar to the first access layer switch mentioned above, and are usually connected to a larger number of processors. The second aggregation layer switch and the third aggregation layer switch are usually connected to a larger number of access layer switches. Exemplarily, the access layer switch can be connected to one or more aggregation layer switches.

[0128] Accordingly, the transmission mode of the second data can be specifically implemented as follows:

[0129] The first processor transmits the second data to the third processor via a third path. The third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence.

[0130] Further, the first processor sends a fifth message, and the fifth message is sequentially forwarded to the second processor through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch;

[0131] Exemplarily, the sender address of the fifth message is the address of the first processor, the destination address is the address of the third processor, and the fifth message carries the second data.

[0132] As described above, the third processor may send the information to the first processor, and accordingly:

[0133] And / or, the first processor receives a sixth message, which is sent by the second processor and is sequentially forwarded to the first processor through the fifth access layer switch, the third aggregation layer switch, the core layer switch, the second aggregation layer switch, and the fourth access layer switch;

[0134] Exemplarily, the sender address of the sixth message is the address of the third processor, the destination address is the address of the first processor, and the sixth message carries the second data.

[0135] refer to Figure 3 In one example, the second data is sent from GPU3 to GPU2. Specifically, GPU3 sends the second data to GPU2 via the following path: starting from GPU3, passing through LA3, LC3, SGLC, LC2, LA2 in sequence, and ending at GPU2.

[0136] To summarize, the method provided in this embodiment realizes data transmission between two processors belonging to the same server by utilizing the communication cable connection between the processor and the switch, and fully utilizes the bandwidth resources of the communication cables between the processor and the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch; uses the bandwidth resources of the communication cables and the bus connections in parallel, and expands a new data transmission method for the two processors that are commonly connected to the core layer switch in addition to the bus connection in the server; avoids the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, and improves the data transmission efficiency between multiple processors belonging to the same server.

[0137] Figure 5 FIG. 1 is a flowchart of an information transmission method provided by an exemplary embodiment of the present application. The method may be executed by a first processor. Figure 4 In the illustrated embodiment, step 510 may be implemented as step 512, and step 520 may be implemented as step 522:

[0138] Step 512: Transmitting first data between the first processor and the second processor based on the interconnection channel provided by the bus connection;

[0139] The interconnection channel is used to provide a data transmission path between at least two processors in the server. Further, the interconnection channel is also used to provide a data transmission path between any two servers in the server.

[0140] In one example, the data transmission path provided by the interconnection channel may be a point-to-point serial path, such as a Peripheral Component Interconnect Express (PCIE), or may be a Figure 2 The data transmission path for communication based on a ring structure (Ring-based) is introduced; taking the first processor sending the first data to the second processor as an example, if the second processor is the right neighbor processor of the first processor, the first data is directly transmitted between the first processor and the second processor. If the second processor is the left neighbor processor of the first processor, the first data sent by the first processor is forwarded to the second processor by at least one processor in the server.

[0141] In one example, the first processor and the second processor are image processors, and the interconnection channel is a high-speed interconnection communication channel corresponding to the image processor (such as an Nvlink channel); further, the high-speed interconnection communication channels of image processors from different manufacturers have different names, all of which are data transmission paths used to provide data transmission between image processors in the same server.

[0142] Step 522: Transmitting second data between the first processor and the third processor based on the network channel provided by the communication cable and the switch;

[0143] The network channel is used to provide a data transmission path between two processors connected by a communication cable. Furthermore, there is a commonly connected switch between any two processors belonging to the same server, and the commonly connected switch is an access layer switch directly connected to the processor, or an aggregation layer switch or a core layer switch indirectly connected through an access layer switch.

[0144] Exemplarily, any two processors among the first processor, the second processor and the third processor are connected via a network channel provided by a communication cable and a switch, so that data can be transmitted between any two processors.

[0145] In an optional implementation, the network channel is a channel based on a ring structure; taking the first processor sending the second data to the third processor as an example, if the third processor is the right neighbor of the first processor, the second data is directly transmitted between the first processor and the third processor. If the third processor is the left neighbor of the first processor, the second data sent by the first processor is forwarded to the third processor by at least one processor and at least one switch in the server. Exemplarily, in a network channel based on a ring structure, each processor only needs to process the corresponding data, and different processors can send or receive data at the same time and transmit data in parallel to achieve shared communication cable and switch bandwidth.

[0146] It should be noted that this embodiment only introduces the situation where step 512 and step 522 are executed simultaneously, but does not exclude the situation where the above two steps are executed separately, such as combining step 512 with step 520 in the above text, or combining step 522 with step 510 in the above text, to form a new embodiment that is implemented separately, and this application does not limit this.

[0147] In summary, the method provided in this embodiment realizes data transmission between two processors belonging to the same server by utilizing the network channel provided by the communication cable connection between the processor and the switch, thereby making full use of the bandwidth resources of the communication cable between the processor and the switch; simultaneously using the bandwidth resources of the communication cable and the bus connection in a parallel manner, expanding the data transmission mode of the network channel in addition to the interconnection channel provided by the bus connection in the server; avoiding the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, thereby improving the data transmission efficiency between multiple processors belonging to the same server.

[0148] Figure 6 A flowchart of an information transmission method provided by an exemplary embodiment of the present application is shown. The method can be executed by a control device. The method includes:

[0149] Step 550: Control the transmission of first data between the first processor and the second processor;

[0150] The first data is transmitted based on a bus connection in the server; the first processor, the second processor and the third processor belong to the same server. The first processor and the second processor are connected via a bus in the server.

[0151] Exemplarily, the control device usually instructs the first processor and / or the second processor to transmit the first data based on the control signaling. The first data may be sent from the first processor to the second processor, or from the second processor to the first processor. Exemplarily, the control signaling may be sent to the first processor and / or the second processor based on a bus connection (such as a PCIE channel provided by the bus connection), and it is not excluded that the control signaling is transmitted through shared memory or the like.

[0152] In an optional implementation manner, this step may be implemented as: controlling the first data to be transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection.

[0153] The interconnection channel is used to provide a data transmission path between at least two processors in the server. Further, the interconnection channel is also used to provide a data transmission path between any two servers in the server.

[0154] The control device controls the transmission of the first data by sending an instruction signaling for calling the interconnection channel to the first processor and / or the second processor.

[0155] Step 560: Control the transmission of the second data between the first processor and the third processor;

[0156] The second data is transmitted based on the communication cable and the switch; the first processor and the third processor are connected to the switch via the communication cable, and the switch is used to provide data forwarding, and forwards the message or data packet carrying the data to the connected processor. For example, the specific method of transmission between the first processor and the third processor can refer to the above Figures 3 to 5 The introduction in will not be repeated here one by one.

[0157] Similar to controlling the transmission of the first data, the control device usually instructs the first processor and / or the third processor to transmit the second data based on the control signaling. The second data can be sent from the first processor to the third processor, or from the third processor to the first processor.

[0158] In an optional implementation, this step may be implemented as: controlling the second data to be transmitted between the first processor and the third processor based on a network channel provided by a communication cable and a switch.

[0159] The network channel is used to provide a data transmission path between two processors connected by a communication cable. Furthermore, there is a commonly connected switch between any two processors belonging to the same server. There is a network channel provided by the communication cable and the switch between any two servers in the server.

[0160] The first data and the second data are transmitted in parallel. A control signal is sent to the first processor, or at least two processors among the first processor, the second processor, and the third processor through a control device in the server to control the parallel transmission of the first data and the second data. In addition to the bus connection in the server, a new way of transmitting data is expanded, the bandwidth of data transmission in the server is increased, and the data transmission efficiency is improved; the bandwidth resources of the communication cable connection between the processor and the switch are fully utilized.

[0161] In summary, the method provided in this embodiment realizes the control of data transmission between two processors belonging to the same server by utilizing the communication cable connection between the processor and the switch, and fully utilizes the bandwidth resources of the communication cable between the processor and the switch; simultaneously uses the bandwidth resources of the communication cable and the bus connection in a parallel manner, and expands a new data transmission method beyond the bus connection within the server; avoids the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, and improves the data transmission efficiency between multiple processors belonging to the same server.

[0162] Figure 7 FIG. 1 is a flowchart of an information transmission method provided by an exemplary embodiment of the present application. The method can be executed by a control device. Figure 6 Based on the embodiment shown, the following steps are also included:

[0163] Step 545: Determine a candidate channel resource ratio between the candidate interconnect channel and the candidate network channel according to a ratio between the first channel resource bandwidth and the second channel resource bandwidth;

[0164] Exemplarily, the first channel resource bandwidth is the channel resource bandwidth corresponding to the bus connection, and the second channel resource bandwidth is the channel resource bandwidth corresponding to the communication cable and the switch. In order to ensure the balance of data transmission between processors in the server, the candidate interconnection channel constructed based on the bus connection and the candidate network channel constructed based on the communication cable and the switch are the same or similar.

[0165] Exemplarily, the first channel resource bandwidth is the maximum bandwidth supported by the bus connection, and the first channel resource bandwidth is evenly divided among the plurality of candidate interconnection channels to ensure that the data transmission rate among the candidate interconnection channels is the same. Similarly, the second channel resource bandwidth is evenly divided among the plurality of candidate network channels to ensure that the data transmission rate among the candidate network channels is the same.

[0166] When the candidate channel resource ratio is equal to the ratio between the first channel resource bandwidth and the second channel resource bandwidth, and the channel resource bandwidth is equally divided as described above, the data transmission rates between the candidate interconnection channels and the candidate network channels can be guaranteed to be the same.

[0167] Step 546: Determine a first number of candidate interconnection channels and a second number of candidate network channels according to a preset number of channels and a channel resource ratio;

[0168] The value of the preset number of channels is greater than 1, and is used to indicate the total number of candidate interconnection channels and candidate network channels. Exemplarily, the preset number of channels can be the maximum number of parallel data transmissions supported by the processor hardware, or can be an empirical value set manually. In a preferred example, the preset number of channels is a positive integer power of 2.

[0169] The values ​​of the first quantity and the second quantity are positive integers greater than or equal to 1.

[0170] The value of the channel resource ratio is added by 1 as the divisor; the preset number of channels is used as the dividend; and a calculated quotient is obtained. If the calculated quotient is an integer, the value of the second number is equal to the calculated quotient;

[0171] If the decimal part of the calculated quotient is not zero, the second number is an integer obtained by rounding up, rounding down, or rounding off the calculated quotient. The first number is the difference between the preset channel number and the second number.

[0172] By constraining the determination of the first number and the second number by the channel resource ratio, on the premise of determining the first number and the second number as integers, it is further ensured that the data transmission rates between the candidate interconnection channels and the candidate network channels are similar.

[0173] In a specific example, the maximum bandwidth supported by the bus connection is 200 GB / s, which is recorded as Bw_for_Nvlink. The maximum bandwidth supported by the communication cable and the switch is 50 GB / s, which is recorded as Bw_for_Net. The candidate channel resource ratio is 4.

[0174] The value of the channel resource ratio (valued as 4) is added by 1 as the divisor; the preset number of channels (for example, 16, recorded as M) is used as the dividend (i.e., calculating M*Bw_for_Net / (Bw_for_Net+Bw_for_Nvlink)), and the calculated quotient is 16 / 5=3.2. The calculated quotient is rounded down to obtain the value of the second value of 3, that is, the number of candidate network channels is 3, the number of candidate interconnection channels is 13, and 16 data transmission channels are constructed in a heterogeneous manner.

[0175] Step 547: In the case where the first processor, the second processor, and the third processor are directly connected to each other, a first channel resource bandwidth corresponding to the bus connection is allocated to at least one candidate interconnection channel, and a second channel resource bandwidth corresponding to the communication cable and the switch is allocated to at least one candidate network channel;

[0176] Exemplarily, allocating a first channel resource bandwidth corresponding to a bus connection to at least one candidate interconnection channel is usually done by equally dividing the first channel resource bandwidth corresponding to the bus connection to a first number of interconnection channels, but other allocation methods are not excluded. Exemplarily, allocating the first channel resource bandwidth includes allocating data transmission resources in the time domain or frequency domain. In order to facilitate encoding and decoding of data transmission, the channel transmission resources of a candidate interconnection channel are usually continuous in the time domain or frequency domain. However, it is not excluded that in order to ensure normal data processing of the processor in the server, multiple transmission resources that are discontinuous in the time domain or frequency domain are allocated to the candidate interconnection channel.

[0177] Exemplarily, when the first processor, the second processor, and the third processor are directly connected to each other, the three processors belong to the same server, and the processors use the tensor parallelism (TP) strategy to perform distributed computing to achieve parallel execution of data processing. This avoids the situation where there is no need to call the interconnection channel to transmit data under the data parallelism (DP) and pipeline parallelism (PP) strategies.

[0178] In an optional implementation, step 547 may be implemented as follows: when any one of the first processor, the second processor, and the third processor obtains the bus identifier of other processors without going through an intermediate device, a first channel resource bandwidth corresponding to the bus connection is allocated to at least one candidate interconnection channel, and a second channel resource bandwidth corresponding to the communication cable and the switch is allocated to at least one candidate network channel.

[0179] Exemplarily, a bus ID is a global ID of a processor in a server. Obtaining the bus ID of another processor without going through an intermediate device is also called direct reachability of the Bus ID; accordingly, the processors communicate directly through the bus without being transferred through additional devices or interfaces. Exemplarily, the bus ID is obtained by calling a Compute Unified Device Architecture (CUDA) driver function.

[0180] In this embodiment, the first data is transmitted through an interconnection channel, and at least one candidate interconnection channel includes an interconnection channel. The second data is transmitted through a network channel, and at least one candidate network channel includes a network channel. For an introduction to the interconnection channel and the network channel, please refer to step 512, step 522, step 552, and step 562 above; they will not be repeated here one by one.

[0181] Step 548: Determine the transmission channel resource ratio of the interconnection channel and the network channel in a load balancing manner according to the data volume of the first data and the second data;

[0182] Exemplarily, the number of interconnection channels and network channels called during the transmission of the first data and the second data may be one or more. Exemplarily, the first data and the second data are split to ensure that the amount of data transmitted on each channel is the same or similar, thereby achieving simultaneous completion of data transmission to improve the parallel transmission efficiency of data.

[0183] It should be noted that this embodiment only describes the case where steps 545 to 548 are all executed, but it does not exclude the case where the above steps are executed separately, such as steps 547 and Figure 6 or steps 545 to 547 and Figure 6 The step 550 and the step 560 in the embodiment are combined to form a new embodiment which is implemented separately, and the present application does not limit this.

[0184] In summary, the method provided in this embodiment realizes the control of data transmission between two processors belonging to the same server by utilizing the communication cable connection between the processor and the switch, and fully utilizes the bandwidth resources of the communication cable between the processor and the switch; simultaneously uses the bandwidth resources of the communication cable and the bus connection in a parallel manner, and expands a new data transmission method beyond the bus connection within the server; avoids the waste of bandwidth resources caused by using switches and communication cables only for data transmission across servers in the related technology, and improves the data transmission efficiency between multiple processors belonging to the same server.

[0185] Figure 8 A schematic diagram showing the connection relationship between a switch and a server provided by an exemplary embodiment of the present application.

[0186] In this embodiment, the server 10 includes eight GPUs (GPU0 to GPU7) for example. The eight GPUs are all connected to access layer switches, totaling eight access layer switches; each GPU is connected to an access layer switch with the same serial number, for example: LA0 is connected to GPU1 through a communication cable; and Figure 3 Similarly, the eight GPUs included in server 10 have one-to-one corresponding NICs.

[0187] It should be noted that LA0 is also connected to GPU1 through a communication cable, LA2 is also connected to GPU3 through a communication cable, LA4 is also connected to GPU5 through a communication cable, and LA6 is also connected to GPU7 through a communication cable.

[0188] Exemplarily, the eight access layer switches are connected to the aggregation layer switches, totaling eight aggregation layer switches; each access layer switch is connected to the aggregation layer switch with the same sequence number, for example: LC0 is connected to LA0 via a communication cable. Exemplarily, the eight aggregation layer switches are all connected to the core layer switch.

[0189] In this embodiment, the candidate network channel is a channel based on a ring structure.

[0190] Specifically: GPU0 communicates with GPU1 through LA0.

[0191] GPU1 communicates with GPU2 through LA1, LC1, SGLC, LC2, LA2.

[0192] GPU2 communicates with GPU3 through LA2.

[0193] GPU3 communicates through LA3, LC3, SGLC, LC4, LA4 and GPU4.

[0194] GPU4 communicates with GPU5 through LA4.

[0195] GPU5 communicates through LA5, LC5, SGLC, LC6, LA6 and GPU6.

[0196] GPU6 communicates with GPU7 through LA6.

[0197] GPU7 communicates through LA7, LC7, SGLC, LC0, LA0 and GPU0.

[0198] In an example, the number of candidate network channels is 3, the number of candidate interconnection channels is 13, and 16 data transmission channels are constructed in a heterogeneous manner; the above 16 data transmission channels transmit data in a parallel manner.

[0199] Fig. 9 A schematic diagram of candidate network channels and candidate interconnection channels provided by an exemplary embodiment of the present application is shown.

[0200] Among them, the first candidate interconnection channel 601, the second candidate interconnection channel 602 to the Nth candidate interconnection channel 603 are interconnection channels provided by bus connections; as described above, the candidate interconnection channels are channels based on a ring structure. In the above example, the number of candidate interconnection channels is 13, that is, the value of N is 13. Each candidate interconnection channel is used to transmit a data block, such as transmitting the first data block 611 based on the first candidate interconnection channel 601. It should be noted that the figure shows that the first data block 611 passes through all processors in GPU0 to GPU7, but in some examples, the first data block can pass through a smaller number of GPUs, such as transmitting the first data block 611 only between GPU0 and GPU1. Exemplarily, the data blocks transmitted on different candidate interconnection channels are different, which can realize the parallel transmission of multiple data blocks without interfering with each other.

[0201] Similarly, the first candidate network channel 604 to the Mth candidate network channel 605 are network channels provided by communication cables and switches; as described above, the candidate network channels are channels based on a ring structure. In the above example, the number of candidate network channels is 3, that is, the value of M is 3. Each candidate network channel is used to transmit a data block, such as transmitting the N+1th data block 614 based on the first candidate network channel 604. Exemplarily, the data blocks transmitted on different candidate network channels are different.

[0202] Constructing N+M data transmission channels in a heterogeneous manner can achieve parallel transmission of up to N+M data blocks without interfering with each other.

[0203] Fig.10 A schematic diagram showing the performance of a single processor provided by an exemplary embodiment of the present application is shown.

[0204] Used to indicate the data transfer rate when the processor uses All Reduce Operation (AllReduce) to perform distributed computing.

[0205] The first data line 701 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 0 and the number of candidate interconnection channels is 16;

[0206] The second data line 702 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 1 and the number of candidate interconnect channels is 15;

[0207] The third data line 703 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 2 and the number of candidate interconnection channels is 14;

[0208] The fourth data line 704 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 3 and the number of candidate interconnection channels is 13;

[0209] The fifth data line 705 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 4 and the number of candidate interconnection channels is 12;

[0210] The sixth data line 706 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 5 and the number of candidate interconnection channels is 11;

[0211] The seventh data line 707 is used to indicate the theoretical upper limit of the data transmission rate of the processor when the candidate network channel and the candidate interconnection channel transmit data in parallel.

[0212] When the number of candidate network channels is 3 and the number of candidate interconnect channels is 13, the optimal data transmission rate of the processor is achieved, and the performance is improved by more than 25% compared with the non-heterogeneous data transmission mode provided by the first data line 701 including only candidate interconnect channels.

[0213] Fig.11 A schematic diagram of the performance of a single processor provided by an exemplary embodiment of the present application is shown, which is used to indicate the data transmission rate when the processor uses data exchange (such as All to All) between all participants to perform distributed computing.

[0214] The eighth data line 708 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 0 and the number of candidate interconnection channels is 16;

[0215] The ninth data line 709 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 1 and the number of candidate interconnection channels is 15;

[0216] The tenth data line 710 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 2 and the number of candidate interconnect channels is 14;

[0217] The eleventh data line 711 is used to indicate the data transmission rate of the processor when the number of candidate network channels is 3 and the number of candidate interconnection channels is 13;

[0218] When the number of candidate network channels is 3 and the number of candidate interconnect channels is 13, the optimal data transmission rate of the processor is achieved, and the performance is improved by more than 20% compared with the non-heterogeneous data transmission method provided by the eighth data line 708 which only includes candidate interconnect channels.

[0219] Fig.12A schematic diagram of the processor execution model training iteration time provided by an exemplary embodiment of the present application is shown, which is used to indicate the training time consumed by the processor under different numbers of iterations.

[0220] The twelfth data line 712 is used to indicate the processor execution model training iteration time when the number of candidate network channels is 0 and the number of candidate interconnection channels is 16;

[0221] The thirteenth data line 713 is used to indicate the processor execution model training iteration time when the number of candidate network channels is 3 and the number of candidate interconnection channels is 13;

[0222] When training large models, data is transmitted in a heterogeneous manner. Compared with the non-heterogeneous data transmission method that only includes candidate interconnect channels, the tensor parallel AllReduce bandwidth is increased by 10% compared to when heterogeneous communication is not enabled, and the model iteration time is reduced by 2.2% compared to when heterogeneous communication is not enabled.

[0223] Those skilled in the art can understand that the above embodiments can be implemented independently, or the above embodiments can be freely combined to form new embodiments to implement the data transmission method of the present application.

[0224] Fig.13 A structural block diagram of a data transmission device provided by an exemplary embodiment of the present application is shown. The first processor, the second processor, and the third processor belong to the same server, the first processor and the second processor are connected via a bus in the server, the first processor and the third processor are connected to a switch via a communication cable, and the device includes:

[0225] A transmission module 810, configured to transmit first data between the first processor and the second processor based on the bus connection in the server;

[0226] The transmission module 810 is further configured to transmit second data between the first processor and the third processor based on the communication cable and the switch;

[0227] The first data and the second data are transmitted in parallel.

[0228] In an optional design of the present application, the first processor and the third processor are respectively connected to a first access layer switch, and the first access layer switch is used to directly connect multiple processors; the transmission module 810 is also used to:

[0229] The first processor transmits the second data to the third processor via a first path, wherein the first path starts from one of the first processor and the third processor and ends at the other processor, passing through the first access layer switch.

[0230] In an optional design of the present application, the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, the first aggregation layer switch is connected to the second access layer switch and the third access layer switch respectively; the transmission module 810 is further used to:

[0231] The first processor transmits the second data to the third processor via a second path, wherein the second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence.

[0232] In an optional design of the present application, the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch respectively; the transmission module 810 is further used to:

[0233] The first processor transmits the second data to the third processor via a third path, wherein the third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence.

[0234] In an optional design of the present application, the transmission module 810 is further used for:

[0235] The second data is transmitted between the first processor and the third processor based on the communication cable and the network channel provided by the switch; the network channel is used to provide a data transmission path between the two processors connected by the communication cable.

[0236] In an optional design of the present application, the network channel is a channel based on a ring structure;

[0237] Among them, a transmission path based on a network channel between any two processors among the first processor, the second processor and the third processor passes through at least one switch.

[0238] In an optional design of the present application, the transmission module 810 is further used for:

[0239] The first data is transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection; the interconnection channel is used to provide a data transmission path between at least two processors in the server.

[0240] Fig.14 A structural block diagram of a data transmission device provided by an exemplary embodiment of the present application is shown. The first processor, the second processor, and the third processor belong to a server, the first processor and the second processor are connected via a bus in the server, the first processor and the third processor are connected to a switch via a communication cable, and the device includes:

[0241] A control module 820, configured to control transmission of first data between the first processor and the second processor, wherein the first data is transmitted based on the bus connection in the server;

[0242] The control module 820 is further used to control the transmission of second data between the first processor and the third processor, wherein the second data is transmitted based on the communication cable and the switch;

[0243] The first data and the second data are transmitted in parallel.

[0244] In an optional design of the present application, the control module 820 is further used for:

[0245] The first data is controlled to be transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection.

[0246] In an optional design of the present application, the control module 820 is further used for:

[0247] The second data is controlled to be transmitted between the first processor and the third processor based on the communication cable and the network channel provided by the switch.

[0248] In an optional design of the present application, the device further includes:

[0249] The processing module 830 is configured to allocate a first channel resource bandwidth corresponding to the bus connection to at least one candidate interconnection channel, and allocate a second channel resource bandwidth corresponding to the communication cable and the switch to at least one candidate network channel, in the case where the first processor, the second processor, and the third processor are directly connected to each other.

[0250] The first data is transmitted via an interconnection channel, and the at least one candidate interconnection channel includes the interconnection channel; the second data is transmitted via a network channel, and the at least one candidate network channel includes the network channel.

[0251] In an optional design of the present application, the processing module 830 is further used for:

[0252] When any one of the first processor, the second processor, and the third processor obtains the bus identifier of the other processor in a manner without going through an intermediate device, the first channel resource bandwidth corresponding to the bus connection is allocated to the at least one candidate interconnection channel, and the second channel resource bandwidth corresponding to the communication cable and the switch is allocated to the at least one candidate network channel.

[0253] In an optional design of the present application, the processing module 830 is further used for:

[0254] Determining a candidate channel resource ratio between the candidate interconnect channel and the candidate network channel according to a ratio between the first channel resource bandwidth and the second channel resource bandwidth;

[0255] According to the preset number of channels and the channel resource ratio, a first number of the candidate interconnection channels and a second number of the candidate network channels are determined.

[0256] In an optional design of the present application, the processing module 830 is further used for:

[0257] According to the data volume of the first data and the second data, a ratio of the number of transmission channels of the interconnection channel and the network channel is determined in a load balancing manner.

[0258] One point that needs to be explained is that the device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example to implement its functions. In actual applications, the above-mentioned functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0259] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method; the technical effects achieved by each module performing operations are the same as those in the embodiment of the method, and will not be elaborated here.

[0260] An embodiment of the present application further provides a computer device, which includes: a processor and a memory, wherein a computer program is stored in the memory; the processor is used to execute the computer program in the memory to implement the data transmission method provided by the above-mentioned method embodiments.

[0261] Optionally, the computer device is a server. Fig.15 It is a structural block diagram of a server provided by an exemplary embodiment of the present application.

[0262] Typically, the server 2300 includes: a processor 2301 and a memory 2302 .

[0263] The processor 2301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 2301 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 2301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 2301 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 2301 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0264] The memory 2302 may include one or more computer-readable storage media, which may be non-transitory. The memory 2302 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 2302 is used to store at least one instruction, which is used to be executed by the processor 2301 to implement the data transmission method provided in the method embodiment of the present application.

[0265] In some embodiments, the server 2300 may also optionally include: an input interface 2303 and an output interface 2304. The processor 2301, the memory 2302, the input interface 2303, and the output interface 2304 may be connected via a bus or a signal line. Each peripheral device may be connected to the input interface 2303 and the output interface 2304 via a bus, a signal line, or a circuit board. The input interface 2303 and the output interface 2304 may be used to connect at least one peripheral device related to input / output (I / O) to the processor 2301 and the memory 2302. In some embodiments, the processor 2301, the memory 2302, the input interface 2303, and the output interface 2304 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 2301, the memory 2302, the input interface 2303, and the output interface 2304 may be implemented on a separate chip or circuit board, which is not limited in the embodiments of the present application.

[0266] Those skilled in the art will appreciate that the structure shown above does not constitute a limitation on the server 2300 , and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0267] In an exemplary embodiment, a chip is also provided. The chip includes a programmable logic circuit and / or program instructions. When the chip runs on a computer device, it is used to implement the data transmission method described in the above aspects.

[0268] In an exemplary embodiment, a computer program product is also provided, the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor reads and executes the computer instructions from the computer-readable storage medium to implement the data transmission method provided by each of the above method embodiments.

[0269] In an exemplary embodiment, a computer-readable storage medium is further provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the data transmission method provided by the above-mentioned method embodiments.

[0270] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0271] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented with hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that a general or special-purpose computer can access.

[0272] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A data transmission method, characterized in that: The method is executed by a first processor, the first processor, the second processor, and the third processor belong to the same server, the first processor and the second processor are connected via a bus in the server, and the first processor and the third processor are connected to a switch via a communication cable, and the method includes: transmitting first data between the first processor and the second processor based on the bus connection in the server; Based on the communication cable and the switch, second data is transmitted between the first processor and the third processor; the first data and the second data are transmitted in parallel; wherein the transmission of the second data passes through the second path or the third path; The second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence; the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, and the first aggregation layer switch is connected to the second access layer switch and the third access layer switch respectively; The third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence; the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch, respectively.

2. The method according to claim 1, characterized in that The transmitting the second data between the first processor and the third processor based on the communication cable and the switch includes: The second data is transmitted between the first processor and the third processor based on the communication cable and the network channel provided by the switch; the network channel is used to provide a data transmission path between the two processors connected by the communication cable.

3. The method according to claim 2, characterized in that The network channel is based on a ring-structured channel; Among them, a transmission path based on a network channel between any two processors among the first processor, the second processor and the third processor passes through at least one switch.

4. The method according to any one of claims 1 to 3, characterized in that: The transmitting first data between the first processor and the second processor based on the bus connection in the server includes: The first data is transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection; the interconnection channel is used to provide a data transmission path between at least two processors in the server.

5. A data transmission method, characterized in that: The method is executed by a control device in a server, a first processor, a second processor, and a third processor belong to the server, the first processor and the second processor are connected via a bus in the server, and the first processor and the third processor are connected to a switch via a communication cable, and the method includes: controlling first data to be transmitted between the first processor and the second processor, wherein the first data is transmitted based on the bus connection in the server; Controlling second data to be transmitted between the first processor and the third processor, wherein the second data is transmitted based on the communication cable and the switch; the first data and the second data are transmitted in parallel; wherein the transmission of the second data passes through the second path or the third path; The second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence; the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, and the first aggregation layer switch is connected to the second access layer switch and the third access layer switch respectively; The third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence; the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch, respectively.

6. The method according to claim 5, characterized in that The controlling the first data to be transmitted between the first processor and the second processor includes: The first data is controlled to be transmitted between the first processor and the second processor based on an interconnection channel provided by the bus connection.

7. The method according to claim 5, characterized in that The controlling the second data to be transmitted between the first processor and the third processor includes: The second data is controlled to be transmitted between the first processor and the third processor based on the communication cable and the network channel provided by the switch.

8. The method according to any one of claims 5 to 7, characterized in that: The method further comprises: In the case where the first processor, the second processor, and the third processor are directly connected to each other, a first channel resource bandwidth corresponding to the bus connection is allocated to at least one candidate interconnection channel, and a second channel resource bandwidth corresponding to the communication cable and the switch is allocated to at least one candidate network channel; The first data is transmitted via an interconnection channel, and the at least one candidate interconnection channel includes the interconnection channel; the second data is transmitted via a network channel, and the at least one candidate network channel includes the network channel.

9. The method according to claim 8, characterized in that In the case where the first processor, the second processor, and the third processor are directly connected to each other, allocating a first channel resource bandwidth corresponding to the bus connection to at least one candidate interconnection channel, and allocating a second channel resource bandwidth corresponding to the communication cable and the switch to at least one candidate network channel, comprising: When any one of the first processor, the second processor, and the third processor obtains the bus identifier of the other processor in a manner without going through an intermediate device, the first channel resource bandwidth corresponding to the bus connection is allocated to the at least one candidate interconnection channel, and the second channel resource bandwidth corresponding to the communication cable and the switch is allocated to the at least one candidate network channel.

10. The method according to claim 8, characterized in that The method further comprises: Determining a candidate channel resource ratio between the candidate interconnect channel and the candidate network channel according to a ratio between the first channel resource bandwidth and the second channel resource bandwidth; According to the preset number of channels and the channel resource ratio, a first number of the candidate interconnection channels and a second number of the candidate network channels are determined.

11. The method according to claim 10, characterized in that The method further comprises: According to the data volume of the first data and the second data, a ratio of the number of transmission channels of the interconnection channel and the network channel is determined in a load balancing manner.

12. A data transmission device, characterized in that: The first processor, the second processor, and the third processor belong to the same server, the first processor and the second processor are connected via a bus in the server, the first processor and the third processor are connected to a switch via a communication cable, and the device includes: a transmission module, configured to transmit first data between the first processor and the second processor based on the bus connection in the server; The transmission module is further used to transmit second data between the first processor and the third processor based on the communication cable and the switch; the first data and the second data are transmitted in parallel; wherein the transmission of the second data passes through the second path or the third path; The second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence; the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, and the first aggregation layer switch is connected to the second access layer switch and the third access layer switch respectively; The third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence; the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch, respectively.

13. A data transmission device, characterized in that: The first processor, the second processor, and the third processor belong to a server, the first processor and the second processor are connected via a bus in the server, the first processor and the third processor are connected to a switch via a communication cable, and the device includes: a control module, configured to control transmission of first data between the first processor and the second processor, wherein the first data is transmitted based on the bus connection in the server; The control module is further used to control the transmission of second data between the first processor and the third processor, wherein the second data is transmitted based on the communication cable and the switch; and the first data and the second data are transmitted in parallel; wherein the transmission of the second data passes through the second path or the third path; The second path starts from one of the first processor and the third processor and ends at the other processor, and passes through the second access layer switch, the first aggregation layer switch, and the third access layer switch in sequence; the first processor is connected to the second access layer switch, the third processor is connected to the third access layer switch, and the first aggregation layer switch is connected to the second access layer switch and the third access layer switch respectively; The third path starts from one of the first processor and the third processor and ends at the other processor, and passes through the fourth access layer switch, the second aggregation layer switch, the core layer switch, the third aggregation layer switch, and the fifth access layer switch in sequence; the first processor is connected to the fourth access layer switch, the third processor is connected to the fifth access layer switch, the second aggregation layer switch is connected to the fourth access layer switch, the third aggregation layer switch is connected to the fifth access layer switch, and the core layer switch is connected to the second aggregation layer switch and the third aggregation layer switch, respectively.

14. A computer device, characterized in that: The computer device comprises: a processor and a memory, wherein the memory stores at least one program; the processor is used to execute the at least one program in the memory to implement the data transmission method as described in any one of claims 1 to 11.

15. A computer-readable storage medium, characterized in that: The readable storage medium stores executable instructions, and the executable instructions are loaded and executed by a processor to implement the data transmission method according to any one of claims 1 to 11.

16. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. The processor reads and executes the computer instructions from the computer-readable storage medium to implement the data transmission method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Communication method, device and system suitable for heterogeneous environment

    CN111680791A

  • Heterogeneous interconnection system and cluster

    CN114968895A