Data transmission method, data processing system, processing chip and server
By deploying a communication engine and a shared configuration queue on the processing chip, the problems of poor flexibility and low efficiency of data transmission between different processing chips are solved, and efficient data transmission between arbitrary processing chips is achieved, which is suitable for scenarios such as AI and HPC.
Patent Information
- Application Number
- CN202311830735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
Data transmission between different processing chips is less flexible and less efficient, especially when the data transmission methods inside and outside the server are inconsistent.
Data transmission between arbitrary processing chips is achieved by deploying a communication engine and a shared configuration queue on the processing chip. The specific method includes the first processing chip writing a communication message carrying the chip memory address to the communication queue of the second processing chip, and the second processing chip transmits data based on the message.
It improves the data transmission efficiency between different processing chips, realizes fast data transmission between any processing chips in the system, and is suitable for AI, HPC and other scenarios.
Smart Images

Figure CN120216220A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data transmission method, a data processing system, a processing chip, and a server. Background Art
[0002] With the rapid development of technologies such as artificial intelligence (AI) and high performance computing (HPC), various types of processing chips have emerged, such as graphics processing units (GPUs), neural network processing units (XPU), intelligent processing units (IPU), tensor processing units (TPU), domain specific architecture (DSA) chips, and so on.
[0003] Generally, different processing chips inside a server are interconnected through a chip bus, and the processing chips of different servers are interconnected through a network switching device. In this way, data can be transmitted between the processing chips inside a server through the chip bus, and data can be transmitted between the processing chips of different servers through remote direct memory access (RDMA).
[0004] However, in the above methods, due to the different data transmission methods used inside and outside the server, the flexibility of data transmission between different processing chips is poor, and the data transmission efficiency is low. Summary of the Invention
[0005] Embodiments of this application provide a data transmission method, a data processing system, a processing chip, and a server, which can realize data transmission between any processing chips in the system and improve the data transmission efficiency between different processing chips. The technical solution is as follows:
[0006] In a first aspect, the present application provides a data transmission method, which is applied to scenarios involving data transmission between multiple processing chips, such as AI, HPC, video encoding and decoding, network packet processing, etc., and the present application does not limit this. Among them, the processing chip can be an acceleration chip, a general-purpose CPU, or a processing chip for video encoding and decoding, network packet processing, etc., and the present application does not limit the type of the processing chip. Schematically, the data processing system involved in the present application includes multiple servers, multiple processing chips are distributed on the multiple servers, each processing chip is interconnected, and there are multiple communication queues on each processing chip. The multiple communication queues are used to store communication messages received by the processing chip. Based on this data processing system, the data transmission method provided by the present application includes: the first processing chip writes a first communication message to the first communication queue on the second processing chip, and the first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located. The first communication queue is used to store the communication messages of the first processing chip; the second processing chip performs data transmission with the first processing chip based on the first communication message.
[0007] Through the above method, data transmission can be quickly realized between any two processing chips in the system, thereby improving the data transmission efficiency between different processing chips. Moreover, even between processing chips deployed on different servers, data transmission can be carried out in this way, thereby effectively improving the data transmission efficiency. For example, in the AI scenario, as the computing power of AI processing chips continues to increase and the models and parameters of AI applications continue to increase, the amount of data transmitted between AI processing chips depends more and more. Adopting the solution provided by the present application can achieve fast data access between AI processing chips, improve the data transmission load, and thus make the data transmission between AI processing chips more efficient.
[0008] In some embodiments, there is a communication engine on the port of each processing chip, and the communication engine is used to operate multiple communication queues on the processing chip. By deploying a communication engine on the port of the processing chip to operate multiple communication queues, since the communication engine is the component that the processing chip first receives external messages, the operation efficiency can be improved.
[0009] In some embodiments, there is also a shared configuration queue on each processing chip. For any processing chip, the shared configuration queue of the processing chip is used to store the mapping relationship between the shared memory space and the chip memory address of the processing chip. That is, multiple processing chips in the data processing system provide a shared memory space. In this case, each processing chip in the data processing system can access the shared memory space.
[0010] In some embodiments, if the data to be transmitted refers to the data to be transmitted from the first processing chip to the second processing chip, the second processing chip performs data transmission with the first processing chip based on the first communication message, including:
[0011] The second processing chip sends a data acquisition request to the first processing chip based on the first communication message;
[0012] The first processing chip determines the shared memory space corresponding to the data to be transmitted based on the data acquisition request and the shared configuration queue on the first processing chip, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted.
[0013] In some embodiments, if the data to be transmitted refers to the data to be transmitted from the second processing chip to the first processing chip, the second processing chip performs data transmission with the first processing chip based on the first communication message, including:
[0014] The second processing chip determines the shared memory space corresponding to the data to be transmitted based on the first communication message and the shared configuration queue on the second processing chip, and transmits the data to be transmitted to the first processing chip based on the shared memory space corresponding to the data to be transmitted.
[0015] Through the above method, for several different situations of the data to be transmitted, the implementation methods of data transmission based on the shared configuration queue are respectively introduced. In this way, in the case where multiple processing chips in the data processing system provide shared memory space, a flexible communication mechanism as described above can also be configured, thereby realizing efficient data transmission between any two processing chips.
[0016] In some embodiments, the first processing chip writes a first communication message to the first communication queue on the second processing chip, including: the first processing chip writes the first communication message to the first communication queue associated with the first chip identifier, and the first chip identifier is used to identify the first processing chip.
[0017] In some embodiments, the first processing chip writes a status message to the first communication queue, and the status message indicates the chip status of the first processing chip. For example, the first processing chip writes the status message to the queue head of the first communication queue. Among them, the chip status refers to the working status of the processing chip. For example, the working status is available or unavailable (such as the processing chip cannot work normally due to an exception, restart, or power-off). In this way, it is convenient for other processing chips to timely know the chip status of the first processing chip.
[0018] In some embodiments, the first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers. In this way, flexible communication between heterogeneous systems can be achieved.
[0019] In some embodiments, the first processing chip and the second processing chip are interconnected through a switching chip. There is routing information on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip.
[0020] In some embodiments, the first processing chip and the second processing chip are interconnected through a chip bus. The second processing chip performs data transmission with the first processing chip based on a first communication message, including: the second processing chip sends an address access request to the first processing chip through the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip. The chip bus is, for example, NvLink, CXL, UCIe, HCCS, CCIX, etc. Since the chip bus can directly transmit the memory access requests issued by the acceleration units in the acceleration chip or the CPU cores in the general CPU, the memory access requests passing through this bus do not need to perform request format conversion or be encapsulated into requests of another format. Therefore, in this way, the data transmission efficiency between different processing chips can be effectively improved.
[0021] In some embodiments, the method further includes: the second processing chip reads the first communication message by using an interrupt mechanism and / or a polling mechanism. That is to say, the communication mechanism provided in this application supports multiple message reading methods and can be flexibly used according to actual needs.
[0022] In a second aspect, an embodiment of the present application provides a data transmission system. The data processing system includes multiple servers, multiple processing chips are distributed on the multiple servers, the processing chips are interconnected with each other, and each processing chip has multiple communication queues. The multiple communication queues are used to store the communication messages received by the processing chip;
[0023] A first processing chip is configured to write a first communication message to a first communication queue on a second processing chip. The first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located, and the first communication queue is used to store the communication messages of the first processing chip;
[0024] The second processing chip is configured to perform data transmission with the first processing chip based on the first communication message.
[0025] In some embodiments, there is a communication engine on the port of each processing chip, and the communication engine is used to operate the multiple communication queues on the processing chip.
[0026] In some embodiments, each of the processing chips further has a shared configuration queue, which is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
[0027] In some embodiments, the first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers.
[0028] In some embodiments, the first processing chip and the second processing chip are interconnected through a switching chip, and routing information is configured on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip.
[0029] In some embodiments, the first processing chip and the second processing chip are interconnected through a chip bus;
[0030] The second processing chip is configured to send an address access request to the first processing chip through the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip.
[0031] In a third aspect, an embodiment of the present application provides a processing chip, which includes a port for interconnecting with other processing chips outside the processing chip, and the processing chip is configured to implement the functions of the processing chip in the data transmission method provided in the foregoing first aspect or any optional manner of the first aspect.
[0032] In a fourth aspect, an embodiment of the present application provides a server, which includes at least one processing chip, and the processing chip is configured to implement the functions of the processing chip in the data transmission method provided in the foregoing first aspect or any optional manner of the first aspect.
[0033] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store at least one program code, and the at least one program code is used to implement the data transmission method provided in the foregoing first aspect or any optional manner of the first aspect. The storage medium includes, but is not limited to, volatile memories such as random access memories, and non-volatile memories such as flash memories, hard disk drives (HDDs), and solid state drives (SSDs).
[0034] In a sixth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a processing chip, it enables the processing chip to implement the functions of the processing chip in the data transmission method provided in the foregoing first aspect or any optional manner of the first aspect. The computer program product can be a software installation package. In the case where it is necessary to implement the functions of the foregoing processing chip, the computer program product can be downloaded and executed on the processing chip. Description of the Drawings
[0035] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0036] Figure 2 is a schematic diagram of another implementation environment provided by an embodiment of the present application;
[0037] Figure 3 is a schematic diagram of the structure of a processing chip provided by an embodiment of the present application;
[0038] Figure 4 is a schematic diagram of the structure of another processing chip provided by an embodiment of the present application;
[0039] Figure 5 is a schematic diagram of a data processing system provided by an embodiment of the present application;
[0040] Figure 6 is a schematic diagram of the structure of a communication engine provided by an embodiment of the present application;
[0041] Figure 7 is a schematic diagram of a data processing system provided by an embodiment of the present application;
[0042] Figure 8 is a schematic diagram of a routing information configuration process provided by an embodiment of the present application;
[0043] Figure 9 is a schematic diagram of a switching chip provided by an embodiment of the present application;
[0044] Figure 10 is a schematic diagram of a data transmission method provided by an embodiment of the present application;
[0045] Figure 11 is a schematic diagram of a communication queue provided by an embodiment of the present application;
[0046] Figure 12 is a schematic diagram of a shared configuration queue provided by an embodiment of the present application;
[0047] Figure 13 is a schematic diagram of writing a status message to a communication queue provided by an embodiment of the present application;
[0048] Figure 14 It is a schematic diagram of the process of a data transmission method provided by an embodiment of the present application;
[0049] Figure 15 It is a schematic diagram of the process of another data transmission method provided by an embodiment of the present application. Detailed implementation manners
[0050] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the communication messages, location information of the processing chip, etc. involved in the present application are obtained under full authorization.
[0051] For the convenience of understanding, the following first explains the key terms and key concepts involved in the present application.
[0052] An acceleration chip, also known as an acceleration device or acceleration card, is a type of dedicated hardware accelerator or computer system designed to accelerate the computing process in AI scenarios and HPC scenarios. Schematically, the acceleration chip is, for example, a graphics processing unit (GPU), a neural network processing unit (XPU), an intelligent processing unit (IPU), a tensor processing unit (TPU), a domain specific architecture (DSA) chip, etc., and the present application is not limited thereto.
[0053] A chip bus, or chip interconnect bus, is a bus that enables direct transmission of memory access requests issued by acceleration units in an acceleration chip or CPU cores in a general-purpose processor (central processing unit, CPU). Memory access requests passing through this bus do not require request format conversion or encapsulation into another format of request. In some embodiments, the bus through which the CPU in a computer system issues requests to access the acceleration chip in the system is referred to as the CPU bus (CPUbus). In the embodiments of this application, the chip bus is, for example, NVIDIA Link (NvLink), Compute Express Link (CXL), Universal Chiplet Interconnect Express (UCIe), Huawei Cache Coherent System (HCCS), Cache Coherent Interconnect for Accelerators (CCIX), etc., and this application is not limited thereto.
[0054] Unified addressing means that in a computer system, the address spaces of internal and external memories are unifiedly addressed, enabling them to use the same address addressing method. For example, all memories in a computer system, including main memory, cache, and external memory, are regarded as a continuous address space, and each storage unit can be addressed and accessed through a unique address. In some embodiments, a computer system adopting unified addressing is simply referred to as a unified addressing system.
[0055] The application scenarios and implementation environments of this application are introduced below.
[0056] This application is applicable to scenarios involving data transmission between multiple processing chips. For example, scenarios such as AI, HPC, video codec, and network packet processing are applicable, and this application is not limited thereto. Among them, the processing chip can be an acceleration chip, a general-purpose CPU, or a processing chip for video codec, network packet processing, etc., and this application is not limited to the type of processing chip. Schematically, this application provides a data processing system and a data transmission method, which can achieve point-to-point communication between any processing chips and effectively improve the data transmission efficiency between different processing chips. For example, in the AI scenario, as the computing power of AI processing chips continuously increases, the models and parameters of AI applications continuously increase, and the data volume transmitted between AI processing chips depends more and more. Adopting the solution provided by this application can achieve fast data access between AI processing chips, increase the data transmission load, and thus make the data transmission between AI processing chips more efficient.
[0057] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application. As Figure 1 shown, the implementation environment includes a data processing system 100, and the data processing system 100 includes a plurality of servers 101. A plurality of processing chips 102 are distributed on the plurality of servers 101, and the respective processing chips 102 are interconnected. In some embodiments, the plurality of servers 101 may be connected to a wireless network or a wired network, and the present application does not limit this.
[0058] For any one server 101, the server 101 includes at least one processing chip 102. Among them, the server 101 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms, etc. The present application does not limit this.
[0059] For any one processing chip 102, the processing chip 102 has storage capabilities and computing capabilities. The processing chip 102 is, for example, an acceleration chip (refer to the foregoing introduction and will not be elaborated here), a general-purpose CPU, etc. The present application is not limited thereto. In some embodiments, the respective processing chips 102 are interconnected through a chip bus. Among them, the chip bus is used to transmit data and various requests (such as address access requests) for the respective processing chips 102. The chip bus is, for example, NvLink, CXL, UCIe, HCCS, CCIX, etc. The present application is not limited thereto.
[0060] In the embodiment of the present application, there are a plurality of communication queues on each processing chip 102, and the plurality of communication queues on each processing chip 102 are used to store communication messages received by the corresponding processing chip 102. That is, for any one processing chip, the processing chip can determine whether to perform data transmission with other processing chips by operating the plurality of communication queues on the processing chip 102. This process will be introduced in detail in the subsequent method embodiments and will not be elaborated here.
[0061] In some embodiments, the data processing system 100 further includes a switching chip 103, and the processing chips 102 are interconnected through the switching chip 103. The switching chip 103 is used to connect different processing chips 102, that is, to establish a data transmission path between different processing chips 102 and provide a routing function for data transmission between different processing chips 102. The switching chip 103 is, for example, any chip with data switching capabilities, and the present application does not limit this. It should be understood that the switching chip 103 is an optional device, and the processing chips 102 can also be directly interconnected with each other.
[0062] In some embodiments, for any server 101, the server 101 includes multiple processing chips 102, and the multiple processing chips 102 include the processing chips of the host and at least one processing chip connected to the host. The host refers to a device for running various computing services and can control at least one processing chip connected thereto to execute corresponding computing tasks. For example, taking the AI scenario as an example, the host is connected to multiple AI processing chips, and an AI model is running on the host. The host can control the multiple AI processing chips to perform distributed parallel training on the AI model, that is, load the distributed parallel training task for the AI model into each AI processing chip for running. Schematically, refer to Figure 2 , Figure 2 which is a schematic diagram of another implementation environment provided by the embodiments of the present application. As Figure 2 shown, in this implementation environment, the server includes the processing chips of the host and at least one processing chip connected to the host, and each processing chip is interconnected with the switching chip through a chip bus. It should be understood that Figure 2 is only for illustrative purposes, and each processing chip can also be directly interconnected through a chip bus, and the present application is not limited thereto.
[0063] It should be noted that the numbers of the server 101, the processing chip 102, and the switching chip 103 shown in Figure 1 and Figure 2 are only schematic, and the numbers of the server 101, the processing chip 102, and the switching chip 103 can be more or less, and the embodiments of the present application do not limit this.
[0064] In some embodiments, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is generally a Transmission Control Protocol / Internet Protocol (TCP / IP) network and an RDMA network in a data center network, such as an RDMA over Converged Ethernet (RoCE) network, an InfiniBand (IB) network, etc., which is not limited thereto. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0065] The following introduces the hardware structure of the processing chip in the above data processing system.
[0066] Figure 3 It is a schematic structural diagram of a processing chip provided by an embodiment of the present application. As Figure 3 shown, the processing chip 300 includes a memory 301, a processor 302, a port 303, and a bus 304. Among them, the memory 301, the processor 302, and the port 303 are communicatively connected to each other through the bus 304.
[0067] The memory 301 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. Schematically, the memory 301 is used to store at least one segment of program code. When the program code stored in the memory 301 is executed by the processor 302, the processing chip 300 executes the steps involved in the processing chip in the following method embodiments.
[0068] The processor 302 may be a network processor (NP), a central processing unit (CPU), an application-specific integrated circuit (ASIC), or an integrated circuit for controlling the execution of the program of the solution of the present application. The processor 302 may be a single-CPU processor or a multi-CPU processor. The number of the processors 302 may be one or more. Schematically, the processor 302 may include one processing unit or multiple processing units. Each processing unit may be composed of a die, and different processing units may be interconnected through a bus. The processor 302 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that processes signals based on operation instructions. Among other functions, the processor 302 is configured to obtain and execute computer-readable instructions and data stored in the memory.
[0069] The port 303 uses a transceiver module such as a transceiver to implement communication between the processing chip 300 and other processing chips or a communication network. For example, data can be obtained through the port 303. In the embodiment of the present application, there is a communication engine on the port 303, and the communication engine is used to operate multiple communication queues on the processing chip 300. Since the communication engine is the component that first receives external messages on the processing chip 300, the operation efficiency can be improved.
[0070] Among them, the memory 301 and the processor 302 may be separately provided or integrated together.
[0071] The bus 304 may include a path for transmitting information between various components of the processing chip 300 (for example, the memory 301, the processor 302, and the port 303).
[0072] Figure 4 It is a schematic structural diagram of another processing chip provided by the embodiment of the present application. As Figure 4 shown, the processing chip 400 includes a memory 401, a processor 402, a port 403, an accelerator 404, and a bus 405. Among them, the memory 401, the processor 402, the port 403, and the accelerator 404 are communicatively connected to each other through the bus 405. In some embodiments, the processing chip 400 is also referred to as an acceleration chip or an AI processing chip.
[0073] The memory 401 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM, or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. Schematically, the memory 401 is used to store at least one segment of program code. When the program code stored in the memory 401 is executed by the processor 402, the processing chip 400 is caused to execute the steps involved in the processing chip in the following method embodiments.
[0074] The processor 402 can be an NP, a CPU, an ASIC, or an integrated circuit for controlling the execution of the program of the solution of the present application. The processor 402 can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The number of the processors 402 can be one or multiple.
[0075] The port 403 uses a transceiver module such as a transceiver to implement the communication between the processing chip 400 and other processing chips or a communication network. For example, data can be obtained through the port 403. In some embodiments, there is a communication engine on the port 403, and the communication engine is used to operate multiple communication queues on the processing chip 400. In some embodiments, the communication engine can also be separately arranged from the port 403, and the present application does not make any limitation thereto.
[0076] The accelerator 404 can be a GPU, an XPU, an IPU, a TPU, a DSA chip, etc. Schematically, the accelerator 404 is used to provide computing power for the computing tasks running on the processing chip 400 to accelerate the computing process. In addition, the number of the accelerators 404 can be one or multiple, and the present application does not make any limitation thereto.
[0077] Among them, the memory 401, the processor 402, and the accelerator 404 can be separately arranged or integrated together.
[0078] The bus 405 can include a path for transmitting information between various components of the processing chip 400 (for example, the memory 401, the processor 402, the port 403, the accelerator 404).
[0079] Based on the above Figures 1 to 4The implementation environment shown. Further, in the present application, a communication mechanism between multiple processing chips is established through the communication engines on each processing chip in the data processing system, enabling data transmission between any processing chips and improving data transmission efficiency. The following refers to Figure 5 , in combination with the above implementation environment, to further introduce the data processing system provided by the present application.
[0080] Figure 5 is a schematic diagram of a data processing system provided by an embodiment of the present application. As Figure 5 shown, the data processing system includes multiple processing chips, and the processing chips are interconnected through a switching chip (only for illustration in the figure, and they may not be interconnected through a switching chip). Each processing chip has multiple communication queues, and there is a communication engine on the port of each processing chip, and the communication engine is used to operate the multiple communication queues on the processing chip. It should be understood that the server is not marked in the figure, but the multiple processing chips shown in the figure can be deployed in the same server or in different servers respectively, and the present application does not limit this.
[0081] For any one processing chip, the functions of the communication engine of the processing chip refer to Figure 6 , Figure 6 is a structural schematic diagram of a communication engine provided by an embodiment of the present application. As Figure 6 shown, the functions of the communication engine include but are not limited to the chip identification management function 601 and the communication queue management function 602.
[0082] The chip identification management function 601 is used to obtain the location information of the processing chip, and generate the chip identification (or simply referred to as chip ID) of the processing chip based on the location information of the processing chip. It should be understood that the chip identification of the processing chip is used to identify the processing chip in the data processing system, so that other processing chips can perform data transmission with it according to the chip identification of the processing chip. Among them, the location information indicates the physical deployment location of the processing chip in the system. Schematically, the location information of the processing chip includes at least one of the following: cabinet ID, frame ID, slot ID, module ID, and this application is not limited thereto. Taking the location information including cabinet ID = A, frame ID = B, slot ID = C, and module ID = D as an example, the way for the communication engine to generate the chip ID through the chip identification management function 601 is, for example: chip ID = A × m + B × n + C × o + D, where m is the maximum number of cabinets planned in the data processing system, n is the maximum number of frames planned in one cabinet, o is the maximum number of slots planned in one frame. For example, m = 2, n = 2, o = 2. Then, for a processing chip with location information including cabinet ID = 0, frame ID = 1, slot ID = 2, and module ID = 3, its chip ID = 0 × 2 + 1 × 2 + 2 × 2 + 3 = 9. In addition, the introduction of the generation method of the chip identification here is only for illustrative purposes and does not constitute a limitation to this application. In practical applications, other methods can also be used to generate the chip identification.
[0083] The communication queue management function 602 is used to initialize multiple communication queues, manage multiple communication queues, and operate multiple communication queues to realize data transmission between the processing chip and other processing chips. In some embodiments, when multiple processing chips in the data processing system provide a shared memory space, the communication queue management function 602 is also used to initialize the shared configuration queue of the processing chip, and the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address. It should be understood that in this case, each processing chip in the data processing system can access the shared memory space with each other. Taking the AI scenario as an example, the shared memory space that is shared out is, for example, the memory space where the AI processing chip or the CPU chip participates in AI computing. This part of the memory space stores the parameters, sample data, etc. that the AI processing chip needs to process during the AI computing process. In some embodiments, the shared memory space of multiple processing chips in the data processing system can be addressed using a unified addressing method, which is convenient for data transmission between each processing chip based on address access requests.
[0084] In addition, the functions of the communication engine are not limited to the above-mentioned chip identification management function 601 and communication queue management function 602. In practical applications, more functions can be set according to the needs of users, and this application does not make any limitations in this regard.
[0085] In some embodiments, continue to refer toFigure 5 , the data processing system further includes a location information providing server and a routing information configuration server. The functions of these two servers will be introduced separately below.
[0086] The location information providing server is used to provide location information for each processing chip in the data processing system, that is, to send the location information to the communication engine of each processing chip. The sending method can be through a low-speed bus such as inter-integrated circuit (I2C), system management bus (SMBUS), etc. This application is not limited thereto. When the communication engine of each processing chip obtains the location information, it can generate a chip identifier of the processing chip according to the location information. For example, for any processing chip, its location information includes at least one of the following: cabinet ID, frame ID, slot ID, module ID. Schematically, the location information providing server obtains the IDs of the cabinet, frame, and slot by reading the hardware memory on the cabinet, frame, and slot, and then sends them to the communication engine of the processing chip. The communication engine obtains the module ID by reading the register of the processing chip through the chip identifier management function, and then generates the chip identifier. Of course, relevant personnel can also directly number each processing chip in the data processing system according to its deployment location, assign cabinet ID, frame ID, slot ID, and module ID, and then send them to the communication engine through the location information providing server, etc. This application does not limit the method for each processing chip to obtain location information.
[0087] In addition, since the forms of different servers often vary, the identifiers of the location information of the processing chips will also vary. A data processing system composed of servers with different forms is also called a heterogeneous system (that is, a complete system composed of multiple processing units or subsystems with different architectures). For example, in a cabinet server, a server cabinet is divided into frames, and a frame is further divided into modules; in addition to cabinet servers, there are also rack servers, and a rack server is composed of multiple rack-mounted servers. There are also tower servers, blade servers, etc. This application is applicable to various types of servers. Taking a cabinet server as an example, a cabinet includes multiple frames, a frame includes multiple slots, each slot has a server, and a slot includes multiple processing chips. When the data processing system includes a switching chip, the switching chip is usually responsible for data exchange between multiple processing chips within a slot, and is also responsible for data exchange between slots and between frames. Of course, the number of switching chips can be more or less. For example, when a data processing system consists of multiple cabinets, the number of switching chips will increase. Each switching chip is responsible for data exchange between processing chips in its own location and also for data exchange between processing chips across locations. Schematically, refer to Figure 7 , Figure 7is a schematic diagram of a data processing system provided in an embodiment of the present application, such as Figure 7 As shown in Figure (a), there are two slots in a frame that are interconnected by a switch chip, and each frame has an independent switch chip. Figure 7 As shown in Figure (b), two cabinets are interconnected through a switching chip, and each cabinet has an independent bus switching chip. It should be understood that Figure 7 This is only an example and does not constitute a limitation on the networking architecture between the various processing chips in this application.
[0088] The routing information configuration server is used to provide routing information for the switching chip when the data processing system includes a switching chip. The routing information indicates the data transmission path between different processing chips. For example, when the shared memory space of each processing chip adopts a unified addressing method, the shared memory space of each processing chip is different, and the range of this shared memory space is associated with the location of the processing chip; an address access request sent to a certain processing chip or an address access request sent from a certain processing chip is routed through the switching chip. Usually, each frame and cabinet has an independent switching chip. When these switching chips are responsible for routing the address access request entering the cabinet and frame, they configure the corresponding routing information based on the range of the shared memory space of the processing chip in the area to which they belong, so that the address access request is sent to the processing chip on the specified cabinet, frame, and slot according to the routing information.
[0089] Schematically, refer to Figure 8 , Figure 8 FIG. 1 is a schematic diagram of a routing information configuration process provided by an embodiment of the present application. Figure 8 As shown, the location information providing server sends the location information of each processing chip and the location information of the switching chip to the routing information configuration server. The routing information configuration server generates routing reference information of the switching chip (such as routing reference information 0 to n, where n is a positive integer) based on the location information of each processing chip and the location information of the switching chip. The routing reference information is then configured to the switching chip, and the switching chip configures the corresponding routing information based on the routing reference information. The routing reference information is, for example, the range of the shared memory space of each processing chip. The routing information indicates the data transmission path between different processing chips. In some embodiments, the data transmission path in the routing information is associated with the chip identifier of the processing chip. It should be understood that after the switching chip configures the routing information, the switching chip in the data processing system has the ability to allow accelerators, CPU cores, etc. between processing chips to be directly accessed using addresses through a unified addressing method. Reference Figure 9 , Figure 9 is a schematic diagram of a switching chip provided in an embodiment of the present application, such as Figure 9As shown, multiple ports of the switching chip are interconnected with multiple processing chips, and routing information is configured on the routing engine of the switching chip, which can indicate the data transmission path between any two processing chips connected to the switching chip. For example, processing chip 0 is interconnected with port 0 of the switching chip, and processing chip 1 is interconnected with port 1 of the switching chip. If the switching chip receives an address access request from processing chip 1, and the address access request indicates data transmission with processing chip 1, then the switching chip routes the address access request to processing chip 1 connected to port 1 based on the routing information. That is, it transmits the data in the specified address space on one processing chip to the specified address space on another processing chip.
[0090] It should be noted that the location information providing server and the routing information configuring server are both optional devices. For example, relevant personnel can also directly configure the location information and routing information to each processing chip and switching chip. In addition, the functions provided by the location information providing server and the routing information configuring server can be implemented by one physical device or by different physical devices respectively. This application is not limited thereto.
[0091] Based on the above Figures 1 to 9 introduced data processing system, the data transmission method provided by this application will be introduced below.
[0092] Figure 10 It is a schematic diagram of a data transmission method provided by an embodiment of this application. As Figure 10 shown, taking the interaction between any two processing chips in the data processing system as an example for introduction, schematically, the data transmission method includes the following steps 1001 to step 1005.
[0093] 1001. Multiple processing chips initialize multiple communication queues on their respective processing chips, and the multiple communication queues are used to store communication messages received by the processing chips.
[0094] In the embodiment of this application, for any one processing chip, a communication engine is deployed on the port of the processing chip, and the processing chip initializes multiple communication queues through the communication engine.
[0095] Schematically, for any processing chip, the communication engine of the processing chip initializes a plurality of communication queues based on the number of processing chips in the data processing system. In some embodiments, the number of processing chips is equal to the number of communication queues. For example, if the data processing system includes 256 processing chips, the communication engine initializes 256 communication queues, and the 256 communication queues correspond one-to-one with the 256 processing chips, or in other words, one communication queue belongs to one processing chip. It should be understood that this is only an example. In some embodiments, the number of communication queues can also be greater than the number of processing chips. For example, if the data processing system includes 256 processing chips, the communication engine initializes 300 communication queues, and the first 256 communication queues correspond one-to-one with the 256 processing chips. In this way, in the case of adding a processing chip to the subsequent data processing system, the communication queues not associated with the processing chip can be associated with the newly added processing chip, thereby improving the processing efficiency.
[0096] In some embodiments, for any processing chip, each communication queue on the processing chip is associated with the chip identifier of each processing chip in the data processing system. Schematically, the chip identifier of the processing chip is determined based on the physical deployment location of the processing chip in the data processing system, that is, based on the location information of the processing chip. For the location information of the processing chip, refer to Figure 5 the relevant content and will not be elaborated here.
[0097] In some embodiments, for any processing chip, the multiple communication queues on the processing chip are stored in the shared memory space of the processing chip. The shared memory space can be located in the storage unit of the upper port of the processing chip or in the memory of the processing chip. This application does not make a limitation on this. It should be understood that the shared memory space for storing multiple communication queues can be accessed by multiple processing chips in the data processing system. For example, the depth of each communication queue is the default value of 2KB. Of course, it can also be configured as other values when initializing multiple communication queues. This application does not make a limitation on this.
[0098] Schematically, refer to Figure 11 , Figure 11 which is a schematic diagram of a communication queue provided by an embodiment of this application. As Figure 11 shown, taking the data processing system including 256 processing chips as an example, each processing chip has 256 communication queues, and each communication queue corresponds one-to-one with each processing chip. The depth of each communication queue is the default value of 2KB, and the shared memory space occupied by the 256 communication queues is 512KB. Taking Figure 11Taking any two of the shown processing chips as an example, after initializing multiple communication queues, if processing chip 1 requests to transfer data with processing chip 2, then processing chip 1 writes a communication message to communication queue 1 corresponding to processing chip 1 on processing chip 2. In this way, when processing chip 2 reads the communication message from communication queue 1, it can perform data transfer with processing chip 1 based on this communication message. That is to say, when a certain processing chip in the data processing system requests to transfer data with other processing chips, it can write the communication message to the communication queue belonging to this processing chip on other processing chips; when a certain processing chip receives a communication message from other processing chips, it reads the corresponding communication message from the communication queue belonging to other processing chips on its own chip.
[0099] After step 1001 above, each processing chip in the data processing system initializes multiple communication queues on its own processing chip. Since the multiple communication queues are used to store communication messages from other processing chips, these multiple communication queues provide technical support for point-to-point communication between each processing chip.
[0100] 1002. Multiple processing chips initialize a shared configuration queue on their own processing chips, and this shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
[0101] In the embodiment of the present application, for any one processing chip, a communication engine is deployed on the port of this processing chip. When multiple processing chips in the data processing system provide a shared memory space, this processing chip initializes its shared configuration queue through the communication engine.
[0102] Schematically, for any one processing chip, the communication engine of this processing chip initializes the shared configuration queue based on the shared memory space corresponding to this processing chip in the data processing system. In some embodiments, the shared memory space of multiple processing chips in the data processing system is evenly allocated based on the target shared memory space of the data processing system. For example, the target shared memory space of the data processing system is 128TB, supporting 256 processing chips to be interconnected in a unified addressing manner. The target shared memory space is equally divided according to the number 256, and the shared memory space that each processing chip can be accessed by other processing chips is 512GB. Based on this, the communication engine of the processing chip establishes the mapping relationship between the shared memory space of this processing chip and the chip memory address of this processing chip based on the address translation unit (ATU), and stores it in the shared configuration queue.
[0103] In some embodiments, for any processing chip, the shared configuration queue on the processing chip is associated with the chip identifier of the processing chip. Schematically, the chip identifier of the processing chip is determined based on the physical deployment location of the processing chip in the data processing system, that is, based on the location information of the processing chip. For the location information of the processing chip, refer to Figure 5 the relevant content, which will not be elaborated here.
[0104] In some embodiments, for any processing chip, the shared configuration queue on the processing chip is stored in the shared memory space of the processing chip. The shared memory space can be located in the storage unit of the upper port of the processing chip or in the memory of the processing chip. This application does not make any limitations in this regard.
[0105] Schematically, refer to Figure 12 , Figure 12 which is a schematic diagram of a shared configuration queue provided by an embodiment of this application. As Figure 12 shown, taking the data processing system including 256 processing chips as an example, each processing chip has a shared configuration queue corresponding to the processing chip. Taking Figure 12 the processing chip 2 shown as an example, its communication engine, based on the ATU, establishes a mapping relationship between the shared memory space of the processing chip 2 and the chip memory address of the processing chip 2, and stores it in the shared configuration queue. In this way, when other processing chips request to access the chip memory address of the processing chip 2, it can be achieved by accessing the shared memory space 2 of the processing chip 2, thereby improving the data transmission efficiency.
[0106] After the above step 1002, each processing chip in the data processing system initializes the shared configuration queue on its own processing chip. Since the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip, the shared configuration queue provides technical support for data transmission between each processing chip based on the address access request.
[0107] Next, in combination with steps 1003 to 1005, taking the interaction between any two processing chips (the first processing chip and the second processing chip) as an example, the process of data transmission between any two processing chips will be introduced. Among them, the first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers. This application does not make any limitations in this regard.
[0108] 1003. The first processing chip writes a first communication message to the first communication queue on the second processing chip. The first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located. The first communication queue is used to store the communication messages of the first processing chip.
[0109] In an embodiment of the present application, there are multiple communication queues on the second processing chip. Among them, the first communication queue is used to store communication messages of the first processing chip. In some embodiments, the first processing chip writes a first communication message to the first communication queue associated with the first chip identifier, and the first chip identifier is used to identify the first processing chip.
[0110] The first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located, including the following situations: First, the first processing chip requests to send the data on the first processing chip to the second processing chip (or there is data on the first processing chip that needs to be obtained by the second processing chip), then the chip memory address is the chip memory address on the first processing chip. That is, the data to be transmitted refers to the data to be transmitted from the first processing chip to the second processing chip. Second, the first processing chip requests to read data from the second processing chip, then the chip memory address is the chip memory address on the second processing chip. That is, the data to be transmitted refers to the data to be transmitted from the second processing chip to the first processing chip. In other words, the present application does not limit the data transmission method indicated by the communication message. It can be that the source processing chip sends data to the destination processing chip, or the source processing chip obtains data from the destination processing chip.
[0111] In some embodiments, before the first processing chip writes a first communication message to the first communication queue on the second processing chip, the first processing chip writes a status message to the first communication queue on the second processing chip, and the status message indicates the chip status of the first processing chip. Among them, the chip status refers to the working status of the processing chip. For example, the working status is available or unavailable (such as the processing chip cannot work normally due to an exception, restart, or power-down). Schematically, the communication engine of the first processing chip writes the status message to the first communication queue. In some embodiments, the first processing chip writes the status message to the head of the first communication queue, and the present application does not limit this. In addition, the first processing chip may write a first status message to the first communication queue on the second processing chip when initializing multiple communication queues, and the first status message indicates that the chip status of the first processing chip is available (for example, marked as active). When an exception, restart, power-down, etc. occur to the first processing chip, a second status message is written to the first communication queue on the second processing chip, and the second status message indicates that the chip status of the first processing chip is unavailable (for example, marked as inactive). In some embodiments, the status message may also carry other content, such as the available time period of the processing chip, the reason for unavailability, etc., and the present application does not limit this.
[0112] In addition, the above description takes the first processing chip writing a status message to the first communication queue on the second processing chip as an example. In some instances, the first processing chip may write a status message to the first communication queues on all processing chips in the data processing system, or may write a status message to the first communication queues on other processing chips in the system except the first processing chip. Schematically, referring to Figure 13 , Figure 13 which is a schematic diagram of writing a status message to a communication queue provided by an embodiment of the present application. As shown in Figure 13 , the first processing chip (such as processing chip 1) writes a status message to the first communication queues on other processing chips in the system except the first processing chip. For example, a status message is written to the queue head of the first communication queue. It should be understood that the first processing chip can update the status message in the first communication queue on other processing chips in a timely manner according to the chip status of the first processing chip, so that other processing chips can learn the chip status of the first processing chip in a timely manner. It should be noted that any processing chip in the data processing system can implement the functions possessed by the first processing chip, which will not be elaborated here.
[0113] 1004. The second processing chip reads the first communication message from the first communication queue.
[0114] In the embodiment of the present application, after the first processing chip writes the first communication message to the first communication queue on the second processing chip, the second processing chip can read the first communication message by using an interrupt mechanism and / or a polling mechanism. The present application does not make any limitation in this regard. That is to say, the communication mechanism provided by the present application supports multiple message reading methods and can be flexibly used according to actual requirements.
[0115] Taking the interrupt mechanism as an example, after the first processing chip writes the first communication message to the first communication queue on the second processing chip, an interrupt message is sent to the second processing chip. After receiving the interrupt message, the second processing chip reads the first communication message from the first communication queue. By adopting the interrupt mechanism, the processing chip can learn the communication message in a timely manner and save the data transmission time.
[0116] Taking the polling mechanism as an example, the second processing chip polls and reads multiple communication queues on the second processing chip. If it reads that there is a communication message stored in the first communication queue, it reads the first communication message from the first communication queue. Schematically, the communication engine of the second processing chip reads the first communication message from the first communication queue. By adopting the polling mechanism, the communication resources between different processing chips can be saved.
[0117] 1005. The second processing chip performs data transmission with the first processing chip based on the first communication message.
[0118] In the embodiments of the present application, since the first communication message carries the chip memory address of the data to be transmitted, the second processing chip can perform data transmission with the first processing chip based on the chip memory address carried in the first communication message. In some embodiments, when multiple processing chips in the data processing system provide a shared memory space, this step includes the following situations:
[0119] First, the data to be transmitted refers to the data to be transmitted from the first processing chip to the second processing chip. Schematically, the second processing chip sends a data acquisition request to the first processing chip based on the first communication message; the first processing chip determines the shared memory space corresponding to the data to be transmitted based on the data acquisition request and the shared configuration queue on the first processing chip, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted.
[0120] Second, the data to be transmitted refers to the data to be transmitted from the second processing chip to the first processing chip. Schematically, the second processing chip determines the shared memory space corresponding to the data to be transmitted based on the first communication message and the shared configuration queue on the second processing chip, and transmits the data to be transmitted to the first processing chip based on the shared memory space corresponding to the data to be transmitted.
[0121] Through the above method, for several different situations of the data to be transmitted, the implementation methods of data transmission based on the shared configuration queue are introduced respectively. In this way, when multiple processing chips in the data processing system provide a shared memory space, a flexible communication mechanism as described above can also be configured, thereby realizing efficient data transmission between any two processing chips.
[0122] In some embodiments, if the first processing chip and the second processing chip are interconnected through a chip bus, the second processing chip sends an address access request to the first processing chip through the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip. Based on the foregoing introduction, the chip bus is, for example, NvLink, CXL, UCIe, HCCS, CCIX, etc. Since the chip bus can directly transmit the memory access requests issued by the acceleration units in the acceleration chip or the CPU cores in the general-purpose CPU, and the memory access requests passing through this bus do not need to be converted in request format or encapsulated into requests of another format, through this method, the data transmission efficiency between different processing chips can be effectively improved.
[0123] In some embodiments, if the first processing chip and the second processing chip are interconnected through a chip bus and a switching chip, the second processing chip sends an address access request to the first processing chip through the chip bus and the switching chip based on a first communication message. The address access request indicates data transmission between the first processing chip and the second processing chip. Among them, there is routing information on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip. For the routing information, refer to the foregoing introduction and will not be elaborated herein.
[0124] The following refers to Figure 14 and Figure 15 , taking the second processing chip reading the first communication message based on the interrupt mechanism or the polling mechanism as an example, to illustrate the data transmission method introduced in the above steps 1003 to 1005.
[0125] Figure 14 is a schematic process diagram of a data transmission method provided by an embodiment of the present application. As Figure 14 shown, taking the first processing chip (processing chip 1) requesting to send the data on the first processing chip to the second processing chip (processing chip 2) as an example (or there is data on the first processing chip that needs to be obtained by the second processing chip), the communication engine of the first processing chip writes a first communication message to the first communication queue (communication queue 1) on the second processing chip. The first communication message carries the chip memory address of the data to be transmitted on the first processing chip, such as the PA0 address. The communication engine of the first processing chip sends an interrupt message to the second processing chip. The communication engine of the second processing chip receives the interrupt message, reads the first communication message from the first communication queue, obtains the chip memory address of the data to be transmitted from it, and sends it to the processor on the second processing chip. The processor sends an address access request to the first processing chip through the switching chip based on the chip memory address of the data to be transmitted, that is, accesses the shared memory space 1 of the first processing chip, so that the first processing chip determines the shared memory space corresponding to the data to be transmitted based on the address access request and the shared configuration queue, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted. It should be understood that the process of the first processing chip requesting to read data from the second processing chip (such as the PA1 address) is the same as the above process, so it will not be elaborated.
[0126] Figure 15 is a schematic process diagram of another data transmission method provided by an embodiment of the present application. As Figure 15As shown in the figure, taking the example that the first processing chip (Processing Chip 1) requests to send the data on the first processing chip to the second processing chip (Processing Chip 2) (or there is data on the first processing chip that needs to be obtained by the second processing chip), the communication engine of the first processing chip writes a first communication message to the first communication queue (Communication Queue 1) on the second processing chip. The first communication message carries the chip memory address of the data to be transmitted on the first processing chip, such as the PA0 address. The communication engine of the second processing chip polls and reads multiple communication queues on the second processing chip. If it reads that there is a communication message stored in the first communication queue, it reads the first communication message from the first communication queue, obtains the chip memory address of the data to be transmitted from it, and sends it to the processor on the second processing chip. The processor, based on the chip memory address of the data to be transmitted, sends an address access request to the first processing chip through the switching chip, that is, accesses the shared memory space 1 of the first processing chip, so that the first processing chip determines the shared memory space corresponding to the data to be transmitted based on the address access request and the shared configuration queue, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted. It should be understood that the process of the first processing chip requesting to read data from the second processing chip (such as the PA1 address) is the same as the above process, so it will not be elaborated here.
[0127] In addition, the above embodiments are introduced by taking the example of the first processing chip writing a communication message to the second processing chip. Similarly, the first processing chip can also read the communication message from the second processing chip from the local communication queue and perform data transmission with the second processing chip, which will not be elaborated here.
[0128] The following combines the foregoing Figure 5 , and gives an example to illustrate the overall process of the data transmission method provided by the embodiments of the present application. As Figure 5 shown, the data processing system includes multiple processing chips, and each processing chip is interconnected with a switching chip through a chip bus (or can be interconnected without passing through a switching chip). There is a communication engine on the port of each processing chip, and the communication engine is used to operate multiple communication queues and shared configuration queues. It should be understood that the server is not marked in the figure, but the multiple processing chips shown in the figure can be deployed in the same server or can be deployed in different servers respectively, and the present application does not limit this. In some embodiments, the data processing system further includes a location information providing server and a routing information configuration server. Among them, the location information providing server is used to provide location information for each processing chip in the data processing system, and the routing information configuration server is used to provide routing information for the switching chip. In Figure 5In the data processing system shown, the communication engine on the processing chip, the location information providing server, and the routing information configuration server cooperate with each other to establish a flexible communication mechanism between any processing chips, enabling data transmission between any processing chips through address access requests, and improving the data transmission efficiency between different processing chips. Moreover, this communication mechanism does not require the management of a heavy communication state machine and does not require complex message encapsulation. Instead, it realizes data transmission between different processing chips in the system through the chip identifiers of each processing chip and the communication queues on each processing chip, thereby achieving fast data access, increasing the data transmission load, and making the data transmission between different processing chips more efficient.
[0129] In summary, the data transmission method provided in the embodiments of the present application is applied to a data processing system. Multiple processing chips are distributed on multiple servers of the system, and each processing chip has multiple communication queues for storing communication messages received by the processing chip. Based on this, when data is transmitted between any two processing chips, the first processing chip writes a first communication message to the first communication queue corresponding to the first processing chip on the second processing chip, enabling the second processing chip to perform data transmission with the first processing chip based on the first communication message. Through this flexible communication mechanism, fast data transmission between any two processing chips in the system is realized, thereby improving the data transmission efficiency between different processing chips. Moreover, even between processing chips deployed on different servers, data can be transmitted in this way, effectively improving the data transmission efficiency.
[0130] In addition, the present application also provides a data transmission device configured on the processing chip, which can implement the steps performed by the processing chip in the above method embodiments. Schematically, the device includes at least one functional unit, and the at least one functional unit is used to implement the steps performed by the processing chip in the foregoing method embodiments. For example, the at least one functional unit includes a communication message writing unit and a data transmission unit. Among them, the communication message writing unit is used to write communication messages to the communication queues on other processing chips, and the data transmission unit is used to perform data transmission with other processing chips based on the communication messages read from the communication queues. It should be understood that when the data transmission device performs data transmission, only the above-mentioned division of each functional unit is used for illustration. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above. In addition, the data transmission device provided in the above embodiments and the data transmission method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0131] In this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various described examples, the first processing chip can be referred to as the second processing chip, and similarly, the second processing chip can be referred to as the first processing chip. Both the first processing chip and the second processing chip can be processing chips, and in some cases, they can be separate and different processing chips.
[0132] In this application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "multiple" refers to two or more. For example, multiple processing chips refer to two or more processing chips.
[0133] The above description is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0134] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of program structure information. This program structure information includes one or more program instructions. When the program instructions are loaded and executed on a computing device, the processes or functions in the embodiments of this application are generated in whole or in part.
[0135] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. This program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, an optical disc, or the like.
[0136] As mentioned above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent substitutions on some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A data transmission method, characterized in that, Applied to a data processing system, the data processing system includes a plurality of servers, on which a plurality of processing chips are distributed, the processing chips are interconnected with each other, and each processing chip has a plurality of communication queues for storing communication messages received by the processing chip. The method includes: A first processing chip writes a first communication message to a first communication queue on a second processing chip, the first communication message carrying the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located, and the first communication queue is used to store the communication messages of the first processing chip; The second processing chip performs data transmission with the first processing chip based on the first communication message.
2. The method according to claim 1, wherein There is a communication engine on the port of each processing chip, and the communication engine is used to operate the plurality of communication queues on the processing chip.
3. The method according to claim 1 or 2, characterized in that, There is also a shared configuration queue on each processing chip, and the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
4. The method according to claim 3, characterized in that, If the data to be transmitted refers to the data to be transmitted from the first processing chip to the second processing chip, the second processing chip performs data transmission with the first processing chip based on the first communication message, including: The second processing chip sends a data acquisition request to the first processing chip based on the first communication message; The first processing chip determines the shared memory space corresponding to the data to be transmitted based on the data acquisition request and the shared configuration queue on the first processing chip, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted.
5. The method according to claim 3, wherein If the data to be transmitted refers to the data to be transmitted from the second processing chip to the first processing chip, the second processing chip performs data transmission with the first processing chip based on the first communication message, including: The second processing chip determines the shared memory space corresponding to the data to be transmitted based on the first communication message and the shared configuration queue on the second processing chip, and transmits the data to be transmitted to the first processing chip based on the shared memory space corresponding to the data to be transmitted.
6. The method according to any one of claims 1 to 5, characterized in that The first processing chip writes the first communication message to the first communication queue on the second processing chip, including: The first processing chip writes the first communication message to the first communication queue associated with the first chip identifier, and the first chip identifier is used to identify the first processing chip.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The first processing chip writes a status message to the first communication queue, and the status message indicates the chip status of the first processing chip.
8. The method according to claim 7, characterized in that, The first processing chip writes a status message to the first communication queue, including: The first processing chip writes the status message to the head of the first communication queue.
9. The method according to any one of claims 1 to 8, characterized in that, The first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers.
10. The method according to any one of claims 1 to 9, characterized in that The first processing chip is interconnected with the second processing chip through a switching chip, and there is routing information on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip.
11. The method according to any one of claims 1 to 10, characterized in that, The first processing chip is interconnected with the second processing chip through a chip bus, and the second processing chip performs data transmission with the first processing chip based on the first communication message, including: The second processing chip sends an address access request to the first processing chip through the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: The second processing chip reads the first communication message by using an interrupt mechanism and / or a polling mechanism.
13. A data transmission system, characterized in that, The data processing system includes multiple servers, multiple processing chips are distributed on the multiple servers, the processing chips are interconnected with each other, and there are multiple communication queues on each processing chip, and the multiple communication queues are used to store communication messages received by the processing chip; A first processing chip is configured to write a first communication message to a first communication queue on a second processing chip, the first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located, and the first communication queue is used to store communication messages of the first processing chip; The second processing chip is configured to perform data transmission with the first processing chip based on the first communication message.
14. The system according to claim 13, wherein There is a communication engine on the port of each processing chip, and the communication engine is used to operate the multiple communication queues on the processing chip.
15. The system according to claim 13 or 14, characterized in that, There is also a shared configuration queue on each processing chip, and the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
16. The system according to any one of claims 13 to 15, characterized in that, The first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers.
17. The system according to any one of claims 13 to 16, characterized in that The first processing chip is interconnected with the second processing chip through a switching chip, and routing information is configured on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip.
18. The system according to any one of claims 13 to 17, characterized in that The first processing chip is interconnected with the second processing chip through a chip bus; The second processing chip is configured to send an address access request to the first processing chip through the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip.
19. A processing chip, characterized in that, The processing chip includes a port, the port is used to be interconnected with other processing chips outside the processing chip, and the processing chip is used to implement the functions of the processing chip in the data transmission method described in any one of the preceding claims 1 to 12.
20. A server, characterized in that, The server includes at least one processing chip, and the processing chip is used to implement the functions of the processing chip in the data transmission method described in any one of the preceding claims 1 to 12.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one program code, and the at least one program code is used to implement the data transmission method described in any one of the foregoing claims 1 to 12.
22. A computer program product, characterized in that, When the computer program product runs on the processing chip, the processing chip is caused to implement the functions of the processing chip in the data transmission method described in any one of the foregoing claims 1 to 12.
Citation Information
Cited By
Data transmission method, data processing system, processing chip, and server
WO2025140221A1