Data transmission method, data processing system, processing chip, and server
By deploying communication engines and queue mechanisms on the processing chips, the problem of flexibility and low efficiency in data transmission between processing chips is solved, and an efficient data transmission mechanism is realized, suitable for scenarios such as AI and HPC.
Patent Information
- Application Number
- PCT/CN2024/142025
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-03
AI Technical Summary
There are problems of poor flexibility and low efficiency in data transmission methods between different processing chips, especially in data transmission between the server and the external environment.
By deploying a communication engine on each processing chip, using multiple communication queues and shared configuration queues, a flexible communication mechanism between processing chips is realized, including writing communication messages, reading status messages and data transmission based on shared memory space, supporting interrupt mechanisms and polling mechanisms, and using chip bus and switching chips for data path routing.
It improves the data transmission efficiency between different processing chips, especially in heterogeneous systems, and realizes fast and efficient data access, which is suitable for data transmission in AI, HPC and other scenarios.
Smart Images

Figure CN2024142025_03072025_PF_FP_ABST
Abstract
Description
Data transmission method, data processing system, processing chip and server
[0001] This application claims priority to Chinese patent application No. 202311830735.2 filed on December 27, 2023, entitled “Data transmission method, data processing system, processing chip and server”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a data transmission method, a data processing system, a processing chip, and a server. Background Art
[0003] With the rapid development of technologies such as artificial intelligence (AI) and high-performance computing (HPC), various types of processing chips have emerged, such as graphics processing units (GPUs), neural network processing units (XPUs), intelligent processing units (IPUs), tensor processing units (TPUs), domain-specific architecture (DSA) chips, and so on.
[0004] Typically, different processing chips within a server are interconnected via a chip bus, and processing chips in different servers are interconnected via network switching devices. In this way, data can be transmitted between processing chips within a server via the chip bus, and data can be transmitted between processing chips in different servers via remote direct memory access (RDMA).
[0005] However, in the above method, since the data transmission methods used inside the server and outside the server are different, the flexibility of data transmission between different processing chips is poor and the data transmission efficiency is low. Summary of the Invention
[0006] The embodiments of the present application provide a data transmission method, a data processing system, a processing chip, and a server, which can realize data transmission between any processing chips in the system and improve the data transmission efficiency between different processing chips. The technical solution is as follows:
[0007] In the first aspect, the present application provides a data transmission method, which is applied to scenarios involving data transmission between multiple processing chips, such as AI, HPC, video encoding and decoding, network message processing, etc., which are not limited in this application. Among them, the processing chip can be an acceleration chip, a general-purpose CPU, or a processing chip for video encoding and decoding, network message processing, etc., and the present application does not limit the type of processing chip. Schematically, the data processing system involved in the present application includes multiple servers, multiple processing chips are distributed on the multiple servers, and each processing chip is interconnected. Each processing chip has multiple communication queues, and the multiple communication queues are used to store communication messages received by the processing chip. Based on this data processing system, the data transmission method provided by the present application includes: a first processing chip writes a first communication message to a first communication queue on a second processing chip, the first communication message carries the chip memory address of the data to be transmitted between the first processing chip and the second processing chip, and the first communication queue is used to store the communication message of the first processing chip; the second processing chip transmits data with the first processing chip based on the first communication message.
[0008] Through the above-mentioned method, data transmission can be quickly realized between any two processing chips in the system, thereby improving the data transmission efficiency between different processing chips. Moreover, data transmission can be carried out even between processing chips deployed on different servers in this way, thereby effectively improving the data transmission efficiency. For example, in the AI scenario, as the computing power of AI processing chips continues to increase, the models and parameters of AI applications continue to increase, and the amount of data transmission between AI processing chips is increasing. The solution provided by this application can achieve fast data access between AI processing chips, increase the data transmission load, and thus make the data transmission between AI processing chips more efficient.
[0009] In some embodiments, each processing chip port includes a communication engine that operates multiple communication queues on the processing chip. Deploying a communication engine on a processing chip port to operate multiple communication queues improves operational efficiency because the communication engine is the first component of the processing chip to receive external messages.
[0010] In some embodiments, each processing chip also includes a shared configuration queue. For any processing chip, the shared configuration queue for that processing chip is used to store the mapping between the shared memory space of that processing chip and the chip's memory addresses. In other words, multiple processing chips in the data processing system provide shared memory space. In this case, each processing chip in the data processing system can access the shared memory space.
[0011] In some embodiments, if the data to be transmitted is data to be transmitted from the first processing chip to the second processing chip, the second processing chip transmits data with the first processing chip based on the first communication message, including:
[0012] The second processing chip sends a data acquisition request to the first processing chip based on the first communication message;
[0013] The first processing chip determines the shared memory space corresponding to the data to be transmitted based on the data acquisition request and the shared configuration queue on the first processing chip, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted.
[0014] In some embodiments, if the data to be transmitted is data to be transmitted from the second processing chip to the first processing chip, the second processing chip transmits data with the first processing chip based on the first communication message, including:
[0015] The second processing chip determines a shared memory space corresponding to the data to be transmitted based on the first communication message and the shared configuration queue on the second processing chip, and transmits the data to be transmitted to the first processing chip based on the shared memory space corresponding to the data to be transmitted.
[0016] Through the above method, for several different situations of data to be transmitted, the implementation method of data transmission based on shared configuration queues is introduced respectively. In this way, when multiple processing chips in the data processing system provide shared memory space, it is also possible to configure a flexible communication mechanism as described above, thereby realizing efficient data transmission between any two processing chips.
[0017] In some embodiments, the first processing chip writes a first communication message to a first communication queue on the second processing chip, including: the first processing chip writes the first communication message to a first communication queue associated with a first chip identifier, where the first chip identifier is used to identify the first processing chip.
[0018] In some embodiments, the first processing chip writes a status message to the first communication queue, where the status message indicates the chip status of the first processing chip. For example, the first processing chip writes a status message to the queue head of the first communication queue. The chip status refers to the working status of the processing chip, for example, the working status is available or unavailable (e.g., the processing chip cannot work normally due to an exception, restart, or power failure). In this way, other processing chips can be informed of the chip status of the first processing chip in a timely manner.
[0019] In some embodiments, the first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are deployed in different servers, thereby enabling flexible communication between heterogeneous systems.
[0020] In some embodiments, the first processing chip and the second processing chip are interconnected via a switching chip. The switching chip carries routing information, and the routing information indicates a data transmission path between the first processing chip and the second processing chip.
[0021] In some embodiments, a first processing chip and a second processing chip are interconnected via a chip bus, and the second processing chip transmits data with the first processing chip based on a first communication message, including: the second processing chip sends an address access request to the first processing chip via the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip. The chip bus is, for example, NvLink, CXL, UCIe, HCCS, CCIX, etc. Since the chip bus allows memory access requests issued by the acceleration unit in the acceleration chip or the CPU core in the general-purpose CPU to be directly transmitted, the memory access requests passing through the bus do not need to be converted into a request format or encapsulated into a request of another format. Therefore, this method can effectively improve the efficiency of data transmission between different processing chips.
[0022] In some embodiments, the method further includes: the second processing chip uses an interrupt mechanism and / or a polling mechanism to read the first communication message. In other words, the communication mechanism provided by this application supports multiple message reading methods and can be flexibly used according to actual needs.
[0023] In a second aspect, an embodiment of the present application provides a data transmission system, wherein the data processing system includes multiple servers, multiple processing chips are distributed on the multiple servers, the processing chips are interconnected, and each processing chip has multiple communication queues, and the multiple communication queues are used to store communication messages received by the processing chip;
[0024] a first processing chip, configured to write a first communication message to a first communication queue on a second processing chip, the first communication message carrying a chip memory address of data to be transmitted between the first processing chip and the second processing chip, the first communication queue being configured to store communication messages of the first processing chip;
[0025] The second processing chip is configured to perform data transmission with the first processing chip based on the first communication message.
[0026] In some embodiments, a communication engine is provided on a port of each processing chip, and the communication engine is configured to operate the plurality of communication queues on the processing chip.
[0027] In some embodiments, each processing chip further includes a shared configuration queue, which is used to store a mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
[0028] In some embodiments, the first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are deployed in different servers respectively.
[0029] In some embodiments, the first processing chip and the second processing chip are interconnected via a switch chip, and routing information is configured on the switch chip, where the routing information indicates a data transmission path between the first processing chip and the second processing chip.
[0030] In some embodiments, the first processing chip and the second processing chip are interconnected via a chip bus;
[0031] The second processing chip is configured to send an address access request to the first processing chip via the chip bus based on the first communication message, where the address access request indicates data transmission with the first processing chip.
[0032] In a third aspect, an embodiment of the present application provides a processing chip, which includes a port, and the port is used to interconnect with other processing chips outside the processing chip. The processing chip is used to implement the functions of the processing chip in the data transmission method provided in the first aspect or any optional method of the first aspect.
[0033] In a fourth aspect, an embodiment of the present application provides a server, which includes at least one processing chip, and the processing chip is used to implement the functions of the processing chip in the data transmission method provided in the first aspect or any optional method of the first aspect.
[0034] In a fifth aspect, embodiments of the present application provide a computer-readable storage medium for storing at least one program code segment, wherein the at least one program code segment is used to implement the data transmission method provided in the first aspect or any optional embodiment of the first aspect. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0035] In a sixth aspect, embodiments of the present application provide a computer program product. When the computer program product is executed on a processing chip, the processing chip implements the functions of the processing chip in the data transmission method provided in the first aspect or any optional embodiment of the first aspect. The computer program product may be a software installation package. When the functions of the processing chip are to be implemented, the computer program product may be downloaded and executed on the processing chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0037] FIG2 is a schematic diagram of another implementation environment provided by an embodiment of the present application;
[0038] FIG3 is a schematic diagram of the structure of a processing chip provided in an embodiment of the present application;
[0039] FIG4 is a schematic diagram of the structure of another processing chip provided in an embodiment of the present application;
[0040] FIG5 is a schematic diagram of a data processing system provided in an embodiment of the present application;
[0041] FIG6 is a structural diagram of a communication engine provided in an embodiment of the present application;
[0042] FIG7 is a schematic diagram of a data processing system provided in an embodiment of the present application;
[0043] FIG8 is a schematic diagram of a routing information configuration process provided in an embodiment of the present application;
[0044] FIG9 is a schematic diagram of a switching chip provided in an embodiment of the present application;
[0045] FIG10 is a schematic diagram of a data transmission method provided in an embodiment of the present application;
[0046] FIG11 is a schematic diagram of a communication queue provided in an embodiment of the present application;
[0047] FIG12 is a schematic diagram of a shared configuration queue provided in an embodiment of the present application;
[0048] 13 is a schematic diagram of writing a status message to a communication queue according to an embodiment of the present application;
[0049] FIG14 is a process diagram of a data transmission method provided in an embodiment of the present application;
[0050] FIG15 is a process diagram of another data transmission method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the communication messages, location information of the processing chip, etc. involved in this application are all obtained with full authorization.
[0052] For ease of understanding, the key terms and key concepts involved in this application are explained below.
[0053] An accelerator chip, also known as an accelerator device or accelerator card, is a specialized hardware accelerator or computer system designed to accelerate computing processes in AI and HPC scenarios. Examples of accelerator chips include graphics processing units (GPUs), neural network processing units (XPUs), intelligent processing units (IPUs), tensor processing units (TPUs), and domain-specific architecture (DSA) chips, though this application is not limited thereto.
[0054] The chip bus, or chip interconnect bus, is a bus that allows the memory access request issued by the acceleration unit in the acceleration chip or the CPU core in the general-purpose processor (central processing unit, CPU) to be directly transmitted. The memory access request passing through the bus does not need to be converted into a request format or encapsulated into a request of another format. In some embodiments, the bus issued by the CPU in the computer system and used to access the acceleration chip in the system is called the CPU bus (CPU bus). In the embodiment of the present application, the chip bus is, for example, NVIDIA Link (NvLink), Compute Express Link (CXL), Universal Chiplet Interconnect Express (UCIe), Huawei Cache Coherent System (HCCS), Cache Coherent Interconnect for Accelerators (CCIX), etc., and the present application is not limited thereto.
[0055] Unified addressing refers to the unified addressing of the address spaces of internal and external memory in a computer system, allowing them to use the same addressing scheme. For example, all memory in a computer system, including main memory, cache, and external memory, is treated as a continuous address space, and each storage unit can be addressed and accessed using a unique address. In some embodiments, a computer system that uses unified addressing is referred to as a unified addressing system.
[0056] The following is an introduction to the application scenarios and implementation environment of this application.
[0057] The present application is applied to scenarios involving data transmission between multiple processing chips, such as AI, HPC, video encoding and decoding, network message processing and other scenarios, which are not limited in the present application. Among them, the processing chip can be an acceleration chip, or a general-purpose CPU, or a processing chip for video encoding and decoding, network message processing, etc. The present application does not limit the type of processing chip. Schematically, the present application provides a data processing system and a data transmission method, which can realize point-to-point communication between any processing chips and effectively improve the data transmission efficiency between different processing chips. For example, in the AI scenario, as the computing power of the AI processing chip continues to increase, the models and parameters of AI applications continue to increase, and the amount of data transmitted between AI processing chips is getting larger and larger. The solution provided by the present application can realize fast data access between AI processing chips, increase the data transmission load, and thus make the data transmission between AI processing chips more efficient.
[0058] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application. As shown in Figure 1, the implementation environment includes a data processing system 100, which includes multiple servers 101. Multiple processing chips 102 are distributed across the multiple servers 101, and each processing chip 102 is interconnected. In some embodiments, the multiple servers 101 can be connected to a wireless network or a wired network, which is not limited in this application.
[0059] For any server 101, the server 101 includes at least one processing chip 102. The server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc. This application does not limit this.
[0060] For any processing chip 102, the processing chip 102 has storage and computing capabilities. The processing chip 102 may be, for example, an accelerator chip (refer to the above description and will not be repeated here), a general-purpose CPU, etc., but the present application is not limited thereto. In some embodiments, the processing chips 102 are interconnected via a chip bus, where the chip bus is used to transmit data and various requests (such as address access requests) for each processing chip 102. The chip bus may be, for example, NvLink, CXL, UCIe, HCCS, CCIX, etc., but the present application is not limited thereto.
[0061] In the embodiment of the present application, each processing chip 102 has multiple communication queues, each of which is used to store communication messages received by the corresponding processing chip 102. That is, any processing chip can determine whether to transmit data to another processing chip by operating the multiple communication queues on that processing chip 102. This process will be described in detail in subsequent method embodiments and will not be repeated here.
[0062] In some embodiments, the data processing system 100 further includes a switch chip 103, through which the processing chips 102 are interconnected. The switch chip 103 is used to connect different processing chips 102, that is, to establish a data transmission path between different processing chips 102 and provide a routing function for data transmission between different processing chips 102. The switch chip 103 can be, for example, any chip with data exchange capabilities, and this application is not limited thereto. It should be understood that the switch chip 103 is optional, and the processing chips 102 can also be directly interconnected.
[0063] In some embodiments, for any server 101, the server 101 includes multiple processing chips 102, and the multiple processing chips 102 include a processing chip of a host and at least one processing chip connected to the host, wherein the host refers to a device for running various types of computing services, and can control at least one processing chip connected to it to perform corresponding computing tasks. For example, taking the AI scenario as an example, the host is connected to multiple AI processing chips, and an AI model is running on the host. The host can control multiple AI processing chips to perform distributed parallel training on the AI model, that is, the distributed parallel training tasks for the AI model are loaded into each AI processing chip for execution. Schematically, referring to Figure 2, Figure 2 is a schematic diagram of another implementation environment provided by an embodiment of the present application. As shown in Figure 2, in this implementation environment, the server includes a processing chip of a host and at least one processing chip connected to the host, and each processing chip is interconnected with a switching chip via a chip bus. It should be understood that Figure 2 is only an example, and each processing chip can also be directly interconnected via a chip bus, and the present application is not limited thereto.
[0064] It should be noted that the number of servers 101, processing chips 102 and switching chips 103 shown in Figures 1 and 2 is only for reference. The number of servers 101, processing chips 102 and switching chips 103 can be more or less, and this embodiment of the application does not limit this.
[0065] In some embodiments, the wireless network or wired network described above uses standard communication technologies and / or protocols. The network is typically a Transmission Control Protocol / Internet Protocol (TCP / IP) network in a data center network and an RDMA network, such as an RDMA over converged Ethernet (RoCE) network or an InfiniBand (IB) network, without limitation. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the above-mentioned data communication technologies.
[0066] The hardware structure of the processing chip in the above data processing system is introduced below.
[0067] Figure 3 is a schematic diagram of the structure of a processing chip provided in an embodiment of the present application. As shown in Figure 3, the processing chip 300 includes a memory 301, a processor 302, a port 303, and a bus 304, wherein the memory 301, the processor 302, and the port 303 are connected to each other through the bus 304.
[0068] The memory 301 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. Schematically, the memory 301 is used to store at least one section of program code. When the program code stored in the memory 301 is executed by the processor 302, the processing chip 300 performs the steps involved in the processing chip in the following method embodiment.
[0069] The processor 302 may be a network processor (NP), a central processing unit (CPU), an application-specific integrated circuit (ASIC), or an integrated circuit for controlling the execution of the program of the present application. The processor 302 may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The number of processors 302 may be one or more. Schematically, the processor 302 may include one processing unit or multiple processing units, each of which may be composed of a die, and different processing units may be interconnected via a bus. The processor 302 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that processes signals based on operating instructions. In addition to other functions, the processor 302 is configured to obtain and execute computer-readable instructions and data stored in a memory.
[0070] Port 303 uses a transceiver module, such as a transceiver, to enable communication between processing chip 300 and other processing chips or a communication network. For example, data can be retrieved through port 303. In an embodiment of the present application, port 303 includes a communication engine that operates multiple communication queues on processing chip 300. Because the communication engine is the first component of processing chip 300 to receive external messages, it can improve operational efficiency.
[0071] The memory 301 and the processor 302 may be separately provided or integrated together.
[0072] The bus 304 may include a path for transmitting information between various components of the processing chip 300 (eg, the memory 301 , the processor 302 , and the port 303 ).
[0073] Figure 4 is a schematic diagram of the structure of another processing chip provided in an embodiment of the present application. As shown in Figure 4, the processing chip 400 includes a memory 401, a processor 402, a port 403, an accelerator 404, and a bus 405, wherein the memory 401, the processor 402, the port 403, and the accelerator 404 are connected to each other via the bus 405. In some embodiments, the processing chip 400 is also called an acceleration chip, or an AI processing chip.
[0074] The memory 401 may be a ROM or other type of static storage device capable of storing static information and instructions, a RAM or other type of dynamic storage device capable of storing information and instructions, an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Schematically, the memory 401 is used to store at least one segment of program code. When the program code stored in the memory 401 is executed by the processor 402, the processing chip 400 performs the steps involved in the processing chip in the following method embodiments.
[0075] The processor 402 may be an NP, a CPU, an ASIC, or an integrated circuit for controlling the execution of the program of the present application. The processor 402 may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The number of the processors 402 may be one or more.
[0076] Port 403 uses a transceiver module, such as a transceiver, to enable communication between processing chip 400 and other processing chips or a communication network. For example, data can be retrieved through port 403. In some embodiments, port 403 includes a communication engine that operates multiple communication queues on processing chip 400. In some embodiments, the communication engine can also be provided separately from port 403, which is not limited in this application.
[0077] Accelerator 404 can be a GPU, XPU, IPU, TPU, DSA chip, etc. Illustratively, accelerator 404 is used to provide computing power for computing tasks running on processing chip 400, thereby accelerating the computing process. Furthermore, there can be one or more accelerators 404, which is not limited in this application.
[0078] The memory 401 , the processor 402 , and the accelerator 404 may be provided separately or integrated together.
[0079] The bus 405 may include a path for transmitting information between various components of the processing chip 400 (eg, the memory 401 , the processor 402 , the port 403 , and the accelerator 404 ).
[0080] Based on the implementation environments shown in Figures 1 to 4 above, this application further establishes a communication mechanism between multiple processing chips through a communication engine on each processing chip in the data processing system, enabling data transmission between any processing chips and improving data transmission efficiency. Referring now to Figure 5 , the data processing system provided by this application will be further described in conjunction with the above implementation environment.
[0081] Figure 5 is a schematic diagram of a data processing system provided by an embodiment of the present application. As shown in Figure 5, the data processing system includes multiple processing chips, each of which is interconnected through a switching chip (the figure is only for example, and the interconnection through the switching chip is also possible). Each processing chip has multiple communication queues, and each processing chip has a communication engine on its port, which is used to operate the multiple communication queues on the processing chip. It should be understood that the server is not marked in the figure, but the multiple processing chips shown in the figure can be deployed in the same server or in different servers, and this application does not limit this.
[0082] For any processing chip, the functions of the communication engine of the processing chip are shown in Figure 6, which is a structural diagram of a communication engine provided by an embodiment of the present application. As shown in Figure 6, the functions of the communication engine include but are not limited to chip identification management function 601 and communication queue management function 602.
[0083] The chip identification management function 601 is used to obtain the location information of the processing chip and generate a chip identification (or simply chip ID) for the processing chip based on the location information of the processing chip. It should be understood that the chip identification of the processing chip is used to identify the processing chip in the data processing system, so that other processing chips can transmit data with the processing chip based on the chip identification of the processing chip. The location information indicates the physical deployment location of the processing chip in the system. Schematically, the location information of the processing chip includes at least one of the following: cabinet ID, frame ID, slot ID, and module ID. The present application is not limited to this. Taking the example of location information including cabinet ID=A, frame ID=B, slot ID=C, and module ID=D, the communication engine generates a chip ID through the chip identification management function 601 in the following manner: chip ID=A×m+B×n+C×o+D, where m is the maximum number of cabinets planned for the data processing system, n is the maximum number of frames planned in a cabinet, and o is the maximum number of slots planned in a frame. For example, if m=2, n=2, and o=2, then for a processing chip whose location information includes cabinet ID=0, frame ID=1, slot ID=2, and module ID=3, its chip ID=0×2+1×2+2×2+3=9. Furthermore, the description of the chip identification generation method here is for illustrative purposes only and does not constitute a limitation of the present application. In actual applications, other methods may also be used to generate chip identifications.
[0084] The communication queue management function 602 is used to initialize multiple communication queues, manage multiple communication queues, and operate multiple communication queues to realize data transmission between the processing chip and other processing chips. In some embodiments, when multiple processing chips in the data processing system provide shared memory space, the communication queue management function 602 is also used to initialize the shared configuration queue of the processing chip, which is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address. It should be understood that in this case, the various processing chips in the data processing system can access each other's shared memory space. Taking the AI scenario as an example, the shared memory space is, for example, the memory space of the AI processing chip or the CPU chip participating in AI calculations. This part of the memory space stores the parameters, sample data, etc. that the AI processing chip needs to process during the AI calculation process. In some embodiments, the shared memory space of multiple processing chips in the data processing system can be addressed using a unified addressing method to facilitate data transmission between the various processing chips based on address access requests.
[0085] In addition, the functions of the communication engine are not limited to the above-mentioned chip identification management function 601 and communication queue management function 602. In actual applications, more functions can be set according to user needs, and this application does not limit this.
[0086] In some embodiments, referring to Figure 5 , the data processing system further includes a location information providing server and a routing information configuration server. The functions of these two servers are described below.
[0087] The location information providing server is used to provide location information for each processing chip in the data processing system, that is, to send the location information to the communication engine of each processing chip. The sending method can be through a low-speed bus such as the inter-integrated circuit I2C, the system management bus (SMBUS), etc., but the present application is not limited to this. When the communication engine of each processing chip obtains the location information, it can generate a chip identification of the processing chip based on the location information. For example, for any processing chip, its location information includes at least one of the following: cabinet ID, frame ID, slot ID, module ID. Schematically, the location information providing server obtains the ID of the cabinet, frame, and slot by reading the hardware memory storage on the cabinet, frame, and slot, and then sends it to the communication engine of the processing chip. The communication engine uses the chip identification management function to read the register of the processing chip to obtain the module ID, and then generates the chip identification. Of course, relevant personnel can also directly number each processing chip in the data processing system according to its deployment location, assign cabinet ID, frame ID, slot ID, module ID, and then send it to the communication engine through the location information providing server, etc. This application does not limit the way in which each processing chip obtains location information.
[0088] In addition, since the morphology of different servers often differs, the identification of the location information of the processing chip will also differ. A data processing system composed of servers with different morphologies is also called a heterogeneous system (that is, a complete system composed of multiple processing units or subsystems with different architectures). For example, in a cabinet server, a server cabinet is divided into frames, and the frames are divided into modules; in addition to cabinet servers, there are rack servers, and a rack server is composed of multiple rack servers. There are also tower servers, blade servers, etc., and this application is applicable to various types of servers. Taking a cabinet server as an example, a cabinet includes multiple frames, a frame includes multiple slots, each slot has a server, and a slot includes multiple processing chips. In the case where the data processing system includes a switching chip, the switching chip is usually responsible for data exchange between multiple processing chips in a slot, and is also responsible for data exchange between slots and frames. Of course, the number of switching chips can be more or less. For example, when a data processing system is composed of multiple cabinets, the number of switching chips will increase. In addition to being responsible for data exchange between processing chips in its own position, each switching chip is also responsible for data exchange between processing chips across positions. Schematically, referring to FIG. 7 , FIG. 7 is a schematic diagram of a data processing system provided in an embodiment of the present application. As shown in FIG. 7 (a), two slots within a frame are interconnected via a switch chip, and each frame has an independent switch chip. As shown in FIG. 7 (b), two cabinets are interconnected via a switch chip, and each cabinet has an independent bus switch chip. It should be understood that FIG. 7 is merely an example and does not constitute a limitation on the networking architecture between the various processing chips in this application.
[0089] When a data processing system includes a switching chip, a routing information configuration server is used to provide routing information to the switching chip. The routing information indicates the data transmission path between different processing chips. For example, when the shared memory space of each processing chip uses a unified addressing method, each processing chip has a different shared memory space, and the scope of this shared memory space is associated with the location of the processing chip. An address access request sent to or from a processing chip is routed through the switching chip. Typically, each frame or cabinet has an independent switching chip. These switching chips are responsible for routing address access requests entering the frame or cabinet. Based on the scope of the shared memory space of the processing chips in the area, they configure corresponding routing information so that the address access request is sent to the processing chip in the designated frame, cabinet, or slot according to the routing information.
[0090] Schematically, referring to FIG8 , FIG8 is a schematic diagram of a routing information configuration process provided by an embodiment of the present application. As shown in FIG8 , the location information providing server sends the location information of each processing chip and the location information of the switching chip to the routing information configuration server. The routing information configuration server generates routing reference information of the switching chip (such as routing reference information 0 to n, where n is a positive integer) based on the location information of each processing chip and the location information of the switching chip, and then configures the routing reference information to the switching chip. The switching chip configures the corresponding routing information based on the routing reference information. Among them, the routing reference information is, for example, the range of the shared memory space of each processing chip, and the routing information indicates the data transmission path between different processing chips. In some embodiments, the data transmission path in the routing information is associated with the chip identifier of the processing chip. It should be understood that after the switching chip configures the routing information, the switching chip in the data processing system has the ability to allow accelerators, CPU cores, etc. between processing chips to directly use addresses for access through a unified addressing method. Refer to Figure 9, which is a schematic diagram of a switching chip provided by an embodiment of the present application. As shown in Figure 9, multiple ports of the switching chip are interconnected with multiple processing chips, and routing information is configured on the routing engine of the switching chip, which can indicate the data transmission path between any two processing chips connected to the switching chip. For example, processing chip 0 is interconnected with port 0 of the switching chip, and processing chip 1 is interconnected with port 1 of the switching chip. If the switching chip receives an address access request from processing chip 1, and the address access request indicates data transmission with processing chip 1, the switching chip routes the address access request to processing chip 1 connected to port 1 based on the routing information, that is, the data in the specified address space on one processing chip is transmitted to the specified address space on another processing chip.
[0091] It should be noted that the location information provision server and the routing information configuration server are both optional devices. For example, relevant personnel can also directly configure the location information and routing information to each processing chip and switching chip. In addition, the functions provided by the location information provision server and the routing information configuration server can be implemented by a single physical device or by different physical devices, and this application is not limited to this.
[0092] Based on the data processing system introduced in Figures 1 to 9 above, the data transmission method provided by this application is introduced below.
[0093] Figure 10 is a schematic diagram of a data transmission method provided by an embodiment of the present application. As shown in Figure 10, taking the interaction between any two processing chips in a data processing system as an example, the data transmission method includes the following steps 1001 to 1005.
[0094] 1001. Multiple processing chips initialize multiple communication queues on their respective processing chips. The multiple communication queues are used to store communication messages received by the processing chips.
[0095] In the embodiment of the present application, for any processing chip, a communication engine is deployed on the port of the processing chip, and the processing chip initializes multiple communication queues through the communication engine.
[0096] Schematically, for any one processing chip, the communication engine of the processing chip initializes multiple communication queues based on the number of multiple processing chips in the data processing system. In some embodiments, the number of multiple processing chips is equal to the number of multiple communication queues. For example, if the data processing system includes 256 processing chips, the communication engine initializes 256 communication queues, and the 256 communication queues correspond one-to-one to the 256 processing chips, or one communication queue belongs to one processing chip. It should be understood that this is only an example. In some embodiments, the number of multiple communication queues may also be greater than the number of multiple processing chips. For example, if the data processing system includes 256 processing chips, the communication engine initializes 300 communication queues, and the first 256 communication queues correspond one-to-one to the 256 processing chips. In this way, when a new processing chip is added to the subsequent data processing system, the communication queues of the unassociated processing chips can be associated with the new processing chips, thereby improving processing efficiency.
[0097] In some embodiments, for any processing chip, each communication queue on the processing chip is associated with a chip identifier of each processing chip in the data processing system. Illustratively, the chip identifier of the processing chip is determined based on the physical deployment location of the processing chip in the data processing system, that is, based on the location information of the processing chip. For details about the location information of the processing chip, refer to FIG. 5 and are not further described here.
[0098] In some embodiments, for any processing chip, multiple communication queues on the processing chip are stored in a shared memory space of the processing chip. The shared memory space may be located in a storage unit of a port on the processing chip or in a memory of the processing chip, and this application does not limit this. It should be understood that the shared memory space used to store multiple communication queues is accessible to multiple processing chips in the data processing system. For example, the depth of each communication queue is a default value of 2KB. Of course, it can also be configured to other values when initializing multiple communication queues, and this application does not limit this.
[0099] Schematically, with reference to FIG11, FIG11 is a schematic diagram of a communication queue provided in an embodiment of the present application. As shown in FIG11, taking a data processing system including 256 processing chips as an example, there are 256 communication queues on each processing chip, and each communication queue corresponds to each processing chip one-to-one. The depth of each communication queue is a default value of 2KB, and the shared memory space occupied by the 256 communication queues is 512KB. Taking any two processing chips shown in FIG11 as an example, after initializing multiple communication queues, if processing chip 1 requests data transmission with processing chip 2, processing chip 1 writes a communication message to the communication queue 1 corresponding to processing chip 1 on processing chip 2. In this way, when processing chip 2 reads the communication message from the communication queue 1, it can perform data transmission with processing chip 1 based on the communication message. That is, when a processing chip in the data processing system requests data transmission with other processing chips, the communication message can be written to the communication queue belonging to the processing chip on the other processing chip; when a processing chip receives a communication message from other processing chips, it reads the corresponding communication message from the communication queue belonging to the other processing chip on this chip.
[0100] After the above step 1001, each processing chip in the data processing system initializes multiple communication queues on its own processing chip. Since the multiple communication queues are used to store communication messages from other processing chips, these multiple communication queues provide technical support for point-to-point communication between the processing chips.
[0101] 1002. A plurality of processing chips initialize a shared configuration queue on each processing chip. The shared configuration queue is used to store a mapping relationship between a shared memory space of the processing chip and a chip memory address of the processing chip.
[0102] In an embodiment of the present application, for any processing chip, a communication engine is deployed on the port of the processing chip. When multiple processing chips in a data processing system provide shared memory space, the processing chip initializes the shared configuration queue of the processing chip through the communication engine.
[0103] Schematically, for any processing chip, the communication engine of the processing chip initializes a shared configuration queue based on the shared memory space corresponding to the processing chip in the data processing system. In some embodiments, the shared memory spaces of multiple processing chips in the data processing system are evenly distributed based on the target shared memory space of the data processing system. For example, the target shared memory space of the data processing system is 128TB, which supports 256 processing chips to be interconnected through a unified addressing method. The target shared memory space is divided into 256 equal parts, and the shared memory space of each processing chip that can be accessed by other processing chips is 512GB. Based on this, the communication engine of the processing chip establishes a mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip based on the address translation unit (ATU), and stores it in the shared configuration queue.
[0104] In some embodiments, for any processing chip, the shared configuration queue on the processing chip is associated with the chip ID of the processing chip. Illustratively, the chip ID of the processing chip is determined based on the physical deployment location of the processing chip in the data processing system, that is, based on the location information of the processing chip. For details about the location information of the processing chip, refer to FIG. 5 and are not further described here.
[0105] In some embodiments, for any processing chip, the shared configuration queue on the processing chip is stored in the shared memory space of the processing chip. The shared memory space can be located in a storage unit of a port on the processing chip or in the memory of the processing chip. This application does not limit this.
[0106] Schematically, referring to Figure 12, Figure 12 is a schematic diagram of a shared configuration queue provided by an embodiment of the present application. As shown in Figure 12, taking the data processing system including 256 processing chips as an example, each processing chip has a shared configuration queue corresponding to the processing chip. Taking the processing chip 2 shown in Figure 12 as an example, its communication engine is based on ATU, establishes a mapping relationship between the shared memory space of the processing chip 2 and the chip memory address of the processing chip 2, and stores it in the shared configuration queue. In this way, if other processing chips request to access the chip memory address of the processing chip 2, it can be achieved by accessing the shared memory space 2 of the processing chip 2, thereby improving data transmission efficiency.
[0107] After the above step 1002, each processing chip in the data processing system initializes the shared configuration queue on its own processing chip. Since the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip, the shared configuration queue provides technical support for data transmission between each processing chip based on address access requests.
[0108] The following describes the process of data transmission between any two processing chips, using the interaction between any two processing chips (a first processing chip and a second processing chip) as an example, in conjunction with steps 1003 to 1005. The first processing chip and the second processing chip may be deployed in the same server, or in different servers, which is not limited in this application.
[0109] 1003. The first processing chip writes a first communication message to a first communication queue on the second processing chip. The first communication message carries a chip memory address where data to be transmitted between the first processing chip and the second processing chip is located. The first communication queue is used to store the communication message of the first processing chip.
[0110] In an embodiment of the present application, the second processing chip has multiple communication queues, wherein the first communication queue is used to store communication messages of the first processing chip. In some embodiments, the first processing chip writes the first communication message to the first communication queue associated with the first chip identifier, where the first chip identifier is used to identify the first processing chip.
[0111] The first communication message carries the chip memory address of the data to be transmitted between the first processing chip and the second processing chip, including the following situations: First, the first processing chip requests to send the data on the first processing chip to the second processing chip (or there is data on the first processing chip that needs to be obtained by the second processing chip), then the chip memory address is the chip memory address on the first processing chip, that is, the data to be transmitted refers to the data to be transmitted from the first processing chip to the second processing chip. Second, the first processing chip requests to read data from the second processing chip, then the chip memory address is the chip memory address on the second processing chip. That is, the data to be transmitted refers to the data to be transmitted from the second processing chip to the first processing chip. In other words, this application does not limit the data transmission method indicated by the communication message. It can be that the source processing chip sends data to the destination processing chip, or it can be that the source processing chip obtains data from the destination processing chip.
[0112] In some embodiments, before the first processing chip writes the first communication message to the first communication queue on the second processing chip, the first processing chip writes a status message to the first communication queue on the second processing chip, the status message indicating the chip status of the first processing chip. Wherein, the chip status refers to the working status of the processing chip, for example, the working status is available or unavailable (such as due to an abnormality, restart, or power failure, causing the processing chip to fail to operate normally). Schematically, the communication engine of the first processing chip writes the status message to the first communication queue. In some embodiments, the first processing chip writes the status message to the queue head of the first communication queue, which is not limited in this application. In addition, when initializing multiple communication queues, the first processing chip can write a first status message to the first communication queue on the second processing chip, and the first status message indicates that the chip status of the first processing chip is available (for example, marked as active). When the first processing chip has an abnormality, restarts, or is powered off, a second status message is written to the first communication queue on the second processing chip, and the second status message indicates that the chip status of the first processing chip is unavailable (for example, marked as inactive). In some embodiments, the status message can also carry other content, such as the available time period of the processing chip, the reason for unavailability, etc., which is not limited in this application.
[0113] In addition, the above is introduced by taking the example of the first processing chip writing a status message to the first communication queue on the second processing chip. In some instances, the first processing chip can write a status message to the first communication queue on all processing chips in the data processing system, or write a status message to the first communication queue on other processing chips in the system except the first processing chip. Schematically, with reference to Figure 13, Figure 13 is a schematic diagram of a method of writing a status message to a communication queue provided by an embodiment of the present application. As shown in Figure 13, the first processing chip (such as processing chip 1) writes a status message to the first communication queue on other processing chips in the system except the first processing chip, for example, writes a status message to the queue head of the first communication queue. It should be understood that the first processing chip can update the status message in the first communication queue on other processing chips in a timely manner according to the chip status of the first processing chip, so that other processing chips can promptly know the chip status of the first processing chip. It should be noted that any processing chip in the data processing system can implement the functions possessed by the first processing chip, which will not be repeated here.
[0114] 1004. The second processing chip reads the first communication message from the first communication queue.
[0115] In the embodiments of the present application, after the first processing chip writes the first communication message to the first communication queue on the second processing chip, the second processing chip can use an interrupt mechanism and / or a polling mechanism to read the first communication message, which is not limited in this application. In other words, the communication mechanism provided by this application supports multiple message reading methods and can be flexibly used according to actual needs.
[0116] Taking the interrupt mechanism as an example, after the first processing chip writes a first communication message to the first communication queue on the second processing chip, it sends an interrupt message to the second processing chip. The second processing chip receives the interrupt message and reads the first communication message from the first communication queue. By using the interrupt mechanism, the processing chip can be notified of the communication message in a timely manner, saving data transmission time.
[0117] Taking the polling mechanism as an example, the second processing chip polls and reads multiple communication queues on the second processing chip. If it finds a communication message stored in the first communication queue, it reads the first communication message from the first communication queue. Illustratively, the communication engine of the second processing chip reads the first communication message from the first communication queue. By adopting a polling mechanism, communication resources between different processing chips can be conserved.
[0118] 1005. The second processing chip transmits data with the first processing chip based on the first communication message.
[0119] In the embodiment of the present application, since the first communication message carries the chip memory address of the data to be transmitted, the second processing chip can transmit data to the first processing chip based on the chip memory address carried in the first communication message. In some embodiments, when multiple processing chips in the data processing system provide shared memory space, this step includes the following situations:
[0120] The first type of data to be transmitted refers to data to be transferred from a first processing chip to a second processing chip. In an illustrative example, the second processing chip sends a data acquisition request to the first processing chip based on a first communication message. The first processing chip determines the shared memory space corresponding to the data to be transmitted based on the data acquisition request and a shared configuration queue on the first processing chip. Based on the shared memory space corresponding to the data to be transmitted, the first processing chip transfers the data to be transmitted to the second processing chip.
[0121] The second type of data to be transmitted refers to data to be transmitted from the second processing chip to the first processing chip. In an illustrative embodiment, the second processing chip determines the shared memory space corresponding to the data to be transmitted based on the first communication message and the shared configuration queue on the second processing chip, and transmits the data to be transmitted to the first processing chip based on the shared memory space corresponding to the data to be transmitted.
[0122] Through the above method, for several different situations of data to be transmitted, the implementation method of data transmission based on shared configuration queues is introduced respectively. In this way, when multiple processing chips in the data processing system provide shared memory space, it is also possible to configure a flexible communication mechanism as described above, thereby realizing efficient data transmission between any two processing chips.
[0123] In some embodiments, if a first processing chip and a second processing chip are interconnected via a chip bus, the second processing chip sends an address access request to the first processing chip via the chip bus based on the first communication message, and the address access request indicates data transmission between the first processing chip and the chip bus. Based on the foregoing description, it can be seen that the chip bus is, for example, NvLink, CXL, UCIe, HCCS, CCIX, etc. Since the chip bus allows memory access requests issued by the acceleration unit in the acceleration chip or the CPU core in the general-purpose CPU to be directly transmitted, the memory access requests passing through the bus do not need to be converted into a request format or encapsulated into a request of another format. Therefore, this method can effectively improve the efficiency of data transmission between different processing chips.
[0124] In some embodiments, if a first processing chip and a second processing chip are interconnected via a chip bus and a switch chip, the second processing chip, based on a first communication message, sends an address access request to the first processing chip via the chip bus and the switch chip. The address access request indicates data transmission between the first processing chip and the switch chip. The switch chip contains routing information that indicates a data transmission path between the first processing chip and the second processing chip. For details about routing information, please refer to the previous description and will not be repeated here.
[0125] 14 and 15 , the data transmission method described in steps 1003 to 1005 will be described below by taking the case where the second processing chip reads the first communication message based on an interrupt mechanism or a polling mechanism as an example.
[0126] FIG14 is a process diagram of a data transmission method provided by an embodiment of the present application. As shown in FIG14, taking the first processing chip (processing chip 1) requesting to send data on the first processing chip to the second processing chip (processing chip 2) as an example (or in other words, there is data on the first processing chip that needs to be obtained by the second processing chip), the communication engine of the first processing chip writes a first communication message to the first communication queue (communication queue 1) on the second processing chip, and the first communication message carries the chip memory address of the data to be transmitted on the first processing chip, such as the PA0 address. The communication engine of the first processing chip sends an interrupt message to the second processing chip, and the communication engine of the second processing chip receives the interrupt message, reads the first communication message from the first communication queue, obtains the chip memory address of the data to be transmitted, and sends it to the processor on the second processing chip. The processor sends an address access request to the first processing chip based on the chip memory address of the data to be transmitted through the exchange chip, that is, accesses the shared memory space 1 of the first processing chip, so that the first processing chip determines the shared memory space corresponding to the data to be transmitted based on the address access request and the shared configuration queue, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted. It should be understood that the process of the first processing chip requesting to read data (for example, address PA1) from the second processing chip is the same as the above process, so it will not be described in detail.
[0127] FIG15 is a process diagram of another data transmission method provided by an embodiment of the present application. As shown in FIG15, taking the first processing chip (processing chip 1) requesting to send data on the first processing chip to the second processing chip (processing chip 2) as an example (or in other words, there is data on the first processing chip that needs to be obtained by the second processing chip), the communication engine of the first processing chip writes a first communication message to the first communication queue (communication queue 1) on the second processing chip, and the first communication message carries the chip memory address of the data to be transmitted on the first processing chip, such as the PA0 address. The communication engine of the second processing chip polls and reads multiple communication queues on the second processing chip. If a communication message is read that the first communication queue stores a communication message, the first communication message is read from the first communication queue, and the chip memory address of the data to be transmitted is obtained therefrom, and the chip memory address is sent to the processor on the second processing chip. The processor sends an address access request to the first processing chip based on the chip memory address of the data to be transmitted through the exchange chip, that is, accesses the shared memory space 1 of the first processing chip, so that the first processing chip determines the shared memory space corresponding to the data to be transmitted based on the address access request and the shared configuration queue, and transmits the data to be transmitted to the second processing chip based on the shared memory space corresponding to the data to be transmitted. It should be understood that the process of the first processing chip requesting to read data (for example, address PA1) from the second processing chip is the same as the above process, so it will not be described in detail.
[0128] In addition, the above embodiment is introduced by taking the example of the first processing chip writing a communication message to the second processing chip. Similarly, the first processing chip can also read the communication message from the second processing chip from the local communication queue and transmit data with the second processing chip, which will not be repeated here.
[0129] Below, in conjunction with the aforementioned Figure 5, the overall process of the data transmission method provided by the embodiment of the present application is illustrated. As shown in Figure 5, the data processing system includes a plurality of processing chips, and each processing chip is interconnected with a switching chip using a chip bus (or may not be interconnected through a switching chip), and there is a communication engine on the port of each processing chip, and the communication engine is used to operate multiple communication queues and shared configuration queues. It should be understood that the server is not marked in the figure, but the multiple processing chips shown in the figure can be deployed in the same server or in different servers respectively, and this application does not limit this. In some embodiments, the data processing system also includes a location information providing server and a routing information configuration server, wherein the location information providing server is used to provide location information for each processing chip in the data processing system, and the routing information configuration server is used to provide routing information for the switching chip. In the data processing system shown in Figure 5, the communication engine, the location information providing server, and the routing information configuration server on the processing chip cooperate with each other to establish a flexible communication mechanism between any processing chips, so that data can be transmitted between any processing chips through address access requests, thereby improving the data transmission efficiency between different processing chips. Moreover, this communication mechanism does not require the use of heavy communication state machine management or complex message encapsulation. Instead, it uses the chip identification of each processing chip and the communication queue on each processing chip to realize data transmission between different processing chips in the system, thereby achieving fast data access and increasing data transmission load, making data transmission between different processing chips more efficient.
[0130] In summary, the data transmission method provided in the embodiment of the present application is applied to a data processing system, in which a plurality of processing chips are distributed on a plurality of servers of the system, and a plurality of communication queues are provided on each processing chip for storing communication messages received by the processing chip. Based on this, when data is transmitted between any two processing chips, the first processing chip writes a first communication message to the first communication queue corresponding to the first processing chip on the second processing chip, so that the second processing chip transmits data with the first processing chip based on the first communication message. Through this flexible communication mechanism, data transmission is realized quickly between any two processing chips in the system, thereby improving the data transmission efficiency between different processing chips. Moreover, even between processing chips deployed on different servers, data transmission can be performed in this way, effectively improving the data transmission efficiency.
[0131] In addition, the present application also provides a data transmission device, which is configured on a processing chip and can implement the steps performed by the processing chip in the above method embodiment. Schematically, the device includes at least one functional unit, and the at least one functional unit is used to implement the steps performed by the processing chip in the above method embodiment. For example, the at least one functional unit includes a communication message writing unit and a data transmission unit, wherein the communication message writing unit is used to write communication messages to the communication queue on other processing chips, and the data transmission unit is used to transmit data between other processing chips based on the communication messages read from the communication queue. It should be understood that when the data transmission device performs data transmission, it only uses the division of the above-mentioned functional units as an example. In actual applications, the above-mentioned functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional units to complete all or part of the functions described above. In addition, the data transmission device provided in the above embodiment and the data transmission method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0132] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items having substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there a limit on quantity and execution order. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various described examples, a first processing chip may be referred to as a second processing chip, and similarly, a second processing chip may be referred to as a first processing chip. Both the first processing chip and the second processing chip may be processing chips, and in some cases, may be separate and different processing chips.
[0133] In this application, the term "at least one" means one or more, and the term "plurality" means two or more. For example, a plurality of processing chips means two or more processing chips.
[0134] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0135] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of program structure information. The program structure information includes one or more program instructions. When the program instructions are loaded and executed on a computing device, all or part of the processes or functions described in the embodiments of the present application are generated.
[0136] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0137] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data transmission method, characterized in that, Applied to a data processing system, the data processing system includes a plurality of servers, a plurality of processing chips are distributed on the plurality of servers, the processing chips are interconnected with each other, each processing chip has a plurality of communication queues, and the plurality of communication queues are used to store communication messages received by the processing chip. The method includes: The first processing chip writes a first communication message to a first communication queue on the second processing chip. The first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located. The first communication queue is used to store the communication messages of the first processing chip. Based on the first communication message, the second processing chip performs data transmission with the first processing chip.
2. The method according to claim 1, wherein There is a communication engine on the port of each processing chip, and the communication engine is used to operate the plurality of communication queues on the processing chip.
3. The method according to claim 1 or 2, characterized in that, Each processing chip also has a shared configuration queue, and the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
4. The method according to claim 3, wherein If the data to be transmitted refers to the data to be transmitted from the first processing chip to the second processing chip, the second processing chip performs data transmission with the first processing chip based on the first communication message, including: The second processing chip sends a data acquisition request to the first processing chip based on the first communication message. Based on the data acquisition request and the shared configuration queue on the first processing chip, the first processing chip determines the shared memory space corresponding to the data to be transmitted, and based on the shared memory space corresponding to the data to be transmitted, transmits the data to be transmitted to the second processing chip.
5. The method according to claim 3, wherein If the data to be transmitted refers to the data to be transmitted from the second processing chip to the first processing chip, the second processing chip performs data transmission with the first processing chip based on the first communication message, including: Based on the first communication message and the shared configuration queue on the second processing chip, the second processing chip determines the shared memory space corresponding to the data to be transmitted, and based on the shared memory space corresponding to the data to be transmitted, transmits the data to be transmitted to the first processing chip.
6. The method according to any one of claims 1 to 5, characterized in that, The first processing chip writing the first communication message to the first communication queue on the second processing chip includes: The first processing chip writes the first communication message to the first communication queue associated with the first chip identifier, and the first chip identifier is used to identify the first processing chip.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: The first processing chip writes a status message to the first communication queue, and the status message indicates the chip status of the first processing chip.
8. The method according to claim 7, characterized in that, The first processing chip writing the status message to the first communication queue includes: The first processing chip writes the status message to the queue head of the first communication queue.
9. The method according to any one of claims 1 to 8, characterized in that, The first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers.
10. The method according to any one of claims 1 to 9, characterized in that The first processing chip and the second processing chip are interconnected through a switching chip, and there is routing information on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip.
11. The method according to any one of claims 1 to 10, characterized in that, The first processing chip and the second processing chip are interconnected through a chip bus, and the second processing chip performs data transmission with the first processing chip based on the first communication message, including: Based on the first communication message, the second processing chip sends an address access request to the first processing chip through the chip bus, and the address access request indicates data transmission with the first processing chip.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: The second processing chip reads the first communication message by using an interrupt mechanism and / or a polling mechanism.
13. A data transmission system, characterized in that, The data processing system includes multiple servers, multiple processing chips are distributed on the multiple servers, each processing chip is interconnected with each other, and there are multiple communication queues on each processing chip, and the multiple communication queues are used to store the communication messages received by the processing chip; A first processing chip is configured to write a first communication message to a first communication queue on a second processing chip, the first communication message carries the chip memory address where the data to be transmitted between the first processing chip and the second processing chip is located, and the first communication queue is used to store the communication messages of the first processing chip; The second processing chip is configured to perform data transmission with the first processing chip based on the first communication message.
14. The system according to claim 13, wherein, There is a communication engine on the port of each processing chip, and the communication engine is used to operate the multiple communication queues on the processing chip.
15. The system according to claim 13 or 14, characterized in that, There is also a shared configuration queue on each processing chip, and the shared configuration queue is used to store the mapping relationship between the shared memory space of the processing chip and the chip memory address of the processing chip.
16. The system according to any one of claims 13 to 15, characterized in that, The first processing chip and the second processing chip are deployed in the same server, or the first processing chip and the second processing chip are respectively deployed in different servers.
17. The system according to any one of claims 13 to 16, characterized in that, The first processing chip and the second processing chip are interconnected through a switching chip, and routing information is configured on the switching chip, and the routing information indicates the data transmission path between the first processing chip and the second processing chip.
18. The system according to any one of claims 13 to 17, characterized in that, The first processing chip and the second processing chip are interconnected through a chip bus; The second processing chip is configured to send an address access request to the first processing chip through the chip bus based on the first communication message, and the address access request indicates data transmission with the first processing chip.
19. A processing chip, characterized in that, The processing chip includes a port, the port is used to be interconnected with other processing chips outside the processing chip, and the processing chip is used to implement the functions of the processing chip in the data transmission method according to any one of the preceding claims 1 to 12.
20. A server, characterized in that, The server includes at least one processing chip, and the processing chip is used to implement the functions of the processing chip in the data transmission method according to any one of the preceding claims 1 to 12.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one program code, and the at least one program code is used to implement the data transmission method described in any one of the foregoing claims 1 to 12.
22. A computer program product, characterized in that, When the computer program product runs on the processing chip, the processing chip is caused to implement the functions of the processing chip in the data transmission method described in any one of the foregoing claims 1 to 12.
Citation Information
Patent Citations
Data transmission method, data processing system, processing chip and server
CN120216220A
Multi-core processor interactive bus design method based on shared memory
CN108959149A
Method and device for data communication among multiple processors
CN110457251A
Shared memory data calling method and device, electronic equipment and storage medium
CN114090289A
Inter-core communication system of multi-core system
CN115080277A