A chip, data moving method and device
Patent Information
- Application Number
- CN202610688745.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-09-25
AI Technical Summary
上述方法一需要占用大量计算引擎资源,导致计算效率下降;方法二与方法三需CPU与芯片频繁同步,增加通信开销与延迟,难以满足超节点对低延迟、高带宽的通信需求
[0018]本发明实施例的技术方案通过内存模块存储数据,计算模块执行控制指令时,如果控制指令为数据搬移指令,将所述数据搬移指令发送给命令队列,数据搬移模块用于根据命令队列中的数据搬移指令在内存模块与外部存储设备之间执行数据搬移操作。由此,可以降低计算模块的功耗,满足低延迟、高带宽的通信需求。
Smart Images

Figure CN122816699A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chip technology, and more particularly to a chip, a data transfer method, and an apparatus. Background Technology
[0002] With the explosive growth in the scale of large model parameters, AI (Artificial Intelligence) supernode servers need to achieve multi-chip collaboration through high-speed inter-chip interconnects, and their communication efficiency directly determines the model training and inference performance. Existing technologies mainly employ three solutions: first, using the Load / Store instructions of the computing engine to move data; second, issuing commands to the communication control module via the CPU (Central Processing Unit); and third, relying on an external RDMA network card for transmission. The first method requires a large amount of computing engine resources, leading to a decrease in computing efficiency; methods two and three require frequent synchronization between the CPU and chips, increasing communication overhead and latency, making it difficult to meet the supernode's requirements for low latency and high bandwidth communication. Summary of the Invention
[0003] In view of this, the purpose of this invention is to provide a chip, a data transfer method and apparatus that can reduce the power consumption of the computing module and meet the communication requirements of low latency and high bandwidth.
[0004] In a first aspect, embodiments of the present invention provide a chip, the chip comprising: Memory modules are used to store data; The communication control module includes a command queue and a data transfer module; At least one computing module is configured to acquire control instructions, and in response to the control instructions being data transfer instructions, send the data transfer instructions to the command queue, wherein the data transfer instructions include at least an operation type, a source address, and a destination address; The data transfer module is used to perform data transfer operations between the memory module and the external storage device according to the data transfer instructions.
[0005] In some embodiments, the data migration instruction further includes at least one of data size, atomic operation type, and the shape of the memory access arrangement.
[0006] In some embodiments, the memory module is a high-bandwidth memory module.
[0007] In some embodiments, the chip further includes: The on-chip network is communicatively connected to the memory module, communication control module, and computing module, respectively.
[0008] In some embodiments, the communication control module further includes: The instruction parsing module is used to read data transfer instructions from the command queue and parse the data transfer instructions to obtain instruction information.
[0009] In some embodiments, the data transfer module is used to read the data to be transferred according to the source address and write the data to be transferred to the destination address.
[0010] In some embodiments, the communication control module further includes: A packet-level processing module is used to generate data packets according to a predetermined data packet format based on the data to be moved when sending data to the external storage device; and to convert the data packets into data to be moved when receiving data packets sent by the external storage device.
[0011] In some embodiments, the computing module is also configured to pause the thread of the currently processed data transfer instruction.
[0012] In some embodiments, the communication control module is further configured to send a communication completion signal to the computing module in response to the completion of the data transfer operation; The computing module is also used to resume the thread executing the currently processed data transfer instruction in response to receiving a communication completion signal.
[0013] In some embodiments, the chip further includes: At least one high-speed interface for communicating with the external storage device.
[0014] Secondly, embodiments of the present invention provide a data migration method, the method comprising: Obtain control commands; In response to the control command being a data transfer command, the data transfer command is sent to the command queue, and the data transfer command includes at least the operation type, source address, and destination address; The data transfer operation is performed between the memory module and the external storage device according to the data transfer instruction.
[0015] Thirdly, embodiments of the present invention provide a data transfer device, the device comprising: The instruction acquisition unit is used to acquire control instructions; The instruction sending unit, in response to the control instruction being a data transfer instruction, sends the data transfer instruction to the command queue, wherein the data transfer instruction includes at least an operation type, a source address, and a destination address; The data transfer unit is used to perform data transfer operations between the memory module and the external storage device according to the data transfer instructions.
[0016] Fourthly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the second aspect.
[0017] Fifthly, embodiments of the present invention provide a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the method described in the second aspect.
[0018] The technical solution of this invention stores data through a memory module. When the computing module executes a control instruction, if the control instruction is a data transfer instruction, it sends the data transfer instruction to a command queue. The data transfer module then performs a data transfer operation between the memory module and an external storage device according to the data transfer instruction in the command queue. This reduces the power consumption of the computing module and meets the requirements for low latency and high bandwidth communication. Attached Figure Description
[0019] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a schematic diagram of a chip system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the Load instruction for inter-chip communication in the first comparison. Figure 3 This is a schematic diagram of the Store instructions for inter-chip communication in the first comparison. Figure 4 This is a schematic diagram of inter-chip communication in the second comparison. Figure 5 This is a schematic diagram of inter-chip communication in the third parallel model; Figure 6 This is a schematic diagram of inter-chip communication according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the communication control module according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the format of the data transfer instruction according to an embodiment of the present invention; Figure 9 This is a flowchart of the data transfer method according to an embodiment of the present invention; Figure 10 This is a schematic diagram of a data transfer device according to an embodiment of the present invention; Figure 11 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0020] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0021] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0022] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0023] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0024] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0025] With the development of artificial intelligence technology, large-scale models have become a core engine driving industry progress. As the scale of model parameters rapidly increases, both model training and inference pose greater challenges to computing power. To address this challenge, data center AI supernode servers have emerged. These servers utilize high-speed interconnect technology to connect multiple AI chips into a powerful computing node with high bandwidth, low latency, and unified memory addressing, becoming a key infrastructure supporting the operation of trillion-parameter large-scale models. In this architecture, inter-chip interconnect communication technology, as the central link connecting the various AI chips, directly determines the computing efficiency and scalability of the entire supernode, making it one of the core technologies in the current AI computing power field.
[0026] Inter-chip interconnect communication refers to the technology that enables efficient data communication and interaction between the same or different AI chips (such as CPUs, GPUs, and NPUs) in electronic devices or computing systems. It is responsible for building high-speed data transmission bridges between different chips, resolving differences in electrical characteristics and communication bottlenecks between chips. By providing ultra-high bandwidth and ultra-low latency physical channels, this technology ensures that multiple chips can work collaboratively, quickly exchanging intermediate computation results and synchronization control information. It is a key foundation for supporting multi-card parallelism and computing power aggregation in complex tasks such as AI large-scale model training and high-performance computing. A supernode is a large-scale computing unit built in a data center using high-speed interconnect technology. It tightly integrates dozens to hundreds of AI chips, forming a logical whole with ultra-high interconnect bandwidth, nanosecond-level low latency, and unified memory addressing. Supernodes break down the physical boundaries of traditional servers, enabling a large number of chips to work collaboratively like a super chip, effectively solving the communication wall bottleneck in AI large-scale model training and inference, and improving the efficiency and cost-effectiveness of distributed parallel computing.
[0027] Figure 1 This is a schematic diagram of a chip system according to an embodiment of the present invention. Figure 1 As shown, the chip system of this embodiment includes a chip 1, a CPU, a network card 2, and an external storage device 3.
[0028] Chip 1 is an artificial intelligence chip, a highly integrated AI computing core that serves as the primary source of computing power for the entire system. Chip 1 interconnects with other chips or storage devices, providing efficient data storage, instruction execution, internal data interaction, and communication scheduling capabilities. Chip 1 works closely with the CPU, network interface card 2, and external storage device 3 to provide robust computing power and high-speed data throughput in large-scale distributed computing and AI model training.
[0029] Furthermore, chip 1 includes a memory module 11, a computing module, a communication control module 13, and an on-chip network 14.
[0030] The memory module 11 is used to store data. Specifically, the memory module 11 can be a high-bandwidth memory (HBM) module. The HBM module adopts 3D stacking technology, which breaks through the bandwidth bottleneck of traditional memory and can provide the computing engine with extremely high data throughput. By significantly shortening the physical distance between data and computing cores, the HBM module effectively reduces data access latency and is a key foundation for supporting high-performance computing tasks such as large AI models.
[0031] The computing module, the computing engine within chip 1, is the core unit that generates actual computing power, responsible for parsing and executing various computing instructions. By executing specific instruction sets, the computing module efficiently completes complex mathematical operations, logical judgments, and related memory access operations. There can be one or more computing modules; this embodiment uses n modules as an example, as shown in figures 12a-12n.
[0032] The communication control module 13 is the core hub responsible for scheduling and controlling the external communication tasks of chip 1. It can accurately map and schedule the communication needs generated by internal computing to the high-speed interface, ensuring that data can be transmitted to the external network efficiently and accurately. Through the coordination of the communication control module 13, chip 1 can directly interconnect with external storage devices at high speed, realizing large-scale data exchange and computing power aggregation across devices. It is a key bridge supporting distributed computing and cluster collaboration.
[0033] The on-chip network 14 is a communication module within the chip responsible for efficient data transmission and interaction between various functional units. Through a distributed and modular routing and switching mechanism, it tightly connects the computing module, memory module, communication control module, and other internal modules. The on-chip network ensures accurate data flow between components within the chip with low latency and high bandwidth, effectively solving the communication bottleneck between multi-core and heterogeneous computing units and guaranteeing the smooth operation of the entire chip. Specifically, the on-chip network 14 is communicatively connected to the memory module 11, communication control module 13, and computing module, providing a communication path for any two of these modules. That is, the on-chip network 14 can provide a communication path between the memory module 11 and communication control module 13, between the memory module 11 and computing module, and between the communication control module and computing module.
[0034] Furthermore, chip 1 also includes at least one high-speed interface 15 for communicating with the external storage device 3. A high-speed interface is a general term for the physical connection channel and supporting communication protocol that enables fast and reliable data transmission between chips, devices, or systems. High-speed interfaces overcome the bandwidth limitations of traditional low-speed interfaces, enabling accurate, low-loss, and low-interference transmission of digital information between the sending and receiving ends at high data rates.
[0035] The CPU is the main control chip of the entire system, undertaking the core responsibilities of system management and control. The CPU is connected to chip 1 through the PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard), and is responsible for issuing control commands to chip 1 in the system, coordinating and driving the orderly operation of the entire computing system.
[0036] Network interface card 2 (NIC 2) is a dedicated hardware component for managing network data transmission. NIC 2 connects to chip 1 via a PCIe interface and provides a high-speed network interface. Furthermore, NIC 2 is an RDMA (Remote Direct Memory Access) NIC. Through RDMA technology, it can achieve efficient, low-latency data transfer and communication with other nodes or external storage devices without consuming CPU resources.
[0037] External storage device 3 is the other end that enables inter-chip interconnection communication with chip 1, and is an important hardware component used by the system for persistent data storage. External storage device 3 can be a chip or other types of storage devices.
[0038] In summary, in such Figure 1 In the system shown, chip 1 and external storage device 3 form an inter-chip communication interconnect. The CPU and network card 2 are connected to chip 1 via a PCIe interface. The CPU is the main control chip of the system and is responsible for issuing commands; network card 2 is responsible for network transmission and has a high-speed network interface.
[0039] There are three main types of existing inter-chip communication schemes: The first type is based on the software Load / Store instruction mechanism. When using Load and Store instructions for complete inter-chip communication, two steps are required. The first step is that the computing module uses the Load instruction to load data from the memory module into the computing module's internal registers. The second step is that the computing module uses the Store instruction to transfer data from the computing module's internal registers to the external storage device.
[0040] Figure 2 This is a schematic diagram of the Load instruction for inter-chip communication in the first proportional representation. (Example) Figure 2 As shown by the dashed line, the computing module uses the Load instruction to load data from the memory module into the computing module's internal register.
[0041] Figure 3 This is a schematic diagram of Store instructions for inter-chip communication, shown in the first example. Figure 3 As shown by the dashed line, the compute module uses the Store instruction to transfer data from the compute module's internal registers to external storage devices.
[0042] When the on-chip network detects that the destination address of the Store instruction does not belong to the memory module but to the external storage device, the on-chip network sends the Store data to the communication control module, which then transmits the data to the external storage device through a high-speed interface to complete the communication operation.
[0043] However, the above method relies on the computing module to actively load and store data to complete the transmission, which seriously consumes the computing module's resources and leads to a decrease in computing throughput.
[0044] The second method is a CPU-driven communication control module mechanism that triggers DMA to move data by issuing commands.
[0045] Figure 4 This is a schematic diagram of inter-chip communication in the second example. (For example...) Figure 4 As shown by the dashed lines, the driver running on the CPU can issue read and write commands to the command queue in the communication control module. When the communication control module receives a read or write command, it reads the data from the memory module through the internal Direct Memory Access (DMA) module and transmits it to the external storage device through a high-speed interface.
[0046] However, since the computation process is completed by the chip and the communication is completed by driver commands running on the CPU, frequent synchronization is required between computation and communication, which reduces communication efficiency.
[0047] The third method is to use an RDMA network card for communication.
[0048] Figure 5 This is a schematic diagram of inter-chip communication in the third example. (Example) Figure 5 As shown by the dashed line, when using an RDMA network card for communication, the CPU sends communication commands to the command queue in the RDMA network card. After receiving the command, the network card reads the data to be transferred from the memory module through the PCIe interface and sends it to the external storage device through the network interface.
[0049] However, this method requires frequent communication and synchronization between the network card, CPU, and chip, and is also limited by the network card bandwidth.
[0050] Therefore, embodiments of the present invention provide a chip and a data transfer method to solve the above problems.
[0051] in, Figure 6 This is a schematic diagram of inter-chip communication according to an embodiment of the present invention. Figure 6 As shown, in the chip of this embodiment of the invention, the computing module sends a data transfer instruction to the communication control module to complete the data transfer.
[0052] Specifically, Figure 7 This is a schematic diagram of the communication control module according to an embodiment of the present invention. Figure 7 The structure of the communication control module 13 is shown, along with the computing module, high-speed interface 15, and external storage device 3 directly or indirectly connected to it.
[0053] like Figure 7As shown, the communication control module 13 includes a command queue 131, an instruction parsing module 132, a data transfer module 133, and a packet-level processing module 134.
[0054] The command queue 131 is used to store data transfer instructions sent by the computing module.
[0055] Specifically, the computing module is used to acquire and process control instructions. Specifically, the computer's ISA (Instruction Set Architecture) compiler sends the compiled ISA instructions (control instructions) to the computing module for execution. These control instructions can be logical operation instructions, data operation instructions, data transfer instructions, decision jump instructions, matrix operation instructions, and other special operation instructions. The computing module stores the received control instructions in memory or a cache.
[0056] Then, the computing module acquires and processes control instructions according to the "Fetch-Decode-Execute" instruction cycle, extracts and parses instructions from the storage medium, and then completes specific data interaction and computation tasks.
[0057] During the instruction fetch phase, the Program Counter (PC) is started. The PC, acting as the core control register, constantly stores the logical address in memory of the next instruction to be executed. Driven by the system clock, the control unit within the computing module sends the address in the PC to memory or cache via the address bus, initiating a read request. Once the target binary machine code is transmitted back via the data bus and loaded into the Instruction Register (IR), the PC automatically increments, pointing to the address of the subsequent instruction, thus laying the foundation for continuous instruction stream processing.
[0058] During the instruction decoding stage, the control unit parses the binary code in the instruction register, breaking it down into opcode and operands. The opcode clarifies the functional attributes of the current instruction (such as reading data, writing data, and whether atomic operations are required), while the operands provide crucial addressing information such as base address, offset, or register index (e.g., source address, destination address, data size, atomic operation type, and memory access layout). This process transforms the abstract machine code into concrete hardware control signals.
[0059] During the instruction execution phase, the computation module activates the corresponding hardware functional units based on the decoded results. In response to the currently processed control instruction, which is a data transfer instruction, the data transfer instruction is sent to the command queue in the communication control module.
[0060] The data transfer instruction includes at least one of the following: operation type, source address, destination address, data size, atomic operation type, and the shape of the memory access arrangement.
[0061] Figure 8 This is a schematic diagram illustrating the format of a data transfer instruction according to an embodiment of the present invention. For example... Figure 8 As shown, the data migration instructions include: The OP Code (Operation Code) field indicates the operation to be performed, such as a read operation, a write operation, and whether an atomic operation is required. A read operation reads data from an external storage device and stores it in the chip's internal memory module. In this case, the source address is the address of the external storage device, and the destination address is the address of the memory module. A write operation reads data from the chip's internal memory module and sends it to an external storage device. In this case, the source address is the address of the memory module, and the destination address is the address of the external storage device. An atomic operation is an operation that cannot be interrupted and must be executed completely in one go. In multi-chip interconnect scenarios, multiple computing modules may simultaneously attempt to access the same shared memory area. If the OP Code indicates that an atomic operation (such as atomic addition, atomic comparison and swap, etc.) is required, the communication control module will lock the access request, ensuring that read and write requests from other computing modules to that address are blocked until the current operation is completely completed and the result is written back. This effectively prevents data contention and state inconsistencies caused by multi-threaded or cross-chip concurrent access, ensuring the data integrity and accuracy of distributed computing tasks.
[0062] The SRC Addr (Source Address) field is used to indicate the source address information of the data transfer.
[0063] The DES Addr (Destination Address) field is used to indicate the destination address information for data relocation.
[0064] The Size field indicates the size of the data to be moved.
[0065] The Atomic OP (Atomic Operation) field indicates the type of atomic operation, such as accumulation or swap.
[0066] The Shape field indicates the shape in which data is arranged during memory access. For example, in AI scene data storage, data can be arranged in memory according to different shapes, such as 1D, 2D, etc.
[0067] In this embodiment, the instruction parsing module 132 reads data transfer instructions from the command queue 131 and parses the data transfer instructions to obtain instruction information. The instruction information includes source address, destination address, operation type, data size, atomic operation type, and the shape of the memory access arrangement. The parsed instruction information is then sent to the data transfer module 133.
[0068] In this embodiment, the data transfer module 133 is used to perform a data transfer operation between the memory module and the external storage device according to the data transfer instruction. Further, the data transfer module is used to read the data to be transferred according to the source address, and write the data to be transferred to the destination address.
[0069] Specifically, when the operation type is a read data operation, data is read from the external storage device according to the source address, and the read data is stored in the internal memory module of the chip according to the destination address. When the operation type is a write data operation, data is read from the internal memory module of the chip according to the source address, and the read data is sent to the external storage device according to the destination address.
[0070] The computing module is further configured to pause the thread currently processing the data transfer instruction. Correspondingly, the communication control module is further configured to send a communication completion signal to the computing module in response to the completion of the data transfer operation. The computing module is further configured to resume the thread executing the currently processing data transfer instruction in response to receiving the communication completion signal.
[0071] In this embodiment, the packet-level processing module 134 is used to convert between data packets and data to be moved. Specifically, when sending data to the external storage device, a data packet is generated based on the data to be moved according to a predetermined data packet format. When receiving a data packet sent by the external storage device, the data packet is converted into data to be moved.
[0072] Specifically, since the granularity of data transfer inside the chip is different from that of data transmission on the high-speed interface, the data to be transferred is segmented by the packet-level processing module 134, packaged according to the data packet format requirements of the high-speed interface to generate data packets, and sent to multiple high-speed interfaces in an interleaved manner.
[0073] In summary, when the computing module executes a control instruction, if the control instruction is a data transfer instruction, it sends the data transfer instruction to the command queue in the communication control module. The instruction parsing module reads the data transfer instruction from the command queue and parses it to obtain instruction information. This instruction information includes the source address, destination address, operation type, data size, atomic operation type, and the shape of the memory access arrangement. The parsed instruction information is then sent to the data transfer module. The data transfer module reads the data to be transferred according to the source address and writes the data to be transferred to the destination address.
[0074] When the operation type is a read data operation, the data transfer module sends the source address to the packet-level processing module. The packet-level processing module generates a read request based on the source address and then sends the read request to the external storage device via a high-speed interface. The read request includes at least the source address so that the external storage device can obtain the data to be transferred based on the source address. The module also receives data packets of the data to be transferred from the external storage device via the high-speed interface. The packet-level processing module converts the data packets into the data to be transferred and sends them to the data transfer module. The data transfer module then stores the data to be transferred in the chip's internal memory module according to the destination address.
[0075] When the operation type is a write data operation, the data transfer module reads the data to be transferred from the memory module inside the chip according to the source address, converts the data to be transferred into a data packet through the packet-level processing module, and sends the data packet to the external storage device through the high-speed interface to write the data to be transferred into the external storage device.
[0076] Unlike the software solution described in the traditional method one, which uses a large number of Load / Store instructions to consume computing engine resources, this embodiment of the invention only requires the computing module to send a small number of communication commands to trigger the hardware to complete the actual data transfer operation. This avoids the communication operation from occupying the computing engine, leaving computing resources for the computing operation.
[0077] Unlike the traditional method two, which uses a CPU driver to issue communication commands and requires frequent communication and synchronization between the CPU and the AI chip, this embodiment of the invention uses the computing module to send data transfer instructions, avoiding the efficiency loss caused by frequent synchronization operations between the chip and the CPU in completing communication.
[0078] Unlike the traditional method three, which uses an RDMA network card to complete the operation, this method is not limited by the network interface bandwidth of the network card (the high-speed interface of the chip usually has ten times the bandwidth of the network card chip interface), and does not require frequent communication between the CPU, chip and network card, thus achieving a larger inter-chip communication bandwidth.
[0079] In addition, existing technologies also include solutions that integrate the RDMA network card (NIC) into the chip. The computing module provides data transfer instructions to the RDMA NIC, which then performs the data transfer operation. However, in this approach, before data transmission, the computing module needs to establish a Queue Pair (QP) between the source and destination nodes via the NIC to confirm the two ends of the transmission. Essentially, this follows the standard RDMA operation procedure, with transmission occurring between QP pairs in a message-based manner. Specifically, the computing module issues a Work Queue Element (WQE) to the RDMA engine's SendQueue (SQ) to indicate the content and length of the transmission. Then, it writes this information to the Doorbell register, effectively "ringing the doorbell" to the RDMA engine and informing it to proceed with the transmission. After executing the corresponding data transfer operation for WQE, the RDMA engine writes CQE (Completion Queue Element) to the Completion Queue (CQ) to confirm the completion of the transfer. The computation module repeatedly queries the status of the Completion Queue, and confirms the completion of the transfer once the CEQ has been written. This approach integrates the traditional RDMA engine into the chip while retaining the traditional RDMA message semantics delivery method and processing flow.
[0080] This invention, based on global memory semantics, allows the ISA instructions of the computing module to access the system's global address space. Data transmission does not rely on the QP pair concept of RDMA message semantics; therefore, no QP pairs need to be established, and the ISA instructions directly access the global address space. Specifically, when the computing module executes a data transfer instruction, this instruction is sent to the communication control module, which can be considered a coprocessor of the computing module, performing the data transfer operation. At this time, the thread that issued the communication instruction from the computing module is paused, while other threads execute normally. After processing the communication operation, the communication control module automatically returns a completion status to the computing module. Subsequently, as with other instruction completions, the thread paused by the communication operation will automatically execute subsequent instructions.
[0081] In summary, the embodiments of this invention are based on native memory semantics access, eliminating the need for RDMA processes, making them more suitable for AI application scenarios. While integrated RDMA engine solutions still maintain the semantic essence of RDMA messages—all messages are transmitted within QP pairs—this embodiment does not involve the concept of QP queues. Remote access does not require establishing QP queues; memory semantics are transmitted using global addresses. Furthermore, this embodiment does not employ the concept of RDMA queues such as SQ / RQ / CQ. Data transmission uses native ISA, is compiled by the compiler, and executed on the computing module. The communication control module, acting as the write processor of the computing module, automatically executes ISA instructions, without the Ring Doorbell process of the RDMA standard. Additionally, computing modules are generally multi-threaded engines. Because execution uses native ISA instructions, threads handling data transfer will automatically pause, waiting for the data transfer instructions to complete. After the communication control module completes data transfer, it automatically returns to a completion status, and the paused threads automatically resume execution. Therefore, there is no CEQ query process required by integrated RDMA engines.
[0082] In this embodiment of the invention, data is stored in a memory module. When the computing module executes a control instruction, if the control instruction is a data transfer instruction, it sends the data transfer instruction to a command queue. The data transfer module then performs a data transfer operation between the memory module and an external storage device according to the data transfer instruction in the command queue. This reduces the power consumption of the computing module and meets the requirements for low latency and high bandwidth communication.
[0083] Figure 9 This is a flowchart of a data transfer method according to an embodiment of the present invention. Figure 9 As shown, the data transfer method of this invention includes the following steps: Step S100: Obtain control commands.
[0084] Step S200: In response to the control instruction being a data transfer instruction, the data transfer instruction is sent to the command queue. The data transfer instruction includes at least the operation type, source address, and destination address.
[0085] Step S300: Perform a data transfer operation between the memory module and the external storage device according to the data transfer instruction.
[0086] In some embodiments, the data migration instruction further includes at least one of data size, atomic operation type, and the shape of the memory access arrangement.
[0087] In some embodiments, performing a data transfer operation between the memory module and the external storage device according to the data transfer instruction includes: Read data transfer instructions from the command queue; The data transfer command is parsed to obtain command information; Read the data to be moved based on the source address and write the data to be moved to the destination address.
[0088] In some embodiments, reading the data to be moved according to the source address and writing the data to be moved to the destination address includes: When sending data to the external storage device, a data packet is generated according to the data to be moved in a predetermined data packet format; Upon receiving a data packet from the external storage device, the data packet is converted into data to be moved.
[0089] In some embodiments, the method further includes: Pause the thread currently processing the data transfer instruction.
[0090] In some embodiments, the method further includes: In response to the completion of the data transfer operation, a communication completion signal is generated; Resume the thread currently executing the data transfer instruction.
[0091] In this embodiment of the invention, data is stored in a memory module. When the computing module executes a control instruction, if the control instruction is a data transfer instruction, it sends the data transfer instruction to a command queue. The data transfer module then performs a data transfer operation between the memory module and an external storage device according to the data transfer instruction in the command queue. This reduces the power consumption of the computing module and meets the requirements for low latency and high bandwidth communication.
[0092] Figure 10 This is a schematic diagram of a data transfer device according to an embodiment of the present invention. Figure 10 As shown, the data transfer device of this embodiment includes an instruction acquisition unit 91, an instruction sending unit 92, and a data transfer unit 93. The instruction acquisition unit 91 is used to acquire control instructions. The instruction sending unit 92 is used to send the data transfer instruction to a command queue in response to the control instruction being a data transfer instruction. The data transfer instruction includes at least an operation type, a source address, and a destination address. The data transfer unit 93 is used to perform a data transfer operation between the memory module and the external storage device according to the data transfer instruction.
[0093] In this embodiment of the invention, data is stored in a memory module. When the computing module executes a control instruction, if the control instruction is a data transfer instruction, it sends the data transfer instruction to a command queue. The data transfer module then performs a data transfer operation between the memory module and an external storage device according to the data transfer instruction in the command queue. This reduces the power consumption of the computing module and meets the requirements for low latency and high bandwidth communication.
[0094] Figure 11 This is a schematic diagram of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device 10 includes a chip, a computer, etc. Figure 11 As shown, the electronic device 10 includes at least one processor 101; a memory 102 communicatively connected to at least one processor 101; and a communication component 103 communicatively connected to a scanning device, wherein the communication component 103 receives and transmits data under the control of the processor 101; wherein the memory 102 stores instructions executable by at least one processor 101, the instructions being executed by at least one processor 101 to implement the above-described data transfer method.
[0095] Specifically, the electronic device includes: one or more processors 101 and a memory 102. Figure 11 Taking a processor 101 as an example, the processor 101 and the memory 102 can be connected via a bus or other means. Figure 11 Taking a bus connection as an example, memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 101 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 102, thereby realizing the above-mentioned data transfer method.
[0096] Memory 102 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 102 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 102 may optionally include memory remotely located relative to processor 101, and these remote memories may be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0097] One or more modules are stored in memory 102 and, when executed by one or more processors 101, perform the data transfer method in any of the above method embodiments.
[0098] The above-mentioned products can perform the methods provided in the embodiments of this application, and have the corresponding functional modules and beneficial effects of performing the methods. For technical details not described in detail in this embodiment, please refer to the methods provided in the embodiments of this application.
[0099] In this embodiment of the invention, data is stored in a memory module. When the computing module executes a control instruction, if the control instruction is a data transfer instruction, it sends the data transfer instruction to a command queue. The data transfer module then performs a data transfer operation between the memory module and an external storage device according to the data transfer instruction in the command queue. This reduces the power consumption of the computing module and meets the requirements for low-latency, high-bandwidth communication. Another embodiment of the invention relates to a non-volatile storage medium for storing a computer-readable program, which is used by a computer to execute some or all of the above-described method embodiments.
[0100] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0101] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A chip, characterized in that, The chip includes: Memory modules are used to store data; The communication control module includes a command queue and a data transfer module; At least one computing module is configured to acquire control instructions, and in response to the control instructions being data transfer instructions, send the data transfer instructions to the command queue, wherein the data transfer instructions include at least an operation type, a source address, and a destination address; The data transfer module is used to perform data transfer operations between the memory module and the external storage device according to the data transfer instructions.
2. The chip according to claim 1, characterized in that, The data move instructions also include at least one of the following: data size, atomic operation type, and the shape of the memory access arrangement.
3. The chip according to claim 1, characterized in that, The memory module is a high-bandwidth memory module.
4. The chip according to claim 1, characterized in that, The chip also includes: The on-chip network is communicatively connected to the memory module, communication control module, and computing module, respectively.
5. The chip according to claim 1, characterized in that, The communication control module further includes: The instruction parsing module is used to read data transfer instructions from the command queue and parse the data transfer instructions to obtain instruction information.
6. The chip according to claim 5, characterized in that, The data transfer module is used to read the data to be transferred according to the source address and write the data to be transferred to the destination address.
7. The chip according to claim 6, characterized in that, The communication control module further includes: A packet-level processing module is used to generate data packets according to a predetermined data packet format based on the data to be moved when sending data to the external storage device; and to convert the data packets into data to be moved when receiving data packets sent by the external storage device.
8. The chip according to claim 1, characterized in that, The computing module is also used to pause the thread of the currently processed data transfer instruction.
9. The chip according to claim 8, characterized in that, The communication control module is also used to send a communication completion signal to the computing module in response to the completion of the data transfer operation; The computing module is also used to resume the thread executing the currently processed data transfer instruction in response to receiving a communication completion signal.
10. The chip according to claim 1, characterized in that, The chip also includes: At least one high-speed interface for communicating with the external storage device.
11. A data migration method, characterized in that, The method includes: Obtain control commands; In response to the control command being a data transfer command, the data transfer command is sent to the command queue, and the data transfer command includes at least the operation type, source address, and destination address; The data transfer operation is performed between the memory module and the external storage device according to the data transfer instruction.
12. A data transfer device, characterized in that, The device includes: The instruction acquisition unit is used to acquire control instructions; The instruction sending unit, in response to the control instruction being a data transfer instruction, sends the data transfer instruction to the command queue, wherein the data transfer instruction includes at least an operation type, a source address, and a destination address; The data transfer unit is used to perform data transfer operations between the memory module and the external storage device according to the data transfer instructions.
13. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in claim 11.
14. A computer-readable storage medium storing computer program instructions thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method as described in claim 11.