Data transmission apparatus, data processing device, system and method, and medium
By designing an address resolution module and multiple memory access modules in the data transmission device, directly accessing the memory address of the remote device, the problem of increasing resource consumption and processing time during data transmission in the prior art is solved, and more efficient data transmission and resource utilization are achieved.
Patent Information
- Application Number
- PCT/CN2024/122631
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art requires multiple migration of data when transmitting acceleration task data between remote clients and servers, resulting in increased resource consumption and processing time.
A data transmission device is designed, including an address resolution module and multiple memory access modules. Each memory access module can directly access the memory address of the remote device and supports time-sharing multiplexing to reduce the number of data transfers.
By directly accessing the memory address of the remote device, the number of data transfers is reduced, resource consumption and accelerated task processing time is reduced, and resource utilization of processors and accelerators is improved.
Smart Images

Figure CN2024122631_05062025_PF_FP_ABST
Abstract
Description
Data transmission device, data processing equipment, system, method and medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on November 29, 2023, with application number 202311607588.2, and application name “A data transmission device, data processing equipment, system, method and medium”, all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a data transmission device, a data processing device, a system, a method and a medium. Background Art
[0004] Currently, remote clients can remotely utilize an accelerator on a server to perform acceleration tasks. Before the server accelerator executes the acceleration task, the relevant data provided by the remote client must be moved from the remote client's memory to the server's memory, and then from the server's memory to the accelerator. Furthermore, the final execution result of the acceleration task must also be moved from the accelerator to the server's memory, and then from the server's memory to the remote client's memory. This process requires multiple data transfers, increasing resource consumption and processing time.
[0005] Summary of the Invention
[0006] In a first aspect, the present application provides a data transmission device, comprising: an address resolution module and a plurality of memory access modules;
[0007] Each memory access module is used to directly connect to a memory address of a corresponding remote device based on a memory request of at least one remote device, and supports time-sharing multiplexing of different remote devices connected thereto, and each remote device directly accessed by each memory access module shares a processor connected to the data transmission device and an accelerator connected to the processor;
[0008] The address parsing module is configured to: determine, based on a received address access request, a target memory access module corresponding to a target remote device to be accessed by the processor, wherein the template memory access module has a mapping relationship with a target memory address of the target remote device; and
[0009] The target memory access module is used to: read the pending data of the acceleration task stored in the target memory address so that the accelerator connected to the processor processes the pending data; and / or store the processing result of the acceleration task output by the accelerator connected to the processor to the target memory address.
[0010] In some embodiments, each memory access module is configured with a memory address of at least one remote device;
[0011] The arbitrary memory access module is configured to: configure the memory address range carried by the address configuration operation in itself according to the address configuration operation sent by the processor, and establish a remote memory access connection with the current remote device, wherein the memory address range corresponds to the arbitrary remote device; and
[0012] The address resolution module is used to record the mapping relationship between the memory address range, the current remote device and the memory access module configured with the memory address range;
[0013] In some embodiments, the arbitrary memory access module is configured to: disconnect the remote memory access connection with the connected remote device according to the address release operation sent by the processor, so that the arbitrary memory access module can directly connect to the memory address of the other remote device;
[0014] The address resolution module is used to delete the corresponding mapping relationship.
[0015] In some embodiments, the invention further includes: a null address processing module;
[0016] The address parsing module is further configured to: in response to determining that there is no target memory access module having a mapping relationship with the target memory address, forward the address access request to the empty address processing module; and
[0017] The empty address processing module is used to construct meaningless response data for the address access request according to a preset strategy, and send the meaningless response data to the processor.
[0018] In some embodiments, further comprising: a high-speed interconnect module;
[0019] The high-speed interconnect module is used to: communicate with the processor through a high-speed interconnect interface; and
[0020] The high-speed interconnect interface includes at least:
[0021] a configuration interface for transmitting at least one of an address release operation and an address configuration operation sent by the processor; and
[0022] Access interface, used to transmit address access requests and corresponding response data.
[0023] In a second aspect, the present application provides a data processing device, comprising: a processor, and an accelerator and a data transmission device connected to the processor;
[0024] The data transmission device is used to: directly connect to a memory address of a corresponding remote device according to a memory request of at least one remote device;
[0025] The processor is configured to: in response to determining that the data transmission device is directly connected to a memory address of a remote device, generate and send an address access request to the data transmission device using an accelerator; and
[0026] The data transmission device is used to: read the target memory address of the target remote device to be accessed by the processor according to the received address access request, and the to-be-processed data of the acceleration task stored in the target memory address with a mapping relationship, so that the accelerator processes the to-be-processed data; and / or store the processing result of the acceleration task output by the accelerator to the target memory address.
[0027] In some embodiments, the processor is configured to: assign a computing request sent by any remote device to a request queue corresponding to the accelerator;
[0028] The accelerator is used to: read a computing request from a request queue, generate an address access request containing the computing request, and send the address access request to a data transmission device to obtain data to be processed; and
[0029] The accelerator is used to: process the data to be processed to obtain a processing result, generate an address access request containing the processing result, and send the address access request to the data transmission device to transmit the processing result back to the corresponding remote device.
[0030] In some embodiments, the processor is further configured to: generate an address configuration operation for an idle memory access module in the data transmission device according to a memory request sent by any remote device, and send the address configuration operation to the data transmission device;
[0031] The data transmission device is used to: enable the idle memory access module to configure the memory address range corresponding to any remote device carried by the address configuration operation in itself according to the address configuration operation, and establish a remote memory access connection with the current remote device; and
[0032] The processor is further configured to: in response to determining that no idle memory access module is found, return an application failure message to the corresponding remote device.
[0033] In some embodiments, the processor is further used to: detect the memory space size of the memory address range based on a memory request sent by any remote device; determine a memory mode that matches the memory space size; and manage the corresponding memory space according to the memory mode.
[0034] In some embodiments, the processor is further configured to: set a configurable address range size for each memory access module in the data transmission device.
[0035] In some embodiments, the data processing device includes multiple processors and multiple data transmission devices; each processor is connected to a data transmission device and multiple accelerators.
[0036] In some embodiments, the data transmission device is also used to: in response to determining that the target memory address does not exist, use the empty address processing module in itself to construct meaningless response data for the current address access request according to a preset strategy, and send the meaningless response data to the processor.
[0037] In a third aspect, the present application provides a data processing system, comprising: a plurality of remote devices, a network device and any one of the data processing devices above; the data processing device is connected to the plurality of remote devices via the network device.
[0038] In some embodiments, data is transmitted between the data processing device and the network device, and between the network device and multiple remote devices in an RDMA manner.
[0039] In a fourth aspect, the present application provides a data processing method, which is applied to a processor in a data processing device, wherein the processor is connected to an accelerator and a data transmission device; the method comprises:
[0040] Detecting whether the data transmission device is directly connected to the memory address of the remote device; and
[0041] In response to determining that the data transmission device is directly connected to the memory address of the remote device, the accelerator is used to generate and send an address access request to the data transmission device, so that the data transmission device reads the target memory address of the target remote device to be accessed by the processor according to the received address access request, and the target memory address has a mapping relationship with the target memory address of the target remote device to be accessed by the processor, and the data to be processed of the acceleration task stored in the target memory address, so that the accelerator processes the data to be processed; and / or stores the processing result of the acceleration task output by the accelerator to the target memory address.
[0042] In some embodiments, further comprising:
[0043] querying an idle memory access module in the data transmission device according to a memory request sent by any remote device;
[0044] In response to determining that an idle memory access module has been queried, an address configuration operation is generated for the idle memory access module, and the address configuration operation is sent to the idle memory access module, so that the idle memory access module configures the memory address range corresponding to any remote device carried by the address configuration operation in itself according to the address configuration operation, and establishes a remote memory access connection with the current remote device.
[0045] In some embodiments, further comprising:
[0046] Detect the memory space size of the memory address range based on the memory request sent by any remote device;
[0047] Determine the memory mode that matches the size of the memory space; and
[0048] Manage the corresponding memory space according to the memory mode.
[0049] In some embodiments, it further includes:
[0050] A configurable address range size is set for each memory access module in the data transmission device.
[0051] In some embodiments, it further includes:
[0052] Receive meaningless response data sent by the data transmission device; in response to determining that the target memory address does not exist, the data transmission device uses the empty address processing module in itself to construct meaningless response data for the current address access request according to a preset strategy.
[0053] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium for storing computer-readable instructions, wherein the computer-readable instructions implement the aforementioned disclosed data processing method when executed by one or more processors. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0055] FIG1 is a schematic structural diagram of a data transmission device disclosed in this application;
[0056] FIG2 is a schematic structural diagram of a data processing device disclosed in this application;
[0057] FIG3 is a schematic structural diagram of another data processing device disclosed in this application;
[0058] FIG4 is a schematic structural diagram of a data processing system disclosed in this application;
[0059] FIG5 is a flow chart of a data processing method disclosed in this application;
[0060] FIG6 is a schematic structural diagram of another data processing system disclosed in this application;
[0061] FIG7 is a schematic structural diagram of another data transmission device disclosed in this application;
[0062] FIG8 is a flowchart of a memory access request processing method disclosed in this application;
[0063] FIG9 is a flow chart of a memory application disclosed in this application;
[0064] FIG10 is a schematic diagram of a memory mode management scenario disclosed in this application;
[0065] FIG11 is a schematic diagram of the structure of a server provided by the present application;
[0066] FIG12 is a schematic structural diagram of a terminal provided by the present application;
[0067] FIG13 is a schematic structural diagram of a non-transitory computer-readable storage medium provided in this application. DETAILED DESCRIPTION
[0068] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other examples obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0069] Currently, before the accelerator on the server executes the acceleration task, the relevant data of the acceleration task provided by the remote client needs to be moved from the remote client memory to the server memory, and then from the server memory to the accelerator. In addition, the final execution result of the acceleration task also needs to be moved from the accelerator to the server memory, and then from the server memory to the remote client memory. It can be seen that this process requires moving a large amount of data multiple times, which increases resource consumption and acceleration task processing time. To this end, the present application provides a data processing solution that can reduce resource consumption and processing time during the execution of acceleration tasks.
[0070] As shown in FIG1 , an embodiment of the present application discloses a data transmission device comprising an address resolution module and multiple memory access modules. Each memory access module is configured to directly access a memory address segment of at least one remote device. Each memory access module is configured to apply for a direct connection to a memory segment of the corresponding remote device based on the memory of the at least one remote device and support time-sharing multiplexing of the different remote devices connected to each memory access module. The remote devices directly accessed by each memory access module share a processor connected to the data transmission device and an accelerator connected to the processor. The remote device is a remote client, which can be a server, for example. Both the address resolution module and the memory access module are hardware logic processing programs located within an integrated circuit and can be implemented using an FPGA (Field-Programmable Gate Array). The address resolution module receives a memory access request and, based on the memory access address, queries an on-chip address mapping table to locate the corresponding memory access module. The memory access module implements RDMA communication with the remote client. After obtaining the memory access address and data length, it reads or writes to the remote memory medium via RDMA (Remote Direct Memory Access).
[0071] The address parsing module is configured to determine, based on a received address access request, a target memory access module corresponding to a target remote device to be accessed by the processor, wherein the target memory access module has a mapping relationship with the target memory address of the target remote device. In some embodiments, the received address access request is parsed to obtain the target memory address of the target remote device to be accessed by the processor, and a target memory access module having a mapping relationship with the target memory address is determined. The target remote device is one or more of the remote devices directly accessed by each memory access module. The target memory address is the memory address of the target remote device configured in the target memory access module.
[0072] The target memory access module is used to: read the pending data of the acceleration task stored at the target memory address so that the accelerator connected to the processor can process the pending data; and / or store the processing results of the acceleration task output by the accelerator connected to the processor at the target memory address. The pending data of the acceleration task includes: the artificial intelligence model for processing the acceleration task and the model input data.
[0073] In some embodiments, a data transmission device with data address resolution and transmission capabilities is provided. The device is located on a server and connected to a processor of the server. Each memory access module in the device can be configured with a memory address segment of at least one remote device (the length of the segment is determined by the server).
[0074] In some embodiments, each memory access module is configured with a memory address range of at least one remote device. Any memory access module is configured to, based on an address configuration operation sent by a processor, configure within itself a memory address range corresponding to any remote device, as carried in the address configuration operation, and establish a remote memory access connection with the current remote device. The remote memory access connection may be, for example, an RDMA connection, in which case the memory access module is an RDMA module. RDMA is a high-performance, low-latency network communication technology that allows direct access to the memory of a remote computer, enabling zero-copy data transfer.
[0075] The address resolution module is used to record the mapping relationship between the memory address range, the current remote device, and the memory access module configured with the memory address range. The address resolution module can be configured with an address mapping relationship table that records the mapping relationship between the memory access module, its configured memory address range, and the IP address of the relevant remote device.
[0076] The same memory access module can be time-shared by multiple connected remote devices. Any memory access module is used to disconnect the remote memory access connection with the connected remote device based on the address release operation sent by the processor, allowing the memory access module to directly connect to the memory address of other remote devices. The address resolution module is used to delete the corresponding mapping relationship.
[0077] The memory address range is flexibly set by the processor for each memory access module. This refers to the configurable address range size for each memory access module and does not represent address information. The memory address ranges (configurable address range sizes) corresponding to different memory access modules can be equal or different. The sum of the memory address ranges corresponding to the memory access modules is configured by the processor. For example, if the processor allocates 1TB of memory to a data transmission device, the memory access modules in the data transmission device will divide up this 1TB of memory.
[0078] When any remote device sends a memory request to the processor, the processor searches for an idle memory access module in the data transmission device based on the memory request sent by the remote device. In response to determining that an idle memory access module has been found, the processor generates an address configuration operation for the idle memory access module according to the configurable address range size set for each memory access module (the address configuration operation is used to establish a binding relationship between the current remote device and the idle memory access module), and sends the address configuration operation to the idle memory access module. The idle memory access module configures the memory address range corresponding to any remote device carried by the address configuration operation in itself according to the address configuration operation, and establishes a remote memory access connection with the current remote device, thereby binding the idle memory access module to the current remote device. The address resolution module is used to record the mapping relationship between the memory address range, the current remote device, and the memory access module configured with the memory address range.
[0079] It should be noted that after the processor sets the memory address range for each memory access module, in response to determining that the processor has queried an idle memory access module for any remote device, the memory address range configured for the idle memory access module can be associated with the current remote device. In response to determining that the processor has not queried an idle memory access module, a request failure message is returned to the corresponding remote device.
[0080] In order to manage the memory space, the processor also detects the memory space size of the memory address range based on the memory request sent by any remote device; determines the memory mode that matches the memory space size; and manages the corresponding memory space according to the memory mode.
[0081] In some embodiments, the data transmission device further includes a null address processing module. The null address processing module is also a hardware logic processing program located within the integrated circuit and can be implemented by an FPGA. In response to determining that the address resolution module has not found a corresponding memory access module, the null address processing module processes the access request of the CPU (Central Processing Unit) to prevent the operating system from hanging.
[0082] The address resolution module is further configured to: in response to determining that there is no target memory access module having a mapping relationship with the target memory address, forward the address access request to the empty address processing module; the empty address processing module is configured to: construct meaningless response data for the address access request according to a preset strategy (such as a garbled character generation strategy and / or an all-zero character generation strategy), and send the meaningless response data to the processor. When the memory access module has not configured an address, the empty address processing module responds to the request of the processor in the server, so that the server can receive a response even without an address configuration, thereby forming a closed-loop request processing to prevent the processor from receiving an error due to prolonged lack of response.
[0083] In some embodiments, the data transmission device further includes a high-speed interconnect module configured to communicate with the processor via a high-speed interconnect interface (e.g., a CXL interface). The CXL interface complies with the Compute Express Link (CXL) protocol, a new high-speed interconnect interface technology that provides higher data throughput and lower latency.
[0084] The high-speed interconnect interface includes at least: a configuration interface for transmitting address release operations and / or address configuration operations sent by the processor; and an access interface for transmitting address access requests and corresponding response data. In some embodiments, the high-speed interconnect interface includes both CXL.io and CXL.mem interfaces. The processor configures each module in the data transmission device via the CXL.io interface and receives and responds to requests via the CXL.mem interface.
[0085] It can be seen that in some embodiments, each memory access module provided by the data transmission device can directly access a memory address of at least one remote device. Therefore, the to-be-processed data of the acceleration task stored in the memory of the remote device can be directly accessed from the memory of the remote device to the memory access module of the data transmission device; and the processing results of the acceleration task output by the accelerator connected to the processor can also be directly accessed to the memory of the remote device through the memory access module, thereby reducing the number of data transfers, simplifying the processing flow, and reducing resource consumption and processing time during the execution of the acceleration task. The remote devices directly accessed by each memory access module can share the processor connected to the data transmission device and the accelerator connected to the processor, thereby improving the resource utilization of the processor and accelerator.
[0086] A data processing device provided in an embodiment of the present application is introduced below. The data processing device described below can be referenced with other embodiments described herein.
[0087] As shown in FIG2 , some embodiments provide a data processing device, including: a processor, and an accelerator and a data transmission device connected to the processor.
[0088] The data transmission device is used for directly connecting to a memory address of a corresponding remote device according to a memory request of at least one remote device.
[0089] The processor is used for: in response to determining that the data transmission device is directly connected to the memory address of the remote device, using the accelerator to generate and send an address access request to the data transmission device.
[0090] The data transmission device is used to: read the target memory address of the target remote device to be accessed by the processor according to the received address access request, and the to-be-processed data of the acceleration task stored in the target memory address with a mapping relationship, so that the accelerator processes the to-be-processed data; and / or store the processing result of the acceleration task output by the accelerator to the target memory address.
[0091] Among them, the data transmission device includes: an address resolution module and multiple memory access modules; each memory access module is used to directly connect to a memory address of the corresponding remote device based on the memory request of at least one remote device, and supports time-sharing multiplexing of different remote devices connected to it. The remote devices directly accessed by each memory access module share a processor and accelerator.
[0092] The processor is used for: in response to determining that the data transmission device is directly connected to the memory address of the remote device, using the accelerator to generate and send an address access request to the data transmission device.
[0093] The address parsing module is used to parse the received address access request, obtain the target memory address of the target remote device to be accessed by the processor, and determine the target memory access module having a mapping relationship with the target memory address.
[0094] The target memory access module is used to: read the to-be-processed data of the acceleration task stored in the target memory address so that the accelerator processes the to-be-processed data; and / or store the processing result of the acceleration task output by the accelerator to the target memory address.
[0095] In some embodiments, after the processor receives a computing request sent by any remote device, it assigns the computing request sent by any remote device to a request queue corresponding to the accelerator; the accelerator is used to: read the computing request from the request queue, generate an address access request containing the computing request, and send the address access request to the data transmission device to obtain the data to be processed; the accelerator is used to: process the data to be processed to obtain a processing result, generate an address access request containing the processing result, and send the address access request to the data transmission device to transmit the processing result back to the corresponding remote device.
[0096] The accelerator is used to: read the computing request from the request queue, generate an address access request containing the computing request, and send the address access request to the address resolution module to obtain the data to be processed; the accelerator is used to: process the data to be processed to obtain the processing result, generate an address access request containing the processing result, and send the address access request to the address resolution module to transmit the processing result back to the corresponding remote device.
[0097] In some embodiments, the processor is further configured to: query the data transmission device for idle memory access modules based on a memory request sent by any remote device; in response to determining that an idle memory access module has been queried, generate an address configuration operation for the idle memory access module and send the address configuration operation to the idle memory access module; the idle memory access module is configured to: configure the memory address range corresponding to any remote device carried in the address configuration operation within itself according to the address configuration operation, and establish a remote memory access connection with the current remote device; the address resolution module is configured to: record the mapping relationship between the memory address range, the current remote device, and the memory access module configured with the memory address range. The processor is further configured to: generate an address configuration operation for the idle memory access module in the data transmission device based on a memory request sent by any remote device, and send the address configuration operation to the data transmission device; the data transmission device is configured to: cause the idle memory access module to configure the memory address range corresponding to any remote device carried in the address configuration operation within itself according to the address configuration operation, and establish a remote memory access connection with the current remote device.
[0098] In some embodiments, the processor is further configured to: in response to determining that no idle memory access module is found in the query, return an application failure message to the corresponding remote device.
[0099] In some embodiments, the processor is further used to: detect the memory space size of the memory address range based on a memory request sent by any remote device; determine a memory mode that matches the memory space size; and manage the corresponding memory space according to the memory mode.
[0100] In some embodiments, the processor is further configured to: set a configurable address range size for each memory access module in the data transmission device.
[0101] As shown in FIG3 , some embodiments provide another data processing device, which includes multiple processors and multiple data transmission devices; each processor is connected to a data transmission device and multiple accelerators.
[0102] In some embodiments, the data transmission device is further configured to: in response to determining that the target memory address does not exist, utilize a null address processing module within the device to construct meaningless response data for the current address access request in accordance with a preset policy, and send the meaningless response data to the processor. The data transmission device further includes: a null address processing module; the address resolution module is further configured to: in response to determining that the target memory address does not exist, forward the address access request to the null address processing module; the null address processing module is configured to: construct meaningless response data for the current address access request in accordance with a preset policy, and send the meaningless response data to the processor.
[0103] It can be seen that in some embodiments, multiple remote devices can share the same processor and the accelerator connected to the processor, which improves the resource utilization of the processor and the accelerator, and can achieve: the to-be-processed data of the acceleration task stored in the memory of the remote device can be directly transferred from the memory of the remote device to the memory access module of the data transmission device; the processing results of the acceleration task output by the accelerator connected to the processor can also be directly transferred to the memory of the remote device through the memory access module, which reduces the number of data movements, simplifies the processing flow, and reduces the resource consumption and processing time during the execution of the acceleration task.
[0104] The following introduces a data processing system provided in an embodiment of the present application. The data processing system described below can be referenced with other embodiments described in this document.
[0105] As shown in FIG4 , some embodiments provide a data processing system comprising: multiple remote devices, a network device, and the aforementioned data processing device; the data processing device is connected to the multiple remote devices via the network device, such as a switch. The data processing device may be a server.
[0106] The data processing equipment includes: a processor, an accelerator and a data transmission device connected to the processor.
[0107] Among them, the data transmission device includes: an address resolution module and multiple memory access modules; each memory access module is used to directly connect to a memory address of the corresponding remote device based on the memory request of at least one remote device, and supports time-sharing multiplexing of different remote devices connected to it. The remote devices directly accessed by each memory access module share a processor and accelerator.
[0108] The processor is used for: in response to determining that the data transmission device is directly connected to the memory address of the remote device, using the accelerator to generate and send an address access request to the data transmission device.
[0109] The address parsing module is used to parse the received address access request, obtain the target memory address of the target remote device to be accessed by the processor, and determine the target memory access module having a mapping relationship with the target memory address.
[0110] The target memory access module is used to: read the to-be-processed data of the acceleration task stored in the target memory address so that the accelerator processes the to-be-processed data; and / or store the processing result of the acceleration task output by the accelerator to the target memory address.
[0111] In some embodiments, after the processor receives a computing request sent by any remote device, it assigns the computing request sent by any remote device to a request queue corresponding to the accelerator; the accelerator is used to: read the computing request from the request queue, generate an address access request containing the computing request, and send the address access request to the address resolution module to obtain the data to be processed; the accelerator is used to: process the data to be processed to obtain a processing result, generate an address access request containing the processing result, and send the address access request to the address resolution module to transmit the processing result back to the corresponding remote device.
[0112] In some implementations, data is transmitted between a data processing device and a network device, and between a network device and multiple remote devices in an RDMA manner.
[0113] In some embodiments, the data processing system includes a plurality of data processing devices, each of which includes a plurality of processors. Each processor in each data processing device is connected to a data transmission device and a plurality of accelerators.
[0114] It can be seen that in some embodiments, multiple remote devices can share the same processor and the accelerator connected to the processor, which improves the resource utilization of the processor and the accelerator, and can achieve: the to-be-processed data of the acceleration task stored in the memory of the remote device can be directly transferred from the memory of the remote device to the memory access module of the data transmission device; the processing results of the acceleration task output by the accelerator connected to the processor can also be directly transferred to the memory of the remote device through the memory access module, which reduces the number of data movements, simplifies the processing flow, and reduces the resource consumption and processing time during the execution of the acceleration task.
[0115] The following introduces a data transmission method provided in an embodiment of the present application. The data transmission method described below can be referenced with other embodiments described in this document.
[0116] Some embodiments provide a data processing method, which is applied to a processor in a data processing device, wherein the processor is connected to an accelerator and a data transmission device; the data transmission device includes: an address resolution module and multiple memory access modules; each memory access module is used to directly connect to a memory address of a corresponding remote device based on a memory request of at least one remote device, and supports time-sharing multiplexing of different remote devices connected to it, and each remote device directly accessed by each memory access module shares the processor and accelerator.
[0117] 5 , some embodiments provide data processing methods including:
[0118] S501: A processor in a data processing device detects whether a data transmission apparatus is directly connected to a memory address of a remote device.
[0119] S502: In response to determining that the data transmission apparatus is directly connected to the memory address of the remote device, the processor in the data processing device generates an address access request using the accelerator.
[0120] S503: The processor in the data processing device sends an address access request to the data transmission device.
[0121] S504. The data transmission device reads the target memory address of the target remote device to be accessed by the processor according to the received address access request, and the to-be-processed data of the acceleration task stored in the target memory address having a mapping relationship, so that the accelerator processes the to-be-processed data; and / or stores the processing result of the acceleration task output by the accelerator to the target memory address.
[0122] In some embodiments, a processor in a data processing device generates and sends an address access request to an address resolution module using an accelerator. The address resolution module resolves the address access request to obtain a target memory address of a target remote device to be accessed by the processor, and determines a target memory access module that has a mapping relationship with the target memory address. The target memory access module reads the to-be-processed data of the acceleration task stored in the target memory address so that the accelerator processes the to-be-processed data; and / or the target memory access module stores the processing result of the acceleration task output by the accelerator in the target memory address.
[0123] In some examples, some embodiments provide a data processing method including: a processor in a data processing device detects whether a data transmission device is directly connected to a memory address of a remote device; in response to determining that the data transmission device is directly connected to a memory address of a remote device, an accelerator is used to generate and send an address access request to the data transmission device, so that the data transmission device reads the target memory address of the target remote device to be accessed by the processor according to the received address access request, and the data to be processed of the acceleration task stored in the target memory address having a mapping relationship, so that the accelerator processes the data to be processed; and / or stores the processing result of the acceleration task output by the accelerator to the target memory address.
[0124] The processor in the data processing device receives the meaningless response data sent by the data transmission device; in response to determining that the target memory address does not exist, the data transmission device uses the empty address processing module in itself to construct meaningless response data for the current address access request according to a preset strategy.
[0125] In some embodiments, a processor in a data processing device queries an idle memory access module in a data transmission device based on a memory request sent by any remote device; in response to determining that an idle memory access module is queried, an address configuration operation is generated for the idle memory access module, and the address configuration operation is sent to the idle memory access module, so that the idle memory access module configures the memory address range corresponding to any remote device carried by the address configuration operation in itself according to the address configuration operation, and establishes a remote memory access connection with the current remote device; and records the mapping relationship between the memory address range, the current remote device and the memory access module configured with the memory address range in the address resolution module.
[0126] In some embodiments, a processor in a data processing device detects a memory space size of a memory address range based on a memory request sent by any remote device; determines a memory mode that matches the memory space size; and manages the corresponding memory space according to the memory mode.
[0127] In some embodiments, a processor in a data processing device sets a configurable address range size for each memory access module in a data transmission apparatus.
[0128] In some embodiments, the data transmission device also includes: an empty address processing module; a processor in the data processing device, in response to using the address resolution module to determine that there is no target memory access module having a mapping relationship with the target memory address, causes the address resolution module to forward the address access request to the empty address processing module; the processor in the data processing device uses the empty address processing module to construct meaningless response data for the current address access request according to a preset strategy; the processor in the data processing device receives the meaningless response data sent by the empty address processing module.
[0129] In some embodiments, the data transmission device further includes: a high-speed interconnection module; and the processor in the data processing device communicates with the data transmission device via a high-speed interconnection interface of the high-speed interconnection module.
[0130] It can be seen that some embodiments can enable multiple remote devices to share the same processor and the accelerator connected to the processor, thereby improving the resource utilization of the processor and the accelerator, and can achieve: the to-be-processed data of the acceleration task stored in the memory of the remote device can be directly transferred from the memory of the remote device to the memory access module of the data transmission device; the processing results of the acceleration task output by the accelerator connected to the processor can also be directly transferred to the memory of the remote device through the memory access module, thereby reducing the number of data movements, simplifying the processing flow, and reducing the resource consumption and processing time during the execution of the acceleration task.
[0131] It should be noted that CXL technology can enable accelerators such as GPU (Graphics Processing Unit) and FPGA (Field-Programmable Gate Array) to better collaborate with processors, thereby improving the speed of artificial intelligence model training and reasoning.
[0132] See Figure 6. The system shown in Figure 6 combines the features of CXL and RDMA remote memory access to create a shared accelerator cluster. This cluster consists of two servers providing GPU / FPGA accelerators, an RDMA switch (i.e., a network device), and multiple clients (i.e., remote devices) sharing the GPU / FPGA. This cluster design fully utilizes the GPU / FPGA while reducing accelerator configuration costs.
[0133] In Figure 6, the processor in the server is connected to a data transmission device (hereinafter referred to as a CXL address decoder). Referring to Figure 7, the CXL address decoder consists of a CXL high-speed interconnect module, an address resolution module, an RDMA module, and a null address processing module.
[0134] The CXL high-speed interconnect module supports two interfaces: CXL.io and CXL.mem. The processor configures or controls the address resolution module, RDMA module, and null address handling module through the CXL.io interface (configuration interface), and receives and responds to processor memory access requests (i.e., address access requests) through the CXL.mem interface (access interface).
[0135] Referring to Figure 8 , the address resolution module receives address access requests from the CXL.mem interface and resolves the memory address carried in the address access request. The address resolution module maintains an address mapping table that stores the mapping relationship between memory address ranges and RDMA modules. The address resolution module first queries the address mapping table based on the memory address. Upon determining that a corresponding mapping relationship exists in the table, the module forwards the memory access request to the RDMA module for processing. The address mapping table is configured via the CXL.io interface. If a memory address issued from the CXL.mem interface does not have a corresponding RDMA module in the address mapping table, the memory access request is forwarded to the null address processing module. Since the processor waits for the access result to be returned after issuing a memory access request to the CXL address decoder, an access timeout will cause the processor to report an error, potentially causing a system crash. Therefore, some embodiments utilize the null address processing module to provide the processor with meaningless response data to avoid processor errors.
[0136] The RDMA module accesses memory on a remote computer. After receiving a memory access request from the address resolution module, the RDMA module accesses the remote client's memory according to pre-configured settings and returns the access result to the address resolution module. The RDMA module is configured through the CXL.io interface.
[0137] During server system initialization, the server processor determines the address space range (i.e., memory address range) for each RDMA module based on the user's configuration of the CXL address decoder and distributes this information to the CXL address decoder via the CXL.mem interface. However, at this point, the RDMA module cannot actually use this address space range because the RDMA module only knows the size of the memory it can access, not the address of this memory range. In other words, the actual memory address that can be accessed has not yet been mapped to the RDMA module. In this case, memory access requests from the CXL address decoder are handled by the null address handling module.
[0138] Figure 6 illustrates an example of the organizational structure of an accelerator sharing cluster. The number of servers and clients in the diagram can be adjusted as needed. CXL address decoders are installed within the servers, and the number of CXL address decoders and accelerators within the servers can be adjusted as needed. Multiple clients can remotely share accelerators within the server via the CXL address decoders.
[0139] As shown in Figures 6 and 7, the CXL address decoder is an intelligent device with multiple RDMA modules connected to the switch. After the processor configures the RDMA modules on the device through CXL.io, it triggers the RDMA modules to automatically connect to the configured clients. Each RDMA module is connected to at least one client. Because there are multiple RDMA modules, the accelerator on the server can be shared by multiple clients through the CXL address decoder. When a single RDMA module is connected to multiple clients, these clients time-share the RDMA module.
[0140] As shown in Figure 9, after the CXL address decoder is initialized, the client can request CXL memory address space (memory address range) and RDMA modules from the server and use the server's accelerator resources. Upon determining that the client requests 1TB of space from the server for computation, the server first checks whether there is sufficient free CXL memory space (the size of which is specified by the server). After requesting the CXL memory space, the client also requests the RDMA module. The CXL memory space range and client information are then assigned to the RDMA module, triggering the RDMA module to automatically establish a connection with the client. Upon successful connection, the address mapping table between the memory address range and the RDMA module is written to the address mapping table. Upon determining that only a free CXL memory range or only a free RDMA module has been requested, a failure notification is returned to the client, and the relevant resources must be released.
[0141] Furthermore, the server maintains a compute request queue for each GPU / FPGA accelerator in the system. When a client requests a computation from the server, it finds the shortest queue and inserts the new compute request at the end of it. When the accelerator has idle resources, it retrieves the first compute request from its queue and processes it. The compute request records the actual memory address of the computational model and data (the memory address of the data on the client), which the GPU / FPGA can directly access. Because RDMA remote memory access is much faster than hard disks, computing efficiency is greatly improved.
[0142] In some embodiments, the client exposes its local memory to the server via RDMA and CXL. The server can access the client memory directly as if it were local memory. The computational model and data in the client memory are directly moved to the GPU / FPGA accelerator, eliminating the need for server resources. The client provides the model and data for computation, while the server's GPU / FPGA accelerator focuses on computation.
[0143] After the client's computing request is completed by the accelerator, the computing results are also directly placed in the client memory, and then the CXL memory range and RDMA module occupied by the computing request are unloaded.
[0144] The CXL memory range on the CXL address decoder is repeatedly allocated and unloaded, which leads to CXL memory fragmentation. Therefore, the CXL memory range needs to be managed. Assuming that the CXL memory range size is 4TB, Figure 10 is an example diagram of memory mode management.
[0145] As shown in Figure 11, there are four memory space modes (i.e., memory modes): 128GB, 256GB, 512GB, and 1TB. Each 1TB of memory space can be configured in any of these four modes. Once a 1TB of memory space is set to one of these modes, it cannot be set to another. When a 1TB of memory space with the same mode is released, the 1TB of memory space is reclaimed, returning to a state without any mode.
[0146] Initially, when a client requests a memory space less than or equal to 128GB, the first 1TB of the CXL address decoder is set to 128GB mode, and the first 128GB is allocated to the client. If a second request is less than or equal to 128GB, the second 128GB is allocated from the 1TB space. If the second request is greater than 128GB but less than or equal to 256GB, the second 1TB is set to 256GB mode, and the first 256GB is allocated to the second request.
[0147] Once a 1TB space is allocated for a pattern, a new 1TB space is set up to represent that pattern, and space equal to the pattern size is allocated from this 1TB space for computational requests. In this example, the pattern size can be adjusted as needed.
[0148] In some embodiments, algorithm calculations can be performed directly on the client memory without switching data again, which can fully utilize the GPU / FPGA accelerator and reduce the cost of model calculation.
[0149] The following describes an electronic device provided by an embodiment of the present application, and the electronic device described below can be cross-referenced with other embodiments described herein. The electronic device can be a data transmission device or a data processing device described herein.
[0150] The present application discloses an electronic device, including:
[0151] a memory for storing computer-readable instructions;
[0152] A processor is configured to execute computer-readable instructions to implement the method disclosed in any of the above embodiments.
[0153] Furthermore, an embodiment of the present application also provides an electronic device. The electronic device can be either a server as shown in FIG11 or a terminal as shown in FIG12 . FIG11 and FIG12 are both structural diagrams of electronic devices according to an exemplary embodiment, and the contents of the diagrams are not to be construed as limiting the scope of use of the present application.
[0154] Figure 11 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server may include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store computer-readable instructions, which are loaded and executed by the processor to implement the relevant steps of the data processing disclosed in any of the aforementioned embodiments.
[0155] In some embodiments, the power supply is used to provide operating voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not limited here; the input and output interface is used to obtain external input data or output data to the outside world, and its interface type can be selected according to application needs and is not limited here.
[0156] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer-readable instructions and data, etc. The storage method can be temporary storage or permanent storage.
[0157] The operating system manages and controls the server's hardware devices and computer-readable instructions, enabling the processor to operate and process data in memory. The operating system can be Windows Server, NetWare, Unix, Linux, or other operating systems. In addition to computer-readable instructions for implementing the data processing methods disclosed in any of the aforementioned embodiments, the computer-readable instructions can also include computer-readable instructions for performing other specific tasks. Data can include application update information and other data, such as information about the application's developer.
[0158] FIG12 is a schematic structural diagram of a terminal provided in an embodiment of the present application. The terminal may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0159] Generally, a terminal in some embodiments includes: a processor and a memory.
[0160] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA, and PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU, which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0161] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the memory is used to store at least the following computer-readable instructions, wherein, after the computer-readable instructions are loaded and executed by the processor, the relevant steps in the data processing method performed by the terminal side disclosed in any of the aforementioned embodiments can be implemented. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to update information of the application.
[0162] In some embodiments, the terminal may further include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.
[0163] Those skilled in the art will appreciate that the structure shown in FIG12 does not limit the terminal and may include more or fewer components than shown in the figure.
[0164] The following is an introduction to a non-transitory computer-readable storage medium provided in an embodiment of the present application. FIG13 is a schematic diagram of the structure of a non-transitory computer-readable storage medium provided in the present application.
[0165] A non-transitory computer-readable storage medium described below may be cross-referenced with other embodiments described herein.
[0166] A readable storage medium for storing computer-readable instructions, wherein the computer-readable instructions, when executed by one or more processors, implement the data processing method disclosed in the aforementioned embodiments. The readable storage medium is a non-transitory computer-readable storage medium, which, as a carrier for resource storage, can be a read-only memory, random access memory, a magnetic disk, or an optical disk, etc. The resources stored thereon include an operating system, computer-readable instructions, and data, etc., and the storage method can be temporary storage or permanent storage.
[0167] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0168] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM (Compact Disc Read-Only Memory), or any other form of readable storage medium known in the art.
[0169] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A data transmission device, characterized in that: include: Address resolution module and multiple memory access modules; Each memory access module is used to directly connect to a memory address of a corresponding remote device according to a memory request of at least one remote device, and supports time-sharing multiplexing of different remote devices connected thereto, and each remote device directly accessed by each memory access module shares the processor connected to the data transmission device and the accelerator connected to the processor; The address parsing module is used to: determine, according to the received address access request, a target memory access module corresponding to the target remote device to be accessed by the processor, wherein the target memory access module has a mapping relationship with the target memory address of the target remote device; and The target memory access module is used to: read the pending data of the acceleration task stored in the target memory address so that the accelerator connected to the processor processes the pending data; and / or store the processing result of the acceleration task output by the accelerator connected to the processor to the target memory address.
2. The device according to claim 1, characterized in that Each memory access module is configured with a memory address of at least one remote device; The arbitrary memory access module is used to: configure the memory address range carried by the address configuration operation in itself according to the address configuration operation sent by the processor, and establish a remote memory access connection with the current remote device, wherein the memory address range corresponds to any remote device; and The address resolution module is further used to record a mapping relationship between the memory address range, the current remote device and a memory access module configured with the memory address range.
3. The device according to claim 2, characterized in that The arbitrary memory access module is used to disconnect the remote memory access connection with the connected remote device according to the address release operation sent by the processor, so that the arbitrary memory access module can directly connect to the memory address of other remote devices.
4. The device according to any one of claims 1 to 3, characterized in that Also includes: Empty address processing module; The address parsing module is further configured to: in response to determining that there is no target memory access module having a mapping relationship with the target memory address, forward the address access request to the empty address processing module; and The empty address processing module is used to construct meaningless response data for the address access request according to a preset strategy, and send the meaningless response data to the processor.
5. The device according to any one of claims 1 to 3, characterized in that: Also includes: High-speed interconnect modules; The high-speed interconnect module is used to: communicate with the processor via a high-speed interconnect interface; as well as The high-speed interconnect interface at least comprises: a configuration interface for transmitting at least one of an address release operation and an address configuration operation sent by the processor; as well as The access interface is used to transmit the address access request and corresponding response data.
6. A data processing device, characterized in that: include: A processor and an accelerator and a data transmission device connected to the processor; The data transmission device is used to: directly connect to a memory address of a corresponding remote device according to a memory request of at least one remote device; The processor is used to: in response to determining that the data transmission device is directly connected to the memory address of the remote device, generate and send an address access request to the data transmission device using the accelerator; and The data transmission device is used to: read the target memory address of the target remote device to be accessed by the processor and the to-be-processed data of the acceleration task stored in the target memory address having a mapping relationship according to the received address access request, so that the accelerator processes the to-be-processed data; And / or storing the processing result of the acceleration task output by the accelerator to the target memory address.
7. The device according to claim 6, characterized in that The processor is used to: allocate a computing request sent by any remote device to a request queue corresponding to the accelerator; The accelerator is used to: read the computing request from the request queue, generate an address access request containing the computing request, and send the address access request to the data transmission device to obtain the data to be processed; and The accelerator is used to: process the data to be processed to obtain the processing result, generate an address access request containing the processing result, and send the address access request to the data transmission device to transmit the processing result back to the corresponding remote device.
8. The device according to claim 6, characterized in that The processor is further configured to: generate an address configuration operation for an idle memory access module in the data transmission device according to a memory request sent by any remote device, and send the address configuration operation to the data transmission device; The data transmission device is used to: enable the idle memory access module to configure the memory address range corresponding to any remote device carried by the address configuration operation in itself according to the address configuration operation, and establish a remote memory access connection with the current remote device; and The processor is further configured to: in response to determining that no idle memory access module is found, return an application failure message to the corresponding remote device.
9. The device according to claim 8, characterized in that The processor is further configured to: Detect the memory space size of the memory address range according to the memory request sent by any remote device; Determine the memory mode that matches the size of the memory space; and The corresponding memory space is managed according to the memory mode.
10. The device according to claim 8, characterized in that The processor is further configured to: set a configurable address range size for each memory access module in the data transmission device.
11. The device according to any one of claims 6 to 10, characterized in that The data processing device comprises a plurality of processors and a plurality of data transmission devices; each processor is connected to a data transmission device and a plurality of accelerators.
12. The device according to any one of claims 6 to 10, characterized in that The data transmission device is also used to: in response to determining that the target memory address does not exist, use the empty address processing module in itself to construct meaningless response data for the current address access request according to a preset strategy, and send the meaningless response data to the processor.
13. A data processing system, characterized in that: include: A plurality of remote devices, a network device and a data processing device as claimed in any one of claims 6 to 12; The data processing device is connected to the multiple remote devices through the network device.
14. The system according to claim 13, characterized in that Data is transmitted between the data processing device and the network device, and between the network device and the plurality of remote devices in a remote direct data access (RDMA) manner.
15. A data processing method, characterized in that: A processor used in a data processing device, wherein the processor is connected to an accelerator and a data transmission device; the method comprises: Detecting whether the data transmission device is directly connected to a memory address of a remote device; and In response to determining that the data transmission device is directly connected to the memory address of the remote device, the accelerator is used to generate and send an address access request to the data transmission device, so that the data transmission device reads the target memory address of the target remote device to be accessed by the processor according to the received address access request, and the target memory address has a mapping relationship with the target memory address of the target remote device to be accessed by the processor, and the to-be-processed data of the acceleration task stored in the target memory address, so that the accelerator processes the to-be-processed data; and / or stores the processing result of the acceleration task output by the accelerator to the target memory address.
16. The method according to claim 15, characterized in that Also includes: querying an idle memory access module in the data transmission device according to a memory request sent by any remote device; as well as In response to determining that an idle memory access module is queried, an address configuration operation is generated for the idle memory access module, and the address configuration operation is sent to the idle memory access module, so that the idle memory access module configures the memory address range corresponding to any remote device carried by the address configuration operation in itself according to the address configuration operation, and establishes a remote memory access connection with the current remote device.
17. The method according to claim 16, characterized in that Also includes: Detect the memory space size of the memory address range according to the memory request sent by any remote device; Determine the memory mode that matches the size of the memory space; as well as The corresponding memory space is managed according to the memory mode.
18. The method according to any one of claims 15 to 17, characterized in that Also includes: A configurable address range size is set for each memory access module in the data transmission device.
19. The method according to any one of claims 15 to 17, characterized in that Also includes: Receive meaningless response data sent by the data transmission device; in response to determining that the target memory address does not exist, the data transmission device uses the empty address processing module in itself to construct meaningless response data for the current address access request according to a preset strategy.
20. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the method according to any one of claims 15 to 19 is implemented.
Citation Information
Patent Citations
Hardware calculation module, device and method, electronic device and storage medium
CN116627888A
DMA method and device between acceleration cards, acceleration cards, acceleration platform and medium
CN117033275A
Data transmission device, data processing equipment, system, method and medium
CN117312229A
Shared memory space among devices
US20200104275A1
Cited By
Edge computing method and system capable of realizing direct exchange of remote memory data
CN121070862A