A data processing system, method and computer system

Through the data processing system directly connected to the processing board and the memory expansion board, the memory shortage and communication bottleneck problems caused by the limited computing power of the GPU chip are solved, and the processing board's direct access to the memory expansion board storage components is realized, which reduces access delay and coupling, and effectively expands memory.

CN119046211BActive Publication Date: 2025-05-09LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411534503.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-09
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

During the training of the artificial intelligence model, due to the limited computing power of a single GPU chip, multiple GPU chips are required to compute together, resulting in insufficient memory, longer access paths, increased delays, and limited capacity of extended memory, resulting in communication bottlenecks.

Method used

By designing a data processing system, the processing board is directly connected to the memory expansion board. The processing core on the processing board recognizes the storage components through a bridge driver thread to establish an address space mapping relationship, so that the processing board can directly access the storage components on the memory expansion board without having to copy memory through the CPU.

Benefits of technology

It reduces the coupling degree between the processing board when accessing extended memory and the server host, shortens the access path, reduces access delay, effectively expands the processing board's memory, and reduces the communication bottleneck in the training process of pre-trained model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046211B_ABST
    Figure CN119046211B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing system, method and computer system, which relates to the field of computer systems. In order to solve the problem of long access path and large access delay in accessing extended memory, the data processing system includes a processing board and a memory extension board, the processing board is provided with a processing core and a first controller, and the memory extension board is provided with a storage component and a control component. The present invention enables the processing board to directly access the storage component on the memory extension board without copying the memory through a server host, reduces the coupling between the processing board and the server host when accessing the extended memory, effectively expands the memory of the processing board, shortens the access path of the processing board to the extended memory, reduces the access delay, and thus reduces the communication bottleneck in the training process of the pre-training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer systems, and in particular to a data processing system, method and computer system. Background Art

[0002] In the current training process of artificial intelligence models, due to the limited computing power of a single GPU (Graphics Processing Unit) chip, multiple GPU chips are required to work together, which involves interconnection and communication between multiple GPU chips, as well as the transmission of intermediate computing data. To solve the problem of insufficient memory, a memory expansion card is usually used.

[0003] Memory expansion cards are usually connected to computer systems through PCIe (Peripheral Component Interconnect Express) interfaces to provide more memory resources for GPU chips to support them in processing more complex data sets and computing tasks. However, when GPU chips access these extended memories, they must perform memory copy operations through the CPU (Central Processing Unit), which makes the access path longer, increases latency, and limits the capacity of the extended memory, thus creating a communication bottleneck for pre-trained model training.

[0004] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the invention

[0005] The object of the present invention is to provide a data processing system, method and computer system, which can enable a processing board to directly access storage components on a memory expansion board without the need to copy the memory through a server host, thereby reducing the coupling between the processing board and the server host when accessing the extended memory, while effectively expanding the memory of the processing board, shortening the access path of the processing board to the extended memory, reducing access delay, and thus reducing the communication bottleneck in the pre-training model training process.

[0006] To solve the above technical problems, the present invention provides a data processing system, including a processing board and a memory expansion board; the processing board includes: a processing core, used to identify a storage component through a bridge driver thread, establish a first mapping relationship between the address space of the processing board and the storage component, and forward an internal access request that complies with the first mapping relationship; the internal access request includes a first internal access request generated by the processing core; a first controller, used to convert the received internal access request into an extended access request and output it; the memory expansion board includes the storage component and a control component, and the control component is used to perform an access operation on the storage component in response to the extended access request.

[0007] Among them, the processing board also includes a second controller, which is used to establish a second mapping relationship between the storage component and the address space of the server host, and forward the first host access request sent by the server host that complies with the second mapping relationship; the first controller is also used to convert the received first host access request into the extended access request and output it.

[0008] The processing board further includes a system bus, which is used to forward the internal access request to the first controller and / or forward the first host access request to the first controller.

[0009] Among them, the processing core is specifically used to generate internal access requests, identify storage components through a bridge driver thread, and establish a first mapping relationship between the address space of the processing board and the physical address of the storage component, forward internal access requests whose destination addresses match the first mapping relationship, and discard internal access requests whose destination addresses do not match the first mapping relationship.

[0010] The processing board further includes at least one computing core for generating a second internal access request corresponding to a current computing requirement; the internal access request further includes the second internal access request.

[0011] Wherein, the memory expansion board also includes a monitoring management module for detecting the operating status of the memory expansion board.

[0012] Among them, the control component includes: an analysis and processing module, which is used to generate and output a target memory access request based on the access address and access type of the received current access request; the current access request includes the extended access request; and a memory control module, which is used to perform an access operation corresponding to the target memory access request on the storage component.

[0013] Among them, the analysis and processing module includes a first hard core module, which is used to generate and output a first memory access request based on the access address and access type of the received extended access request; the target memory access request includes the first memory access request; the memory control module is specifically used to respond to the first memory access request received and perform a corresponding access operation on the storage component.

[0014] Among them, the analysis and processing module also includes a second hard-core module, which is used to generate and output a second memory access request based on the access address and access type of the received second host access request; the current access request also includes the second host access request sent by the server host, and the target memory access request also includes the second memory access request; the memory control module is specifically used to respond to the received second memory access request and perform corresponding access operations on the storage component.

[0015] Among them, the analysis and processing module also includes: a network interface module, used to receive a remote access request; a network protocol stack, used to generate and output a third memory access request based on the access address and access type of the remote access request; the target memory access request also includes the third memory access request, and the current access request also includes the remote access request; the memory control module is specifically used to respond to the received third memory access request and perform a corresponding access operation on the storage component.

[0016] The analysis and processing module further includes an access arbitration module for arbitrating the received first memory access request and / or the second memory access request and / or the third memory access request, and outputting a target memory access request that has been successfully arbitrated.

[0017] The memory expansion board further includes a network optical module connected to the network interface module, and is used to receive the remote access request sent by the remote computing node.

[0018] The control component further includes a control register module for managing registers in the memory control module and the analysis and processing module.

[0019] Among them, the storage component includes a non-volatile storage device; the memory control module includes: a shared memory controller, used to set the access state of the access address of the received target memory access request to a locked state, and forward the target memory access request; a memory access controller, used to perform corresponding access operations on the non-volatile storage device according to the access address and access type of the target memory access request.

[0020] Among them, the storage component also includes memory particles; the memory access controller is specifically used to parse and forward the access information in the target memory access request; the access information includes the access type and access address corresponding to the target memory access request; the memory control module also includes: a cache controller, which is used to respond to the access type in the access information as a read operation type. If the access address in the access information hits the cache, the corresponding response data is read from the memory particle and returned to the memory access controller; otherwise, the access information is forwarded, and the response data returned by the storage controller is cached in the memory particle; the storage controller is used to read the corresponding response data in the non-volatile storage device according to the access address in the access information sent by the cache controller and return it.

[0021] Among them, the cache controller is also used to forward the access information in response to the access type in the access information being a write operation type; the storage controller is also used to write the access data in the access information into the non-volatile storage device according to the access address in the access information sent by the cache controller.

[0022] Among them, the storage controller is specifically used to determine a read address range including the access address based on the access address in the access information sent by the cache controller, read the data to be cached in the non-volatile storage device according to the read address range and return it; the cache controller is specifically used to respond to the access type in the access information being a read operation type, if the access address in the access information hits the cache, read the corresponding response data from the memory particle and return it to the memory access controller, otherwise, forward the access information, cache the data to be cached in the memory particle, and return the response data in the data to be cached to the memory access controller.

[0023] The memory access controller is also used to return the response data to the shared memory controller so that the response data is returned to the sender of the target memory access request through the shared memory controller; the sender is the server host or the processing board or the remote computing node.

[0024] In order to solve the above technical problems, the present invention further provides a computer system, comprising a server host and a data processing system as described in any one of the above.

[0025] In order to solve the above technical problems, the present invention further provides a data processing method, which is applied to the data processing system as described in any one of the above, wherein the data processing system comprises a processing board and a memory expansion board provided with a storage component, and the data processing method comprises:

[0026] The processing core on the processing board identifies the storage component through a bridge driver thread, establishes a first mapping relationship between the address space of the processing board and the storage component, and forwards an internal access request that complies with the first mapping relationship; the internal access request includes a first internal access request generated by the processing core; the first controller on the processing board converts the received internal access request into an extended access request and outputs it; and the control component on the memory expansion board responds to the extended access request to perform an access operation on the storage component.

[0027] The present application provides a data processing system, which directly connects a memory expansion board and a processing board through a connector and a cable, thereby reducing the coupling between the processing board and a server host when accessing the extended memory, and using a thread on a processing core of the processing board to implement a bridge driver function, so as to identify and address-map the processing components on the memory expansion board connected to itself, so that the processing board can directly access the storage components on the memory expansion board without the need for memory copying through the CPU. While effectively expanding the memory of the processing board, the access path of the processing board to the extended memory is shortened, and the access delay is reduced, thereby reducing the communication bottleneck in the training process of the pre-trained model.

[0028] The present application also provides a data processing method and a computer system, which have the same beneficial effects as the above-mentioned data processing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0030] Figure 1 A schematic diagram of the structure of a first data processing system provided by the present invention;

[0031] Figure 2 A schematic diagram of the structure of a second data processing system provided by the present invention;

[0032] Figure 3 A schematic diagram of the structure of a processing board provided by the present invention;

[0033] Figure 4 A schematic diagram of the structure of a memory expansion board provided by the present invention;

[0034] Figure 5 A schematic diagram of the structure of a control component provided by the present invention;

[0035] Figure 6A schematic diagram of the structure of a third data processing system provided by the present invention;

[0036] Figure 7 This is a flow chart of the steps of the data processing method provided by the present invention. DETAILED DESCRIPTION

[0037] The core of the present invention is to provide a data processing system, method and computer system, which enable a processing board to directly access storage components on a memory expansion board without the need to copy memory through a server host, thereby reducing the coupling between the processing board and the server host when accessing the extended memory. While effectively expanding the memory of the processing board, the access path of the processing board to the extended memory is shortened, the access delay is reduced, and the communication bottleneck in the pre-training model training process is reduced.

[0038] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] First, please refer to Figure 1 The present invention provides a data processing system, including a processing board 1 and a memory expansion board 2, wherein: the processing board 1 includes: a processing core 11, which is used to identify a storage component 21 through a bridge driver thread, establish a first mapping relationship between the address space of the processing board 1 and the storage component 21, and forward an internal access request that complies with the first mapping relationship; the internal access request includes a first internal access request generated by the processing core 11; a first controller 12, which is used to convert the received internal access request into an extended access request and output it; the memory expansion board 2 includes a storage component 21 and a control component 22, and the control component 22 is used to perform an access operation on the storage component 21 in response to the extended access request.

[0040] In this embodiment, considering that a single GPU chip can access limited memory, in order to solve the problem of insufficient memory, a memory expansion card in the form of a PCIe (Peripheral Component Interconnect Express) gold finger interface can be made based on technologies such as DDR (Double Data Rate) memory particles or PMem (Persistent Memory) particles. However, when the GPU chip accesses these extended memories, it must perform memory copy operations through the CPU (Central Processing Unit), which makes the access path longer, increases the delay, and the capacity of the extended memory is also limited. Based on this, this embodiment designs a new processing board 1 and a memory expansion board 2, and the two are used to build a data processing system.

[0041] The data processing system includes at least one processing board 1 and at least one memory expansion board 2. One processing board 1 can be connected to multiple memory expansion boards 2, and one memory expansion board 2 can also be connected to multiple processing boards 1. It can be set according to actual engineering needs. To facilitate understanding of the memory expansion solution provided in this embodiment, this embodiment is explained by taking a processing board 1 connected to a memory expansion board 2 as an example. The processing board 1 in this embodiment can specifically be a GPU board, and the processing board 1 and the memory expansion board 2 can be detachably connected, and can be specifically connected through a connecting device. As an optional embodiment, refer to Figure 2 As shown, the connection device includes but is not limited to a first connector 13 provided on the processing board 1, a second connector 23 provided on the memory expansion board 2, and a cable for connecting the first connector 13 and the second connector 23. The first connector 13 on the processing board 1 and the second connector 23 on the memory expansion board 2 can both use MCIO (Mini Cool Edge IO, Mini Cool Edge Input / Output) connectors. It can be understood that MCIO is a new type of high-speed cable interface technology, which is designed to meet the needs of high-speed data transmission in data centers and high-performance computing environments. It can support multiple protocols including PCIe, CXL (Compute Express Link, Compute Express Link), etc., has good signal integrity and a small package size, and is suitable for long-distance inter-board data transmission. Of course, in addition to the combination of connectors and cables, the connection device can also use other connection methods to meet actual engineering needs. This embodiment is not specifically limited here. The processing board 1 and the memory expansion board 2 are directly connected via an MCIO high-speed differential cable. On the one hand, this facilitates reducing the coupling between the processing board 1 and the server host when using the extended memory, allowing the processing board 1 to directly access the mounted extended memory. On the other hand, it is conducive to system integration, and can improve system bandwidth and reduce access latency.

[0042] The processing core 11 may be a GPU core on the processing board 1. In this embodiment, the bridge driver uses the processing core 11 to run on a GPU thread. The bridge driver may be a RC (Root Complex) bridge driver, that is, the function of the RC bridge driver is unloaded from the CPU to the processing core 11 of the processing board 1 to realize the extended memory function of the processing board 1. At the same time, the CPU load can be reduced, allowing the CPU to focus more on other computing tasks and improve the overall system efficiency. Specifically, when the first connector 13 on the processing board 1 and the second connector 23 on the memory expansion board 2 are connected through a cable, the bridge driver thread on the processing core 11 starts to execute. The bridge driver thread sends a signal or performs a detection operation to identify the connected memory expansion board 2 and the storage component 21 thereon. After the storage component 21 is identified, the bridge driver thread establishes a first mapping relationship between the address space of the processing board 1 and the storage component 21. The first mapping relationship defines how the logical address on the processing board 1 corresponds to the physical address of the storage component 21. After the first mapping relationship is established, the processing board 1 can identify the address range of the storage component 21 and use it as an available memory resource. The processing core 11 generates an internal access request according to the current application scenario. In this embodiment, the internal access request generated by the processing core 11 is determined as the first internal access request. The internal access request includes a target address. The bridge driver thread determines whether the internal access request satisfies the first mapping relationship, that is, whether the target address requested to be accessed by the internal access request matches the address of the storage component 21 on a certain memory expansion board 2 in the first mapping relationship. If so, it is determined that the internal access request satisfies the first mapping relationship, and the internal access request is forwarded to the first controller 12.

[0043] The first controller 12 performs format conversion on the internal access request, converts it into an extended access request that can be transmitted through the first connector 13, and then transmits it to the second connector 23 through the first connector 13 and the cable. In this embodiment, the first controller 12 may be a CXL RC (Compute Express Link Root Complex) controller, and accordingly, the first controller 12 is specifically used to convert the internal access request into an extended access request that complies with the CXL protocol.

[0044] After receiving the extended access request through the second connector 23, the control component 22 on the memory expansion board 2 performs corresponding operations on the storage component 21 on the memory expansion board 2 based on the extended access request. In this embodiment, the storage component 21 is connected to the control component 22, and the storage component 21 includes at least one storage device. This embodiment does not limit the type of the storage device.

[0045] It can be seen that in this embodiment, the memory expansion board 2 and the processing board 1 are directly connected through a connector and a cable, which reduces the coupling between the processing board 1 and the server host when accessing the extended memory. A thread is used on a processing core 11 of the processing board 1 to implement a bridge driver function, and the processing component on the memory expansion board 2 connected to itself is identified and the address is mapped, so that the processing board 1 can directly access the storage component 21 on the memory expansion board 2 without the need for memory copying through the CPU. While effectively expanding the memory of the processing board 1, the access path of the processing board 1 to the extended memory is shortened, the access delay is reduced, and the communication bottleneck in the training process of the pre-trained model is reduced.

[0046] In an exemplary embodiment, referring to Figure 3 The processing board 1 also includes a second controller 14, which is used to establish a second mapping relationship between the storage component 21 and the address space of the server host, and forward the first host access request sent by the server host that complies with the second mapping relationship; the first controller 12 is also used to convert the received first host access request into an extended access request and output it.

[0047] In this embodiment, the second controller 14 may be a CXL EP (End Point) controller, which is used to receive and send data between the processing board 1 and the CPU of the server host, so as to enable the server host to access the extended memory of the processing board 1. Specifically, the second controller 14 is used to establish a second mapping relationship between the storage component 21 and the address space of the server host. The second mapping relationship allows the CPU of the server host to access the extended memory of the processing board 1 through the address space. When the server host sends a first host access request, the second controller 14 checks whether these first host access requests satisfy the second mapping relationship. If satisfied, the first host access request is forwarded to the first controller 12. The server host can increase its effective memory capacity through the extended memory of the processing board 1 without directly upgrading the physical memory of the host.

[0048] In an exemplary embodiment, the processing board 1 further includes a system bus 15 for forwarding an internal access request to the first controller 12 and / or forwarding a first host access request to the first controller 12 .

[0049] In this embodiment, the processing board 1 also includes a system bus 15, and the first controller 12, the second controller 14 and the processing core 11 are all mounted on the system bus 15. Transmitting data through the system bus 15 can reduce the complexity and delay of data transmission between various components, thereby improving transmission efficiency, simplifying board-level design, and reducing the number of required hardware interfaces and connectors.

[0050] In an exemplary embodiment, the processing board 1 further includes at least one computing core 16 for generating a second internal access request corresponding to a current computing demand; the internal access request further includes a second internal access request.

[0051] It can be understood that in addition to the processing core 11, the processing board 1 also includes other GPU cores, namely the computing core 16 in this embodiment. The computing core 16 generates a second internal access request based on the current computing demand, and the second internal access request is also included in the internal access request. The second internal access request is transmitted to the processing core 11, and the processing core 11 determines whether to forward the second internal access request, so that the processing core 11 can coordinate the memory access between multiple GPU cores to ensure that they can efficiently share the extended memory resources and avoid resource conflicts and waste. In addition, the data path can be optimized to reduce the transmission delay of data between GPU cores and improve the data transmission efficiency. By centrally processing memory access requests through the processing core 11, the system can more easily expand the number of GPU cores without large-scale modifications to the memory management architecture. The processing core 11 can ensure that high-performance computing tasks obtain sufficient memory bandwidth and low-latency access, thereby improving computing performance.

[0052] In an exemplary embodiment, referring to Figure 3 The processing board 1 also includes a first gold finger interface 17 detachably connected to a first slot on a server host; the server host is used to allocate a server host address space to the processing board 1 when the first gold finger interface 17 is connected to the first slot.

[0053] In this embodiment, the first gold finger interface 17 is a gold finger interface that supports the PCIe communication standard. The PCIe communication standard includes but is not limited to PCIe Gen4 x16, PCIe Gen5x16 standards, etc. It can be selected based on the compatibility of the server motherboard chipset, processor and related equipment. This embodiment does not make specific limitations here.

[0054] In an exemplary embodiment, referring to Figure 4 The memory expansion board 2 also includes a second gold finger interface 24 that is detachably connected to a second slot on the server host.

[0055] In this embodiment, the second gold finger interface 24 is a gold finger interface that supports the PCIe communication standard. The PCIe communication standard includes but is not limited to the PCIe Gen4 x16, PCIe Gen5x16 standard, etc., which can be selected based on the compatibility of the chipset, processor and related equipment of the server motherboard. This embodiment does not make specific restrictions here. The gold fingers of the two cards are respectively inserted into the corresponding PCIe slots on the server host, which facilitates the server host to access the memory expansion board 2 and the processing board 1.

[0056] In an exemplary embodiment, referring to Figure 4 The memory expansion board 2 also includes a monitoring management module 25 for detecting the operating status of the memory expansion board 2 .

[0057] The present embodiment also includes a monitoring management module 25 , which can be constructed by an MCU (Microcontroller Unit) and related sensors to monitor the operating status of the memory expansion board 2 , such as voltage, current, temperature, etc.

[0058] In an exemplary embodiment, referring to Figure 4 The memory expansion board 2 also includes a network optical module 26 for receiving a remote access request sent by a remote computing node.

[0059] In an exemplary embodiment, the control component 22 includes: an analysis and processing module for generating and outputting a target memory access request based on an access address and an access type of a received current access request; the current access request includes an extended access request; and a memory control module for performing an access operation corresponding to the target memory access request on the storage component 21.

[0060] In this embodiment, the control component 22 includes an analysis and processing module and a memory control module, which work together to process the received current access request, and the current access request at least includes the extended access request sent by the processing board 1. The analysis and processing module parses the access address and access type of the current access request, and generates a target memory access request based on the access address and access type to ensure that valid information in the current access request is transmitted to the memory control module. Of course, the analysis and processing module can also perform access rights checks based on the access address and access type. If the sender of the current access request does not have the corresponding access rights, the access request is not transmitted downward to improve system security, while ensuring that only valid information is transmitted to the memory control module, thereby improving the quality and efficiency of data processing.

[0061] The memory control module directly interacts with the storage component 21. When the memory control module receives the target memory access request generated by the analysis and processing module, it performs a specific access operation on the storage component 21 according to the content of the target memory access request, such as reading, writing or modifying the data in the storage component 21. In this embodiment, by dividing the processing of the access request into two stages, namely, parsing and execution, the system design becomes more modular, reducing the overall complexity and facilitating maintenance and upgrading.

[0062] Based on the above embodiments, Figure 5 , the memory expansion board 2 is described in detail.

[0063] In an exemplary embodiment, the analysis and processing module includes a first hard core module 221, which is used to generate and output a first memory access request based on the access address and access type of the received extended access request; the target memory access request includes the first memory access request; the memory control module is specifically used to perform a corresponding access operation on the storage component 21 in response to the received first memory access request.

[0064] In an exemplary embodiment, the analysis and processing module also includes a second hard core module 222, which is used to generate and output a second memory access request based on the access address and access type of the received second host access request; the current access request also includes the second host access request sent by the server host, and the target memory access request also includes the second memory access request; the memory control module is specifically used to respond to the received second memory access request and perform a corresponding access operation on the storage component 21.

[0065] In this embodiment, a second hard core module 222 for data interaction with the server host is also included, so that the server host can directly access the storage component 21 on the memory expansion board 2 without going through other intermediate devices or complex communication paths, thereby reducing access delays. In addition, the task of data interaction is handled by the hard core module, which reduces the burden on the CPU of the server host, allowing the CPU to focus more on computing tasks and improve overall system efficiency. Among them, the first hard core module 221 and the second hard core module 222 are both hard core IP (Intellectual Property).

[0066] In this embodiment, the first hard core module 221 and the second hard core module 222 are both in CXL EP mode. In the CXL EP mode, the hard core module can be used as a memory expansion device, allowing the server host to directly access it.

[0067] In an exemplary embodiment, the analysis and processing module also includes: a network interface module 223, used to receive a remote access request; a network protocol stack 224, used to generate and output a third memory access request based on the access address and access type of the remote access request; the target memory access request also includes a third memory access request, and the current access request also includes a remote access request; the memory control module is specifically used to perform a corresponding access operation on the storage component 21 in response to the received third memory access request.

[0068] In this embodiment, the analysis and processing module also includes a network interface module 223, including but not limited to 400G MAC (Media Access Control) 1, 400G MAC2, 100G MAC, etc., for receiving a remote access request transmitted by the network optical module 26 on the memory expansion board 2. The network protocol stack 224 is specifically a protocol stack capable of implementing remote access to the storage component 21, which can be an Ethernet-based RDMA (Remote Direct Memory Access) protocol stack, which is used to implement the packetization and unpacking operations of the data packets transmitted between the memory expansion board 2 and the remote computing node, including but not limited to parsing the access address and access type of the remote access request, and generating and outputting a third memory access request based on the access address and access type of the remote access request. Using the Ethernet-based RDMA protocol stack, access to the storage component 21 can be initialized remotely, which greatly improves the shared memory capacity and reduces the data parallel splitting degree and data communication volume. Compared with the solution based on ordinary network cards, using the Ethernet-based RDMA protocol stack, the relevant drivers do not need to run on the local server, and the FPGA (Field-Programmable Gate Array) and GPU can be configured and initialized on the remote host through the network, making resource pooling more flexible.

[0069] Of course, in addition to the Ethernet-based RDMA protocol stack, the network protocol stack 224 can also be FCoE (Fibre Channel over Ethernet), iSCSI (Internet Small Computer System Interface), etc. The selection of which network protocol stack 224 depends on the specific network environment, performance requirements, cost considerations, and existing infrastructure, etc. This embodiment does not make any specific limitations here.

[0070] In an exemplary embodiment, the analysis and processing module further includes an access arbitration module 225 for arbitrating the received first memory access request and / or second memory access request and / or third memory access request, and outputting a target memory access request that has been successfully arbitrated.

[0071] In this embodiment, considering that the current access request has multiple senders, the analysis and processing module also includes an access arbitration module 225, which is used to arbitrate the received first memory access request and / or second memory access request and / or third memory access request to determine the processing order of each target memory access request. When multiple requests access the same memory resource at the same time, the arbitration module can avoid access conflicts and ensure data consistency and integrity, thereby improving the efficiency of memory access and reducing waiting time.

[0072] Specifically, the access arbitration module 225 determines the priority of the received first memory access request and / or the second memory access request and / or the third memory access request according to preset rules and strategies. Specifically, arbitration can be performed in the order of the time when the access arbitration module 225 receives the above-mentioned memory access requests. After arbitration, the winning request will be forwarded to the corresponding memory control module for subsequent data access operations, which is suitable for high-performance computing scenarios with high-concurrency access requests.

[0073] It is understandable that in different application scenarios, a priority may also be set for the sender of the above memory access request, such as a memory access request sent by a host will be processed first.

[0074] In an exemplary embodiment, the control component 22 further includes a control register module 226 for managing registers in the memory control module and the analysis processing module.

[0075] In an exemplary embodiment, the storage component 21 includes a non-volatile storage device 211; the memory control module includes: a shared memory controller 227, which is used to set the access state of the access address of the received target memory access request to a locked state, and forward the target memory access request; a memory access controller 228, which is used to perform corresponding access operations on the non-volatile storage device 211 according to the access address and access type of the target memory access request.

[0076] In this embodiment, the storage component 21 includes a non-volatile storage device 211, and the non-volatile storage device 211 has the characteristic that the stored data will not be lost after power failure, which means that in the model training process, even if a power failure occurs, the weight parameters and training data that have been loaded will not be lost, thereby avoiding the need to reload data. Since there is no need to reload the initial data at each startup, the computing resources of the GPU are more effectively utilized, which helps to reduce the overall time of model training, eliminates the step of data reloading, and simplifies the process of model training. In this embodiment, the memory control module includes a shared memory controller 227, which is used to parse the access address of the received target memory access request, and sets the access state of the access address to a locked state, ensuring that when processing the target memory access request, other senders cannot operate on the access address at the same time, maintaining the consistency and integrity of the data. After the access is completed, the unlocked state allows other senders to access, thereby improving the flexibility of the system. And after the target memory access request is processed, the access state of the access address is set to an unlocked state so that other senders can access the segment access address. In this embodiment, the shared memory controller 227 may specifically be a CXL.mem-based shared memory controller 227 .

[0077] The shared memory controller 227 is also used to forward the target memory access request to the memory access controller 228, so that the memory access controller 228 parses the access type and access address in the target memory access request and performs an access operation corresponding to the access type on the non-volatile storage device 211 according to the access address. The memory access controller 228 can be a DMA controller.

[0078] Among them, the non-volatile storage device 211 can be NVMe SSD (Non-Volatile Memory Express Solid State Drive), HDD (Hard Disk Drive), NAND Flash Memory (flash memory chip), EEPROM (Electrically Erasable Programmable Read-Only Memory), etc., which can be selected according to actual engineering needs, and this embodiment does not make specific limitations here.

[0079] In an exemplary embodiment, the storage component 21 also includes a memory particle 212; the memory access controller 228 is specifically used to parse and forward the access information in the target memory access request; the access information includes the access type and access address corresponding to the target memory access request; the memory control module also includes: a cache controller 229, which is used to respond to the access type in the access information as a read operation type. If the access address in the access information hits the cache, the corresponding response data is read from the memory particle 212 and returned to the memory access controller 228; otherwise, the access information is forwarded, and the response data returned by the storage controller 230 is cached in the memory particle 212; the storage controller 230 is used to read the corresponding response data in the non-volatile storage device 211 according to the access address in the access information sent by the cache controller 229 and return it.

[0080] In this embodiment, the storage component 21 also includes memory particles 212, which can be specifically DDR DIMM (Dual Inline Memory Module) memory bars. DDR is a memory particle 212 that can transmit data twice per clock cycle, i.e., sampling at the rising edge and the falling edge respectively. It is usually in the form of DIMM, with multiple memory particles 212 concentrated on a circuit board, and is applied to a server host or various acceleration cards using address, data and control buses. Similar packaging forms include RDIMM (Registered DIMM), UDIMM (Unbuffered DIMM), Mini-DIMM, etc.

[0081] Correspondingly, the memory control module also includes a cache controller 229 and a storage controller 230. The cache controller 229 is used to cache the non-volatile storage device 211 into the memory particle 212. Specifically, the cache controller 229 is used to first check whether the cache is hit when receiving access information of the read operation type, that is, whether the access address is in the memory particle 212. If it is hit, the corresponding response data is directly read from the memory particle 212 and returned to the memory access controller 228. If it is not hit, the request is forwarded to the storage controller 230, and the response data returned by the storage controller 230 is cached to improve the efficiency of subsequent access. The storage controller 230 is used to read the corresponding response data from the non-volatile storage device 211 according to the access address, and return the data to the cache controller 229, ensuring fast access to the data.

[0082] It can be understood that after the data in the non-volatile storage device 211 is first cached to the DDR DIMM memory stick, the DMA controller can perform single-byte access, which greatly improves the flexibility of access, and has higher flexibility for application scenarios that require random access or modification of a small amount of data. When processing non-continuous or random data access, the single-byte access capability can reduce unnecessary block read operations, thereby improving the efficiency of data access. Considering that frequent data access may cause wear and tear on the non-volatile storage device 211, the number of direct accesses to the non-volatile storage device 211 can be reduced through DDR caching, thereby extending its service life.

[0083] The memory access controller 228 is also used to return the response data to the shared memory controller 227, so that the response data is returned to the sender of the target memory access request through the shared memory controller 227; the sender is the server host or the processing board 1 or the remote computing node.

[0084] In an exemplary embodiment, the cache controller 229 is also used to forward the access information in response to the access type in the access information being a write operation type; the storage controller 230 is also used to write the access data in the access information into the non-volatile storage device 211 according to the access address in the access information sent by the cache controller 229.

[0085] In this embodiment, if the access type is a write operation type, the cache controller 229 and the storage controller 230 cooperate to write the access data in the access information into the corresponding access address in the non-volatile storage device 211 .

[0086] In an exemplary embodiment, the storage controller 230 is specifically used to determine a read address range including an access address based on the access address in the access information sent by the cache controller 229, and read and return the data to be cached in the non-volatile storage device 211 according to the read address range; the cache controller 229 is specifically used to respond to the access type in the access information being a read operation type, if the access address in the access information hits the cache, read the corresponding response data from the memory particle 212 and return it to the memory access controller 228; otherwise, forward the access information, cache the data to be cached in the memory particle 212, and return the response data in the data to be cached to the memory access controller 228.

[0087] In this embodiment, when the storage controller 230 reads data from the non-volatile storage device 211, it reads data of a preset length. For example, if the access address is 5, the storage controller 230 determines the read address range to be 0-10. This range is usually based on a preset cache line size or data block size to ensure data continuity and reading efficiency. The data to be cached is read in the non-volatile storage device 211 according to the read address range. The data to be cached includes the response data corresponding to the access address 5 and other data near the address. The data to be cached is returned to the cache controller 229, and the cache controller 229 stores the received data to be cached in the memory particle 212 (such as DDR). In this way, when subsequent access requests fall within the same read address range, data can be directly obtained from the cache without accessing the non-volatile storage device 211 again. By pre-reading and caching data, the system can quickly respond to subsequent access requests. Caching frequently accessed data in DDR can significantly reduce access delays and significantly improve data access efficiency and overall system performance.

[0088] In summary, the memory expansion board 2 and the processing board 1 in this embodiment are both built based on FPGA. The GPU card is directly connected to the FPGA-based memory expansion processing board 1 through an MCIO cable, which reduces the coupling degree between the GPU board and the server when using the extended memory, so that the GPU chip can directly mount and access the extended memory and NMVe solid-state hard disk; based on the GPU memory expansion technology of the present invention, there is no need to pass through the server motherboard PCIe Chipset (PCIe chipset) related chips, and the physical link has a lower access delay; compared with the solution based on the ordinary network card, the RDMA communication solution based on FPGA can not run the relevant driver on the local server, and can configure and manage the initialization of FPGA and GPU on the remote host through the network, and the resource pooling flexibility is higher; using the Ethernet-based RDMA protocol stack, the NMVe solid-state hard disk can be remotely initialized and accessed, and the NVMe solid-state hard disk is used as the extended memory that the GPU instance can access, which greatly improves the shared memory capacity and reduces the data parallel splitting degree and data communication volume; in the GPU instance based on FPGA, on one core of the GPU, a thread is used to implement the PCIe RC driver function, so that the GPU instance can directly access the mounted extended memory and NVMe solid-state hard disk. The memory expansion FPGA board and the GPU instance FPGA board are directly inserted into the standard server PCIe slot, and the two cards are directly connected through the MCIO high-speed differential cable, which is conducive to system integration and has the characteristics of high bandwidth and low latency; traditional GPU cards communicate with resources outside the node through the RDMA network card and PCIe Switch for P2P communication; the FPGA card of the present invention has a 400G optical network interface, which can be connected to the RoCEv2 network or standard Ethernet, and communicate with other computing nodes or storage nodes. The GPU can communicate with resources outside the node through the directly connected FPGA, without the traditional solution of connecting to the RDMA network card through the PCIe Switch.

[0089] In a specific application, taking the interconnection between a memory expansion card and a processing card as an example, the connection relationship is as follows: Figure 6As shown, the data flow when the host accesses the memory expansion card, the processing card accesses the memory expansion card, and the remote computing node accesses the memory expansion card and the solid-state storage is as follows: the memory expansion card automatically loads the configuration file to complete the initialization when it is powered on, the processing card is powered on, the GPU is initialized through the host system and the driver, the GPU core 1 starts the GPU thread 1 to load the RC bridge driver, and completes the link initialization of the extended memory EP device. The initial training data or model weight parameters can be loaded through the local PCIe to load the local storage training data (that is, the training data stored in the local memory of the server host) or the data received from the remote computing node through the 400G Ethernet RDMA protocol to the NVMe SSD mounted on the memory expansion card; other GPU cores start the model training program and access the DDR4 extended memory on the memory expansion board 2 card through the system bus 15 and the RC controller. The memory controller loads data from the NVMe solid-state storage to the DDR4 according to the cache algorithm, and then continuously updates the cache data in the DDR4 as the access address changes, so as to realize the transmission of intermediate training data in the model training process, which can reduce the time for the model to load weight parameters and training data, effectively expand the GPU computing memory, reduce the communication bottleneck, increase the GPU computing resource utilization, and shorten the model training time.

[0090] In a second aspect, the present invention further provides a computer system, comprising a server host and a data processing system as described in any one of the above embodiments.

[0091] For an introduction to a computer system provided in the present application, please refer to the above embodiments, and the present application will not go into details here.

[0092] The computer system provided in the present application has the same beneficial effects as the above-mentioned data processing system.

[0093] Third, please refer to Figure 7 The present invention also provides a data processing method, which is applied to a data processing system as described in any one of the embodiments above, wherein the data processing system includes a processing board and a memory expansion board provided with a storage component, and the data processing method includes: S101: identifying the storage component through a bridge driver thread by a processing core on the processing board, establishing a first mapping relationship between the address space of the processing board and the storage component, and forwarding an internal access request that complies with the first mapping relationship; the internal access request includes a first internal access request generated by the processing core; S102: converting the received internal access request into an extended access request and outputting it through a first controller on the processing board; S103: performing an access operation on the storage component in response to the extended access request through a control component on the memory expansion board.

[0094] It can be seen that in this embodiment, the present application provides a data processing method, which directly connects the memory expansion board and the processing board, reduces the coupling between the processing board and the server host when accessing the extended memory, and uses a thread on a processing core of the processing board to implement a bridge driver function, and identifies and addresses the processing components on the memory expansion board connected to itself, so that the processing board can directly access the storage components on the memory expansion board without the need for memory copying through the CPU. While effectively expanding the memory of the processing board, it shortens the access path of the processing board to the extended memory, reduces access delays, and thereby reduces communication bottlenecks during the training process of the pre-trained model.

[0095] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0096] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing system, characterized in that: The invention comprises a processing board, a memory expansion board and a connecting device, wherein the connecting device comprises a first connector provided on the processing board, a second connector provided on the memory expansion board, and a cable for connecting the first connector and the second connector; The processing board comprises: a processing core, configured to, when the first connector and the second connector are connected by the cable, send a signal or perform a detection operation through a bridge driver thread to identify a storage component, establish a first mapping relationship between the address space of the processing board and the storage component, and forward an internal access request that complies with the first mapping relationship; the internal access request includes a first internal access request generated by the processing core; A first controller, configured to convert the received internal access request into an extended access request and output the extended access request; The memory expansion board includes the storage component and a control component, and the control component is used to perform an access operation on the storage component in response to the extended access request.

2. The data processing system according to claim 1, characterized in that The processing board also includes a second controller, which is used to establish a second mapping relationship between the storage component and the address space of the server host, and forward the first host access request sent by the server host and conforming to the second mapping relationship; The first controller is further configured to convert the received first host access request into the extended access request and output the request.

3. The data processing system according to claim 2, characterized in that: The processing board further comprises a system bus for forwarding the internal access request to the first controller and / or forwarding the first host access request to the first controller.

4. The data processing system according to claim 1, characterized in that: The processing core is specifically used to generate internal access requests, identify storage components by sending signals through a bridge driver thread or performing detection operations, and establish a first mapping relationship between the address space of the processing board and the physical address of the storage component, forward internal access requests whose destination addresses match the first mapping relationship, and discard internal access requests whose destination addresses do not match the first mapping relationship.

5. The data processing system according to claim 1, characterized in that: The processing board also includes at least one computing core, which is used to generate a second internal access request corresponding to the current computing demand; the internal access request also includes the second internal access request.

6. The data processing system according to claim 1, characterized in that: The memory expansion board also includes a monitoring management module for detecting the operating status of the memory expansion board.

7. The data processing system according to any one of claims 1 to 6, characterized in that: The control component comprises: An analysis and processing module, configured to generate and output a target memory access request based on an access address and an access type of a received current access request; the current access request includes the extended access request; A memory control module is used to perform an access operation corresponding to the target memory access request on the storage component.

8. The data processing system according to claim 7, characterized in that: The analysis and processing module includes a first hard core module, which is used to generate and output a first memory access request based on an access address and an access type of a received extended access request; the target memory access request includes the first memory access request; The memory control module is specifically configured to execute a corresponding access operation on the storage component in response to the received first memory access request.

9. The data processing system according to claim 8, characterized in that: The analysis and processing module also includes a second hard core module, which is used to generate and output a second memory access request based on the access address and access type of the received second host access request; the current access request also includes the second host access request sent by the server host, and the target memory access request also includes the second memory access request; The memory control module is specifically configured to execute a corresponding access operation on the storage component in response to the received second memory access request.

10. The data processing system according to claim 9, characterized in that: The analysis and processing module also includes: A network interface module, used for receiving a remote access request; A network protocol stack, configured to generate and output a third memory access request based on an access address and an access type of the remote access request; the target memory access request also includes the third memory access request, and the current access request also includes the remote access request; The memory control module is specifically configured to execute a corresponding access operation on the storage component in response to the received third memory access request.

11. The data processing system according to claim 10, characterized in that: The analysis and processing module also includes an access arbitration module, which is used to arbitrate the received first memory access request and / or the second memory access request and / or the third memory access request, and output a target memory access request that has been successfully arbitrated.

12. The data processing system according to claim 10, characterized in that: The memory expansion board also includes a network optical module connected to the network interface module, which is used to receive the remote access request sent by the remote computing node.

13. The data processing system according to claim 7, characterized in that: The control component also includes a control register module for managing registers in the memory control module and the analysis and processing module.

14. The data processing system according to claim 7, characterized in that: The storage component includes a non-volatile storage device; The memory control module includes: A shared memory controller, configured to set the access state of the access address of the received target memory access request to a locked state, and forward the target memory access request; A memory access controller is used to perform corresponding access operations on the non-volatile storage device according to the access address and access type of the target memory access request.

15. The data processing system according to claim 14, characterized in that: The storage component also includes memory particles; The memory access controller is specifically used to parse and forward the access information in the target memory access request; the access information includes the access type and access address corresponding to the target memory access request; The memory control module also includes: A cache controller, configured to respond to the access type in the access information being a read operation type, if the access address in the access information hits the cache, read the corresponding response data from the memory particle and return it to the memory access controller; otherwise, forward the access information and cache the response data returned by the storage controller in the memory particle; The storage controller is used to read and return corresponding response data in the non-volatile storage device according to the access address in the access information sent by the cache controller.

16. The data processing system according to claim 15, characterized in that: The cache controller is further configured to forward the access information in response to the access type in the access information being a write operation type; The storage controller is further configured to write access data in the access information sent by the cache controller into the non-volatile storage device according to the access address in the access information.

17. The data processing system according to claim 15, characterized in that: The storage controller is specifically used to determine a read address range including the access address based on the access address in the access information sent by the cache controller, and read the data to be cached in the non-volatile storage device according to the read address range and return it; The cache controller is specifically used to respond to the access type in the access information being a read operation type. If the access address in the access information hits the cache, the corresponding response data is read from the memory particle and returned to the memory access controller; otherwise, the access information is forwarded, the data to be cached is cached in the memory particle, and the response data in the data to be cached is returned to the memory access controller.

18. The data processing system according to claim 17, characterized in that: The memory access controller is also used to return the response data to the shared memory controller, so that the response data is returned to the sender of the target memory access request through the shared memory controller; the sender is the server host or the processing board or the remote computing node.

19. A computer system, characterized in that: It comprises a server host and a data processing system as described in any one of claims 1-18.

20. A data processing method, characterized in that: Applied to a data processing system according to any one of claims 1 to 18, the data processing system comprising a processing board and a memory expansion board provided with a storage component, and further comprising a connecting device, the connecting device comprising a first connector provided on the processing board, a second connector provided on the memory expansion board, and a cable for connecting the first connector and the second connector, the data processing method comprising: After the first connector and the second connector are connected by the cable, the processing core on the processing board sends a signal or performs a detection operation to identify the storage component through a bridge driver thread, establishes a first mapping relationship between the address space of the processing board and the storage component, and forwards an internal access request that complies with the first mapping relationship; the internal access request includes a first internal access request generated by the processing core; Converting the received internal access request into an extended access request and outputting the request through a first controller on the processing board; The control component on the memory expansion board responds to the extended access request by executing an access operation on the storage component.

Citation Information

Patent Citations

  • Device and method for improving RDMA transmission efficiency

    CN112597094A

  • Method and device for expanding memory and related equipment

    CN115794669A

  • Memory expansion device, server and server cluster

    CN117807013A