Data processing system, method, device, medium and program product

By collecting the actual operation information of the accelerator card through the management device, predicting the memory demand and allocating the memory type, the problem of multiple accelerator cards competing for the host memory is solved, and the accelerator card's expanded memory is realized without affecting the system performance.

CN120144326BActive Publication Date: 2025-09-12LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510630314.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-12
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

When multiple accelerator cards request memory from the host at the same time, competition for resources such as bandwidth is likely to occur, resulting in increased data transmission delays and affecting system performance.

Method used

The management device collects the actual operation information of the accelerator card, predicts the memory demand, and allocates the memory type on demand, so that the memory device can be directly used as the extended memory of the accelerator card, avoiding host allocation.

Benefits of technology

Without affecting the system's performance, the memory of the accelerator card is expanded to avoid bandwidth resource competition and reduce data transmission delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144326B_ABST
    Figure CN120144326B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing system, method, device, medium and program product in the field of computer technology. In the present application, multiple accelerator cards are directly connected to multiple memory devices through a management device, and the actual operation information of each accelerator card is collected through the management device to allocate memory space to each accelerator card on demand according to the actual operation information of the accelerator card. In other words, each memory device is not perceived by the host, and these memory devices are directly used as the extended memory of each accelerator card, so there is no need for the host to allocate memory for each accelerator card, thereby avoiding the competition for resources such as bandwidth that is easy to occur when multiple accelerator cards apply for memory at the same time, and data transmission delay is not easy to occur, thereby achieving the purpose of expanding the memory for the accelerator card without affecting the system operation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing system, method, device, medium and program product. Background Art

[0002] The demand for accelerator card memory capacity is increasing. Host memory can generally be used as extended memory for accelerator cards, thereby expanding the accelerator card memory. However, host memory capacity is limited. To provide more memory space for accelerator cards, host extended memory can be provided to accelerator cards. However, the host needs to allocate memory to each accelerator card. If multiple accelerator cards request memory from the host at the same time, bandwidth and other resources are likely to be competed for, resulting in increased data transmission delays and affecting overall performance.

[0003] Therefore, how to expand the memory for the accelerator card without affecting the system performance is a problem that those skilled in the art need to solve. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a data processing system, method, device, medium and program product to expand the memory of the accelerator card without affecting the system performance.

[0005] In a first aspect, the present application provides a data processing system, comprising: a management device, multiple accelerator cards, and multiple memory devices; the multiple accelerator cards are connected to the multiple memory devices through the management device; the management device is used to: collect actual operating information of each accelerator card; the management device is also used to: for any target accelerator card among the multiple accelerator cards, predict the memory requirement of the target accelerator card based on the actual operating information of the target accelerator card, select at least one memory type, and determine the memory allocation result of the target accelerator card based on the memory requirement and the at least one memory type; the target accelerator card is used to: update the memory configuration according to the memory allocation result.

[0006] In a second aspect, the present application provides a data processing method, which is applied to a management device, wherein the management device is connected to multiple accelerator cards and multiple memory devices, including: collecting actual operating information of each accelerator card; for any target accelerator card among the multiple accelerator cards, predicting the memory requirement of the target accelerator card based on the actual operating information of the target accelerator card, and selecting at least one memory type; determining a memory allocation result of the target accelerator card based on the memory requirement and the at least one memory type, so that the target accelerator card updates the memory configuration according to the memory allocation result.

[0007] In a third aspect, the present application provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the aforementioned disclosed data processing method.

[0008] In a fourth aspect, the present application provides a non-volatile storage medium for storing a computer program, wherein the computer program implements the aforementioned disclosed data processing method when executed by a processor.

[0009] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which implements the steps of the aforementioned disclosed data processing method when executed by a processor.

[0010] From the above scheme, it can be seen that the present application provides a data processing system, including: a management device, multiple accelerator cards and multiple memory devices; the multiple accelerator cards are connected to the multiple memory devices through the management device; the management device is used to: collect actual operation information of each accelerator card; the management device is also used to: for any target accelerator card among the multiple accelerator cards, predict the memory requirement of the target accelerator card according to the actual operation information of the target accelerator card, and select at least one memory type, and determine the memory allocation result of the target accelerator card based on the memory requirement and the at least one memory type; the target accelerator card is used to: update the memory configuration according to the memory allocation result.

[0011] It can be seen that the beneficial effects of this application are: multiple accelerator cards are directly connected to multiple memory devices through a management device, and the management device collects the actual operating information of each accelerator card, so that memory space is allocated to each accelerator card on demand based on the actual operating information of the accelerator card. In other words, each memory device is not perceived by the host, and these memory devices directly serve as the extended memory of each accelerator card, so there is no need for the host to allocate memory for each accelerator card. This avoids the competition for resources such as bandwidth that is prone to occur when multiple accelerator cards apply for memory at the same time, and is less likely to cause data transmission delays, thus achieving the purpose of expanding the memory of the accelerator card without affecting the system's operating performance.

[0012] Correspondingly, the data processing method, device, medium and program product provided by this application also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0014] Figure 1 A schematic diagram of a data processing system disclosed in this application;

[0015] Figure 2 A flow chart of a data processing method disclosed in this application;

[0016] Figure 3 A schematic diagram of another data processing system disclosed in this application;

[0017] Figure 4 Disclosed in this application Figure 3 Corresponding system logic diagram;

[0018] Figure 5 A schematic diagram of an electronic device disclosed in this application;

[0019] Figure 6 A server structure diagram provided for this application;

[0020] Figure 7 This is a terminal structure diagram provided for this application. DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other examples obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] At present, in order to provide more memory space for the accelerator card, the host's extended memory can be provided to the accelerator card for use. However, the host needs to allocate memory for each accelerator card. If multiple accelerator cards apply for memory from the host at the same time, it is easy for them to compete for resources such as bandwidth, resulting in increased data transmission delays and affecting overall performance. To this end, the present application provides a data processing solution that can prevent each memory device from being perceived by the host. These memory devices are directly used as extended memory for each accelerator card, so there is no need for the host to allocate memory for each accelerator card. This avoids the competition for resources such as bandwidth that is easy to occur when multiple accelerator cards apply for memory at the same time, and is less likely to cause data transmission delays, thus achieving the purpose of expanding memory for the accelerator card without affecting the system's operating performance.

[0023] See also Figure 1 As shown, an embodiment of the present application discloses a data processing system, including: a management device, multiple acceleration cards: acceleration card 1 to acceleration card N, and multiple memory devices: memory device 1 to memory device N; multiple acceleration cards are connected to multiple memory devices through the management device.

[0024] The management device is configured to collect actual operating information of each accelerator card. The management device is further configured to, for any target accelerator card among the multiple accelerator cards, predict the target accelerator card's memory requirement based on the target accelerator card's actual operating information, select at least one memory type, and determine a memory allocation result for the target accelerator card based on the memory requirement and the at least one memory type. The target accelerator card is configured to update its memory configuration based on the memory allocation result.

[0025] It should be noted that each memory device can provide the same or different memory types. Furthermore, a memory device can provide only DRAM (Dynamic Random Access Memory) resources or SSD (Solid State Drive) resources, or both DRAM and SSD resources. Therefore, memory types can be DRAM, SSD, and so on.

[0026] The memory allocation result may record: which memory device's memory address portion is allocated to the target accelerator device, which memory type this memory address portion belongs to, etc. Specifically, the memory allocation result may include: the identification information of the memory device allocated to the target accelerator card, the device port, the port of the management device to which the device port is connected, the corresponding address segment, the memory type of the address segment, etc. Therefore, in one embodiment, the management device is configured to: determine the identification information of the memory device allocated to the target accelerator card based on the memory allocation result; and record the mapping relationship between the port of the identification information (i.e., the port of the management device to which the memory device port is connected) and the port of the target accelerator card (i.e., the port of the management device to which the target accelerator card port is connected) in a preset routing table.

[0027] Accordingly, the target accelerator card is configured to: determine the identification information and corresponding address segments of the memory devices allocated to the target accelerator card based on the memory allocation result; and record the identification information and address segments in a preset memory mapping table. The preset memory mapping table is a table in the accelerator card that records the available memory addresses. Each accelerator card has a memory mapping table.

[0028] In this embodiment, the management device can collect actual operating information from each accelerator card to provide real-time insights into each accelerator card's actual usage, such as the accelerator card's current memory usage (which may include memory usage rate and memory usage size), and the accelerator card's currently pending task queue information (which may include the type and number of unfinished tasks). Therefore, in one embodiment, the management device is configured to collect memory usage information, task queue information, and throughput from each accelerator card as the corresponding accelerator card's actual operating information. Task types may include model inference tasks, model training tasks, batch data processing, and background data backup.

[0029] Accordingly, different task types may correspond to different priorities and resource allocation rules, for details, see the resource allocation table shown in Table 1. Therefore, in one embodiment, the management device is configured to query the preset resource allocation table for the unfinished task types and their corresponding allocatable memory types in the task queue information to obtain at least one memory type.

[0030] Table 1

[0031]

[0032] If tasks of different priorities are running simultaneously in an accelerator card, the accelerator card can use a weighted fair queue algorithm to allocate bandwidth to tasks of different priorities in the target accelerator card. If the current available memory of the target accelerator card is lower than the memory required by the highest priority task when the highest priority task starts running, the lowest priority task in the target accelerator card is suspended, and the current status information of the lowest priority task is stored. If the management device finds that tasks of different priorities are running in the accelerator cards that apply for memory at the same time, the management device can also use a weighted fair queue algorithm to allocate bandwidth to different accelerator cards for their memory applications. In addition, when the management device finds that the amount of memory available for the accelerator card where the high-priority task is located is small among all current memory devices, the memory application of the accelerator card where the low-priority task is located can be suspended, and the current status information of its memory application can be stored so that the memory application of the accelerator card where the low-priority task is located can be continued when more memory is available.

[0033] In one example, corresponding trigger conditions can be set for the memory allocation of the accelerator card. For example: the management device is used to: if the memory occupancy information exceeds a preset first threshold value, and the number of unfinished tasks in the task queue information exceeds a preset second threshold value, then the memory demand for the target accelerator card is predicted. The prediction method can be: using a trained prediction model with the historical memory occupancy information of the accelerator card, the type of unfinished tasks in the task queue information, and the number of unfinished tasks as the data basis. Therefore, in one embodiment, the management device is used to: query the historical memory occupancy information of the target accelerator card; and predict the memory demand based on the historical memory occupancy information, the type of unfinished tasks in the task queue information, and the number of unfinished tasks. The management device can record the memory occupancy information collected each time for each accelerator card, thereby forming the historical memory occupancy information of an accelerator card.

[0034] To facilitate the management of each memory device, a memory record table can be set up in the management device to record the memory type (DRAM resource type, SSD resource type, etc.), total memory capacity, used memory, remaining available memory, and even which portion of each memory device's memory address is allocated to which accelerator card. Therefore, in one embodiment, the management device is used to: query the remaining available memory and memory type of each memory device in a preset memory record table to obtain a query result; determine a memory allocation result based on the query result, memory demand, and at least one memory type; based on the memory allocation result, determine the identification information and corresponding address segment of the memory device allocated to the target accelerator card; and update the memory record table based on the identification information and address segment.

[0035] The memory capacity of each memory device can be allocated to each accelerator card for use, and can also be released by each accelerator card, so that flexible use of extended memory can be achieved. In one embodiment, the management device is used to: if the memory occupancy information in the target accelerator card is lower than the preset third threshold, then the target size of memory resources is recovered from the target accelerator card, and the memory record table and routing table are updated accordingly, specifically: the mapping relationship between the port of the target accelerator card and the port of the corresponding memory device recorded in the routing table is deleted, and the used memory amount, remaining available memory amount, etc. of the corresponding memory device in the memory record table are changed. The target size is an integer multiple of the set memory block size, wherein the set memory block size is the preset minimum memory size that can be applied each time, such as set to 4K, 1G, etc. Correspondingly, the target accelerator card is used to: update the memory mapping table according to the target size of memory resources, specifically: delete the mapping relationship between the port of the target accelerator card and the port of the corresponding memory device recorded in the memory mapping table.

[0036] In one embodiment, a management device includes a management controller and a switch controller. The management controller connects to the switch controller via a first protocol (e.g., PCIe) and a second protocol (e.g., I2C). The switch controller connects to multiple accelerator cards and multiple memory devices via a third protocol (e.g., CXL). The management controller connects to a host computer and multiple accelerator cards via a fourth protocol (e.g., PCIe). PCIe (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard used to connect various peripherals and expansion cards within a computer, enabling high-bandwidth, low-latency data transmission. The management controller uses I2C (Inter-Integrated Circuit Bus), a serial bus protocol for connecting integrated circuits, to perform management operations on the switch controller, multiple accelerator cards, and multiple memory devices. It also collects operational information from multiple accelerator cards and multiple memory devices via PCIe. The switch controller uses CXL (Compute ExpressLink), an interconnect protocol for high-performance computing, to enable memory read / write and memory allocation between multiple accelerator cards and multiple memory devices, improving efficiency. The management controller connects the host and multiple accelerator cards via PCIe, building a management network. This management network allows the host to dispatch various tasks to each accelerator card, enabling the host to manage the devices of each accelerator card. CXL is a high-speed interface protocol that enables data exchange through low-latency, high-bandwidth connections. It includes three dynamically multiplexed sub-protocols: CXL.io, CXL.cache, and CXL.mem. CXL.io is similar to PCIe's IO protocol and is enabled for operations such as discovery and enumeration, error reporting, and host physical address lookup. CXL.cache is a protocol for accessing caches that defines interactions between processors and devices, allowing connected CXL devices to use a request and response method to efficiently cache processor memory with extremely low latency. CXL.mem is a protocol for accessing memory that uses load and store commands to provide the processor with access to device-attached memory. The processor acts as the master and the CXL device as the slave, supporting both volatile and persistent memory architectures. Based on this, each accelerator card can perform a read operation and / or a write operation on a memory device connected to a port that has a mapping relationship with a port of the target accelerator card through the management device.

[0037] As can be seen, in this application, multiple accelerator cards are directly connected to multiple memory devices through a management device, and the management device collects the actual operating information of each accelerator card to allocate memory space to each accelerator card on demand based on the actual operating information of the accelerator card. In other words, each memory device is not perceived by the host computer, and these memory devices directly serve as the extended memory of each accelerator card, so there is no need for the host computer to allocate memory for each accelerator card. This avoids the competition for resources such as bandwidth that is prone to occur when multiple accelerator cards apply for memory at the same time, and is less likely to cause data transmission delays, thus achieving the purpose of expanding the memory of the accelerator card without affecting the system's operating performance.

[0038] The following introduces a data processing method provided in an embodiment of the present application. The data processing method described below can be referenced with other embodiments described in this document.

[0039] See also Figure 2 As shown, an embodiment of the present application discloses a data processing method, which is applied to a management device, where the management device is connected to multiple acceleration cards and multiple memory devices.

[0040] The method provided in this embodiment includes:

[0041] S201: Collect actual operation information of each accelerator card.

[0042] S202: For any target accelerator card among the multiple accelerator cards, predict the memory requirement of the target accelerator card according to actual operation information of the target accelerator card, and select at least one memory type.

[0043] S203: Determine a memory allocation result of the target accelerator card based on the memory requirement and at least one memory type, so that the target accelerator card updates a memory configuration according to the memory allocation result.

[0044] In one embodiment, collecting actual operation information of each accelerator card includes collecting memory usage information, task queue information, and throughput in each accelerator card as the actual operation information of the corresponding accelerator card.

[0045] In one embodiment, if the memory usage information exceeds a preset first threshold and the number of unfinished tasks in the task queue information exceeds a preset second threshold, a memory requirement prediction is performed for the target accelerator card, and the following steps are performed: the memory requirement of the target accelerator card is predicted based on actual operating information of the target accelerator card, and at least one memory type is selected; and a memory allocation result of the target accelerator card is determined based on the memory requirement and the at least one memory type, so that the target accelerator card updates the memory configuration according to the memory allocation result.

[0046] In one embodiment, the memory requirement of the target accelerator card is predicted based on the actual operation information of the target accelerator card, including: querying the historical memory usage information of the target accelerator card; and predicting the memory requirement based on the historical memory usage information and the type and number of unfinished tasks in the task queue information.

[0047] In one embodiment, selecting at least one memory type includes: querying a preset resource allocation table for unfinished task types and their corresponding allocatable memory types in task queue information to obtain at least one memory type.

[0048] In one embodiment, determining a memory allocation result for a target accelerator card based on memory requirements and at least one memory type includes: querying a preset memory record table for the remaining available memory and memory type of each memory device to obtain a query result; determining a memory allocation result based on the query result, the memory requirements, and the at least one memory type; and then, further determining identification information and a corresponding address segment of the memory device allocated to the target accelerator card based on the memory allocation result; and updating the memory record table based on the identification information and address segment.

[0049] In one embodiment, the management device further determines identification information of the memory device allocated to the target accelerator card based on the memory allocation result; and records the mapping relationship between the port of the identification information and the port of the target accelerator card in a preset routing table to update the routing table.

[0050] In one embodiment, the target accelerator card updates the memory configuration according to the memory allocation result, including: determining the identification information and the corresponding address segment of the memory device allocated to the target accelerator card based on the memory allocation result; recording the identification information and the address segment in a preset memory mapping table to update the memory mapping table.

[0051] In one embodiment, if the management device determines that the memory usage information in the target accelerator card is lower than a preset third threshold, the target size of memory resources is reclaimed from the target accelerator card, and the memory record table and routing table are updated accordingly; the target size is an integer multiple of the set memory block size; accordingly, the target accelerator card is used to: update the memory mapping table according to the target size of memory resources.

[0052] In one embodiment, if the current available memory capacity of the target accelerator card is lower than the memory capacity required by the highest-priority task when the highest-priority task begins running, the lowest-priority task in the target accelerator card is suspended, and the current status information of the lowest-priority task is stored. The target accelerator card may also utilize a weighted fair queuing algorithm to allocate bandwidth to tasks of different priorities in the target accelerator card. A management device is used to perform read and / or write operations on a memory device connected to a port mapped to a port of the target accelerator card.

[0053] Among them, for more specific working processes of each step in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0054] As can be seen, this embodiment provides a data processing method in which multiple accelerator cards are directly connected to multiple memory devices through a management device. The management device collects actual operating information of each accelerator card and allocates memory space to each accelerator card on demand based on the actual operating information. In other words, each memory device is not perceived by the host computer; it directly serves as extended memory for each accelerator card. This eliminates the need for the host computer to allocate memory for each accelerator card. This avoids the contention for resources such as bandwidth that can occur when multiple accelerator cards simultaneously request memory, reduces data transmission delays, and achieves the goal of expanding memory for accelerator cards without affecting system performance.

[0055] See Figure 3 In a data processing system, management devices include a management node (i.e., a management controller) and a CXL Switch (a switch controller). The management node connects to one port of the CXL Switch via the PCIe protocol. Other ports of the CXL Switch connect to multiple custom accelerator cards and CXL memory expansion devices. The management node also connects to the CXLSwitch via I2C, enabling rapid management of these cards, such as configuring memory devices. The management node and the custom accelerator cards form a management network, which connects to hosts, enabling them to distribute tasks and manage accelerator cards within the network.

[0056] Specifically, the CXL Switch's multiple uplink ports connect to N-1 customized accelerator cards, which integrate logic for memory request and allocation. These cards logically constitute the compute subsystem. The CXL Switch's multiple downlink ports connect to M CXL memory expansion devices, which use DRAM or SSDs and support dynamic media type identification and mixed addressing. The uplink and downlink port ratios can be flexibly adjusted based on needs. A typical configuration is 8 uplink ports (7 customized accelerator cards + 1 management node) and 8 downlink CXL memory expansion devices, supporting up to 16 interconnects. Fabric Management (FM), the resource management and control software in the management node, is responsible for system topology management, resource scheduling, and policy coordination. It can run on a separate management node or be embedded in the CXL Switch controller.

[0057] The management node connects to the CXL switch through a PCIe port, while the accelerator card and CXL memory expansion device connect to the CXL switch through a CXL port. The management node and the BMC in the CXL memory expansion device jointly control and manage the CXL switch, implementing a CXL memory expansion pool. The management node also monitors the local memory usage of each accelerator card and triggers the dynamic application, allocation, and release of expanded memory when appropriate.

[0058] Reference Figure 3 The system, correspondingly the logic control system can be Figure 4 . Figure 3 The accelerator cards shown can be configured Figure 4 The computing subsystem in Figure 3 The memory devices and accelerator card memories shown can constitute Figure 4 The storage subsystem in Figure 3 The management node, CXL Switch, and related modules responsible for address management in the accelerator card (such as the simplified firmware core) can constitute Figure 4 The CXL subsystem in Figure 3 PCIe port configuration shown Figure 4 The computing subsystem, CXL subsystem, PCIe subsystem, and storage subsystem are interconnected through the system bus.

[0059] The CXL subsystem includes a simplified firmware core and a CXL root complex.

[0060] The CXL root complex serves as a bridge between the compute subsystem and CXL devices (such as CXL switches and CXL memory expansion devices). It is responsible for establishing and maintaining physical and logical connections between the management node and CXL devices, recording the memory address range of each CXL device's root port, processing memory requests from accelerator cards, converting them into the format required by the CXL protocol, and forwarding these requests to the corresponding CXL memory expansion devices. The simplified firmware core supports and enhances the root complex's functionality, enabling the system to more efficiently manage memory and data transfers. It is responsible for initializing storage devices connected to the root complex, initializing the Host Managed Device Memory (HDM) decoder in the accelerator card, and managing the associated physical memory address ranges to ensure the system can correctly access and use these memory resources. CXL endpoints are identified by examining the endpoint's configuration space and PCIe BAR address, and the HDM registry is used to summarize the memory address space of each endpoint. The compute subsystem includes the accelerator card's main compute core or stream processor. The PCIe subsystem primarily consists of PCIe ports and is responsible for establishing communication with the management node or host. The storage subsystem also includes a DDR controller and a DDR storage unit in the accelerator card.

[0061] Reference Figure 3 and Figure 4 As shown in Figure 1, the management node dynamically allocates extended memory based on the actual needs of the accelerator card, and the accelerator card can read and write to the extended memory. For example, when the local memory usage of accelerator card 1 exceeds the threshold, the FM dynamically allocates 2GB of DRAM resources from the extended memory pool to the accelerator card through the CXL switch. The accelerator card then directly accesses this memory through the CXL subsystem.

[0062] The management node is connected to the CXL Switch uplink port, monitors the memory usage of each accelerator card in real time, and dynamically allocates memory through FM. The specific steps include the following:

[0063] Step 1: Data collection and monitoring.

[0064] The management node collects the local memory usage of each accelerator card through the management network, such as accelerator card 1's local memory usage of 95% and accelerator card 2's local memory usage of 60%, as well as task queue information, such as accelerator card 1's task queue loading a large batch of data.

[0065] Step 2: The memory expansion policy is triggered.

[0066] Trigger an expansion request based on preset rules. For example, if the accelerator card's local memory usage or the predicted accelerator card's memory usage exceeds 90% and the task currently running on the accelerator card has a high priority, an expansion request is triggered.

[0067] In some implementations, FM can employ a memory allocation strategy based on load forecasting, triggering scaling policies based on the forecast results. The forecasting model is pre-trained and outputs future memory requirements. Data such as historical accelerator card memory usage (based on a time window), task type characteristics (such as model training, image inference, and batch data processing), queue depth, and throughput (such as the number of instructions queued per processor core) can be collected to train the forecasting model. Models such as LSTM (Long Short Term Memory) and random forest models can be used. After training, the trained forecasting model is used to predict future memory requirements, triggering scaling policies based on the forecast results.

[0068] In some implementations, FM can use a priority scheduling algorithm to balance resource contention among multiple accelerator cards. As shown in Table 1, memory resources are allocated and prioritized based on the types of tasks running on different accelerator cards. Specifically, higher-priority tasks are allocated more bandwidth to ensure low latency, while lower-priority tasks share the remaining bandwidth. Algorithms such as weighted fair queuing can be used for scheduling. Resource preemption policies can also be set. For example, when the highest-priority task is triggered, it can preempt the extended memory already allocated to lower-level tasks and save the intermediate states of the lower-level tasks to the SSD.

[0069] Step 3: Address space allocation.

[0070] The FM on the management node checks the global memory pool status, selects available memory areas (based on the memory table, locating specific available memory devices), generates allocation instructions, and allocates the corresponding memory areas to the corresponding accelerator cards. The expansion results are sent to the simplified firmware core in the corresponding accelerator card via the management network.

[0071] Step 4: HDM decoder update.

[0072] The accelerator card modifies the HDM decoder by simplifying the firmware core, expanding its physical address range to cover the newly expanded memory area. Specifically, the HDM decoder configuration parameters can be modified to enable it to recognize and access the expanded physical memory area. For example, the base address and size in the HDM decoder's register definition file can be modified to adjust the decoder's supported upper physical address limit. The device memory mapping table is updated, and the BARs in the PCIe configuration space are adjusted to expand the PCIe BAR address window, allowing a larger range of memory to be mapped to the accelerator card.

[0073] Step 5: Synchronize the CXL Switch routing table.

[0074] The CXL Switch updates the port mapping in the routing table and binds the downlink port connected to the corresponding memory device to the uplink port corresponding to accelerator card 1 to ensure that the request is correctly routed.

[0075] In one embodiment, the memory release steps are as follows:

[0076] Step 1: Data collection and monitoring.

[0077] The management node collects the local memory usage and task queue information of each accelerator card through the management network.

[0078] Step 2: The memory release policy is triggered.

[0079] Memory release requests are triggered at the appropriate time based on preset rules. For example, if the local memory usage is below 50% and the task priority is low, a release request is triggered.

[0080] Step 3: Address space release.

[0081] The FM of the management node checks the status of the global memory pool, generates a release instruction to reclaim the extended memory allocated to the accelerator card, and sends the reclaim result to the simplified firmware core of the corresponding accelerator card through the management network.

[0082] Step 4: HDM decoder update.

[0083] The simplified firmware core of the accelerator card modifies the HDM decoder of the accelerator card and reduces its physical address range to remove the addresses of the released extended memory area.

[0084] Step 5: Synchronize the CXL Switch routing table.

[0085] The CXL Switch updates the port mapping relationship, unbinding the downlink port of the corresponding memory device from the uplink port corresponding to the accelerator card, and releasing the routing relationship between the two.

[0086] In one embodiment, the process of the accelerator card reading the extended memory is as follows:

[0087] Step 1: Initiate a read request: The accelerator card computing unit generates a read operation.

[0088] Step 2: The storage subsystem determines the read operation address and finds the DDR range, and transfers it to the CXL subsystem for processing.

[0089] Step 3: The root complex resolves the address, determines that the DDR range belongs to the mapped range of the extended memory, converts the request into CXL protocol format, and adds the target endpoint and transaction tag.

[0090] Step 4: The CXL Switch routes the read request to the extended memory connected to the corresponding downstream port based on the port mapping table.

[0091] Step 5: The memory extension performs a read operation, returning the data block from the target address.

[0092] Step 6: The response data is transmitted back to the accelerator card's root complex through the CXL Switch, converted to the accelerator card's local bus format, and then sent to the accelerator card's compute unit.

[0093] In one embodiment, the process of the accelerator card writing to the extended memory is as follows:

[0094] Step 1: Initiate a write request: The accelerator card computing unit generates a write operation.

[0095] Step 2: The storage subsystem determines the write operation address and finds the DDR range, and transfers it to the CXL subsystem for processing.

[0096] Step 3: The root complex resolves the address, determines that the DDR range belongs to the mapped range of the extended memory, converts the request into CXL protocol format, and adds the target endpoint and transaction tag.

[0097] Step 4: The CXL Switch routes the read request to the extended memory connected to the downstream port based on the port mapping table.

[0098] Step 5: The memory expansion performs a write operation, writes the data to the target address, and returns a confirmation signal.

[0099] Step 6: The confirmation signal is transmitted back to the accelerator card's root complex through the CXL Switch, converted to the accelerator card's local bus format, and then notified to the accelerator card's compute unit.

[0100] As can be seen in this embodiment, a memory expansion pool shared by multiple accelerator cards is constructed based on the CXL high-speed interconnect protocol, which can realize dynamic memory expansion of the accelerator card and improve the computing resource utilization of the accelerator card in computing scenarios such as machine learning. This solution design can bypass the host memory and realize direct memory expansion of the accelerator card, reducing data handling and effectively reducing the latency of extended memory access. The management node can dynamically allocate extended memory resources based on real-time monitoring of accelerator card memory usage and subsequent task queues, which can effectively avoid idle resources and improve the overall resource utilization of the system.

[0101] In terms of architectural design, by integrating the CXL subsystem (including the root complex and simplified firmware core) with a single-layer CXLSwitch topology, a shared extended memory pool is constructed for multiple accelerator cards, enabling direct expansion without requiring host memory. The root complex is responsible for translating and routing local memory requests to and from the CXL protocol. The simplified firmware core dynamically manages the address space and HDM decoder. Combined with the switch's flexible port mapping, this resolves traditional PCIe bus contention and improves bandwidth utilization for extended memory.

[0102] A collaborative control mechanism for dynamic memory capacity management is provided for memory allocation. The management node monitors accelerator card memory usage in real time to trigger memory expansion or release. This, combined with dynamic adjustments to the HDM decoder (expanding or reducing the physical address range) and synchronized updates to the CXL Switch routing table, enables on-demand allocation of pooled memory. This mechanism, through firmware and hardware collaboration (such as FM global scheduling and address space binding and unbinding), ensures dynamic and flexible scheduling of expanded memory resources across multiple accelerator cards, avoiding storage resource waste.

[0103] The following describes an electronic device provided in an embodiment of the present application. The electronic device described below can be cross-referenced with other embodiments described herein. The electronic device can be any device described in the aforementioned embodiments, such as a management device, an accelerator card, etc.

[0104] See also Figure 5 As shown, an embodiment of the present application discloses an electronic device, including: a memory 501 for storing a computer program; a processor 502 for executing the computer program to implement the method disclosed in any of the above embodiments.

[0105] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: collect actual operating information of each accelerator card; for any target accelerator card among the multiple accelerator cards, predict the memory requirement of the target accelerator card based on the actual operating information of the target accelerator card, and select at least one memory type; and determine a memory allocation result for the target accelerator card based on the memory requirement and the at least one memory type, so that the target accelerator card updates the memory configuration according to the memory allocation result.

[0106] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: collecting memory usage information, task queue information and throughput in each accelerator card as actual operation information of the corresponding accelerator card.

[0107] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: if the memory usage information exceeds a preset first threshold and the number of unfinished tasks in the task queue information exceeds a preset second threshold, the memory demand is predicted for the target accelerator card.

[0108] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: query the historical memory usage information of the target accelerator card; predict the memory demand based on the historical memory usage information, the type of unfinished tasks and the number of unfinished tasks in the task queue information.

[0109] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: in a preset resource allocation table, query the unfinished task types and their corresponding allocatable memory types in the task queue information to obtain at least one memory type.

[0110] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: query the remaining available memory and memory type of each memory device in a preset memory record table to obtain a query result; determine a memory allocation result based on the query result, memory demand and at least one memory type; based on the memory allocation result, determine the identification information and corresponding address segment of the memory device allocated to the target accelerator card; and update the memory record table based on the identification information and address segment.

[0111] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: determine the identification information of the memory device allocated to the target accelerator card based on the memory allocation result; and record the mapping relationship between the port of the identification information and the port of the target accelerator card in a preset routing table.

[0112] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: based on the memory allocation result, determine the identification information and corresponding address segment of the memory device allocated to the target accelerator card; and record the identification information and address segment in a preset memory mapping table.

[0113] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: if the memory usage information in the target accelerator card is lower than a preset third threshold, reclaiming memory resources of a target size from the target accelerator card and updating the memory record table and routing table accordingly; the target size is an integer multiple of the set memory block size.

[0114] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: updating a memory mapping table according to the target size of the memory resource.

[0115] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: if, when the highest priority task starts running, the current available memory amount of the target accelerator card is lower than the memory amount required by the highest priority task, the lowest priority task in the target accelerator card is paused, and the current status information of the lowest priority task is stored.

[0116] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: using a weighted fair queuing algorithm to allocate bandwidth to tasks of different priorities in the target accelerator card.

[0117] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: performing a read operation and / or a write operation on a memory device connected to a port that has a mapping relationship with a port of the target accelerator card through a management device.

[0118] Furthermore, the embodiment of the present application also provides an electronic device. The electronic device can be Figure 6 The server shown can also be Figure 7 The terminal shown. Figure 6 and Figure 7 Each of the diagrams is a structural diagram of an electronic device according to an exemplary embodiment, and the contents in the diagrams cannot be considered as any limitation on the scope of use of the present application.

[0119] Figure 6 This is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, which is loaded and executed by the processor to implement the relevant steps of the data processing disclosed in any of the aforementioned embodiments.

[0120] In this embodiment, the power supply is used to provide operating voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world. The specific interface type can be selected according to specific application needs and is not specifically limited here.

[0121] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.

[0122] The operating system is used to manage and control the hardware devices and computer programs on the server, enabling the processor to operate and process data in the memory. It can be Windows Server, NetWare, Unix, Linux, etc. In addition to computer programs capable of performing the data processing methods disclosed in any of the aforementioned embodiments, computer programs can also include computer programs capable of performing other specific tasks. Data can include data such as application update information and other data such as application developer information.

[0123] Figure 7This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may specifically include but is not limited to a smartphone, tablet computer, laptop computer or desktop computer.

[0124] Generally, the terminal in this embodiment includes: a processor and a memory.

[0125] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor may be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content required to be displayed on the display. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0126] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is used to store at least the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the data processing method performed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to update information of the application.

[0127] In some embodiments, the terminal may further include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0128] Those skilled in the art will understand that Figure 7The structure shown in the figure does not constitute a limitation to the terminal, and may include more or fewer components than shown in the figure.

[0129] A non-volatile storage medium provided in an embodiment of the present application is introduced below. The non-volatile storage medium described below can be referenced with other embodiments described herein.

[0130] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the data processing method disclosed in the aforementioned embodiment. The non-volatile storage medium is a computer-readable non-volatile storage medium that, as a carrier for resource storage, may be a read-only memory, random access memory, a magnetic disk, or an optical disk. The resources stored thereon include an operating system, a computer program, and data, and the storage method may be either temporary or permanent.

[0131] A computer program product provided in an embodiment of the present application is introduced below. The computer program product described below can be referenced with other embodiments described herein.

[0132] A computer program product comprises a computer program / instruction, which implements the steps of the aforementioned data processing method when executed by a processor.

[0133] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0134] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.

[0135] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A data processing system, characterized in that: include: A management device, multiple accelerator cards, and multiple memory devices; the multiple accelerator cards are connected to the multiple memory devices via the management device; The management device is used to: collect actual operation information of each accelerator card; The management device is further configured to: for any target accelerator card among the multiple accelerator cards, predict a memory requirement of the target accelerator card based on actual operation information of the target accelerator card, select at least one memory type, and determine a memory allocation result for the target accelerator card based on the memory requirement and the at least one memory type; The target accelerator card is used to: update the memory configuration according to the memory allocation result; Wherein, the management device includes a management controller and a switching controller; The management controller is connected to the switching controller via a first protocol and a second protocol; The switch controller connects the plurality of accelerator cards and the plurality of memory devices via a third protocol; The management controller connects the host and the multiple acceleration cards via a fourth protocol; each memory device is not perceived by the host and directly serves as an extended memory for each acceleration card, without the host having to allocate memory for each acceleration card.

2. The data processing system according to claim 1, wherein: The management device is used to: The memory usage information, task queue information, and throughput of each accelerator card are collected as the actual operation information of the corresponding accelerator card.

3. The data processing system according to claim 2, wherein: The management device is used to: If the memory usage information exceeds a preset first threshold and the number of unfinished tasks in the task queue information exceeds a preset second threshold, a memory requirement prediction is performed for the target accelerator card.

4. The data processing system according to claim 2, wherein: The management device is used to: Query the historical memory usage information of the target accelerator card; The memory requirement is predicted based on the historical memory usage information, the type of unfinished tasks, and the number of unfinished tasks in the task queue information.

5. The data processing system according to claim 2, wherein: The management device is used to query the unfinished task types and their corresponding allocatable memory types in the task queue information in a preset resource allocation table to obtain the at least one memory type.

6. The data processing system according to claim 1, wherein: The management device is used to: In the preset memory record table, query the remaining available memory and memory type of each memory device to obtain the query result; Determining the memory allocation result according to the query result, the memory requirement and the at least one memory type; Determining identification information of a memory device allocated to the target accelerator card and a corresponding address segment based on the memory allocation result; The memory record table is updated according to the identification information and the address segment.

7. The data processing system according to claim 1, wherein: The management device is used to: Determining identification information of a memory device allocated to the target accelerator card based on the memory allocation result; The mapping relationship between the port of the identification information and the port of the target accelerator card is recorded in a preset routing table.

8. The data processing system according to claim 1, wherein: The target accelerator card is used for: Determining identification information of a memory device allocated to the target accelerator card and a corresponding address segment based on the memory allocation result; The identification information and the address segment are recorded in a preset memory mapping table.

9. The data processing system according to claim 1, wherein: The management device is used to: If the memory usage information in the target accelerator card is lower than a preset third threshold, reclaiming memory resources of a target size from the target accelerator card and updating the memory record table and the routing table accordingly; the target size is an integer multiple of the set memory block size; Correspondingly, the target accelerator card is used to update the memory mapping table according to the memory resources of the target size.

10. The data processing system according to claim 1, wherein: The target accelerator card is used for: If the current available memory amount of the target accelerator card is lower than the memory amount required by the highest priority task when the highest priority task starts to run, the lowest priority task in the target accelerator card is paused, and the current status information of the lowest priority task is stored.

11. A data processing method, characterized in that: Applicable to a management device connected to multiple accelerator cards and multiple memory devices, including: Collect the actual operation information of each accelerator card; For any target accelerator card among the multiple accelerator cards, predict the memory requirement of the target accelerator card according to actual operation information of the target accelerator card, and select at least one memory type; determining a memory allocation result of the target accelerator card based on the memory requirement and the at least one memory type, so that the target accelerator card updates a memory configuration according to the memory allocation result; Wherein, the management device includes a management controller and a switching controller; The management controller is connected to the switching controller via a first protocol and a second protocol; The switch controller connects the plurality of accelerator cards and the plurality of memory devices via a third protocol; The management controller connects the host and the multiple acceleration cards via a fourth protocol; each memory device is not perceived by the host and directly serves as an extended memory for each acceleration card, without the host having to allocate memory for each acceleration card.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to claim 11.

13. A non-volatile storage medium, characterized in that: Used for storing a computer program, wherein the computer program implements the method according to claim 11 when executed by a processor.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method of claim 11 is implemented.

Citation Information

Patent Citations

  • Data processing method, product, server and medium

    CN118409871A