Distributed task memory allocation method, device, medium and product in heterogeneous system

By calculating the minimum memory access bandwidth in a heterogeneous system and allocating remote memory space, the problem of difficult to balance the memory allocation of multiple computing power devices is solved, and efficient memory allocation and optimization of computing performance is achieved.

CN119621355BActive Publication Date: 2025-05-13LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162976.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In heterogeneous systems, when multiple computing power devices participate in the same task, it is difficult for each device to balance the memory characteristics and computing power capabilities of computing power devices, resulting in waste of memory allocation.

Method used

By obtaining the computing power peak information of distributed tasks and each heterogeneous computing power device, the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing power device is calculated, and an appropriate memory device is selected in the heterogeneous system to allocate remote memory space for each computing power device.

Benefits of technology

It realizes balancing memory characteristics and computing power equipment computing capabilities in heterogeneous systems, avoiding waste of memory allocation, and ensuring efficient processing of distributed tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621355B_ABST
    Figure CN119621355B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, medium and product for memory allocation of distributed tasks in a heterogeneous system in the field of computer technology. The present invention implements memory allocation for multiple heterogeneous computing devices participating in the same distributed task without reducing the minimum memory access bandwidth constraint of the computing performance of each heterogeneous computing device, which can not only ensure the computing performance of each heterogeneous computing device when executing the distributed task, but also complete the memory allocation, thereby achieving reasonable memory allocation in the heterogeneous system under the premise of balancing the memory characteristics and the computing power of the computing device, and can make full use of the computing performance of the heterogeneous computing devices to accelerate the processing efficiency of distributed tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a distributed task memory allocation method, device, medium and product in a heterogeneous system. Background Art

[0002] In a heterogeneous system, the memory space of a single computing device is limited, and additional memory devices are often needed to store the data to be processed and the corresponding processing results. In this separated memory scenario, due to the differences in memory performance of different memory devices and the differences in computing power of different computing devices, when multiple computing devices participate in the execution of the same task, the memory allocation of each computing device is difficult to balance the memory characteristics and computing power of the computing device, resulting in waste of memory allocation.

[0003] Therefore, how to balance memory characteristics and computing power of computing devices in heterogeneous systems and achieve reasonable memory allocation is a problem that technical personnel in this field need to solve. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a distributed task memory allocation method, device, medium and product in a heterogeneous system, so as to balance the memory characteristics and the computing power of computing devices in the heterogeneous system and realize reasonable memory allocation.

[0005] In a first aspect, the present invention provides a method for allocating memory for distributed tasks in a heterogeneous system, comprising: obtaining a distributed task; determining a plurality of heterogeneous computing devices that execute the distributed task in the heterogeneous system; calculating the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device based on the computing power peak value of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed; taking each heterogeneous computing device as a device to be allocated, and selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from a plurality of available memory devices included in the heterogeneous system, and allocating off-site memory space for the device to be allocated in the target memory device.

[0006] Optionally, determining multiple heterogeneous computing devices that execute the distributed task in a heterogeneous system includes: obtaining computing power deployment information of the distributed task; and determining the multiple heterogeneous computing devices based on the computing power deployment information.

[0007] Optionally, before calculating the minimum memory access bandwidth that does not affect the computing performance of each heterogeneous computing device based on the computing power peak value of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, the method also includes: obtaining the computing power peak value of each heterogeneous computing device; obtaining the computing power deployment information of the distributed task; and determining the execution complexity of each heterogeneous computing device for the distributed task, the amount of data to be processed, and the amount of parameters to be processed based on the computing power deployment information.

[0008] Optionally, obtaining the peak computing power of each heterogeneous computing power device includes: querying the performance parameters of each heterogeneous computing power device to obtain the peak computing power of each heterogeneous computing power device; or, determining the historical peak computing power of each heterogeneous computing power device; or, using a performance analysis tool to obtain the real-time peak computing power of each heterogeneous computing power device.

[0009] Optionally, based on the peak computing power of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, the minimum memory access bandwidth that does not affect the computing performance of each heterogeneous computing device is calculated separately, including: taking each heterogeneous computing device as a target device, calculating the sum of the amount of data to be processed and the amount of parameters to be processed of the target device for the distributed task, and obtaining the off-site memory allocation amount of the target device; calculating the product of the off-site memory allocation amount and the peak computing power of the target device; and taking the quotient of the execution complexity of the target device for the distributed task and the product as the minimum memory access bandwidth of the target device.

[0010] Optionally, each heterogeneous computing power device is used as a device to be allocated, including: arranging each heterogeneous computing power device in ascending order according to the size of the minimum memory access bandwidth to obtain a first device sequence; starting from the last device in the first device sequence, each heterogeneous computing power device in the first device sequence is used as the device to be allocated in turn; or arranging each heterogeneous computing power device in descending order according to the size of the minimum memory access bandwidth to obtain a second device sequence; starting from the first device in the second device sequence, each heterogeneous computing power device in the first device sequence is used as the device to be allocated in turn.

[0011] Optionally, after allocating off-site memory space to the device to be allocated in the target memory device, the method further includes: if the device to be allocated has completed memory allocation, marking the device to be allocated as an allocated device; and recording the allocated device and the corresponding memory allocation result into a memory allocation list.

[0012] Optionally, after allocating off-site memory space to the device to be allocated in the target memory device, it also includes: if the device to be allocated has completed memory allocation, marking the device to be allocated as an allocated device; recording the allocated device and the corresponding memory allocation result into a memory allocation list, and confirming that the allocated device recorded in the first device sequence or the second device sequence has been deleted.

[0013] Optionally, it also includes: if the number of devices recorded in the memory allocation list is equal to the number of the multiple heterogeneous computing devices, sending the memory allocation list to a preset management end so that the preset management end displays the memory allocation list.

[0014] Optionally, among the multiple available memory devices included in the heterogeneous system, selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated, including: determining, among the multiple available memory devices, at least one to-be-selected memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated; and selecting, among the at least one to-be-selected memory device, the to-be-selected memory device whose memory access bandwidth with the device to be allocated is the smallest as the target memory device.

[0015] Optionally, it also includes: if there is no to-be-selected memory device among the multiple available memory devices whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated, then, among the multiple available memory devices, select the available memory device with the largest memory access bandwidth with the device to be allocated as the target memory device.

[0016] Optionally, allocating off-site memory space to the device to be allocated in the target memory device includes: calculating the sum of the amount of data to be processed and the amount of parameters to be processed by the device to be allocated for the distributed task, and obtaining the off-site memory allocation amount of the device to be allocated; if the available space of the target memory device is not less than the off-site memory allocation amount, dividing the space of the off-site memory allocation amount in the target memory device as the off-site memory space, and confirming that the memory allocation of the device to be allocated has been completed; if the available space of the target memory device is less than the off-site memory allocation amount, allocating the available space of the target memory device to the device to be allocated, and calculating the difference between the off-site memory allocation amount and the available space of the target memory device, taking the difference as the off-site memory allocation amount, and executing the steps of selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from multiple available memory devices included in the heterogeneous system, and allocating off-site memory space to the device to be allocated in the target memory device until it is confirmed that the memory allocation of the device to be allocated has been completed.

[0017] In a second aspect, the present invention provides a distributed task memory allocation device in a heterogeneous system, comprising: an acquisition module for acquiring distributed tasks; a determination module for determining multiple heterogeneous computing devices that execute the distributed tasks in the heterogeneous system; a calculation module for calculating the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device based on the computing power peak of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed; an allocation module for treating each heterogeneous computing device as a device to be allocated, and selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from the multiple available memory devices included in the heterogeneous system, and allocating off-site memory space for the device to be allocated in the target memory device.

[0018] Optionally, the determination module is specifically used to: obtain computing power deployment information of the distributed task; and determine the multiple heterogeneous computing power devices based on the computing power deployment information.

[0019] Optionally, it also includes: an information processing module, which is used to obtain the computing power peak value of each heterogeneous computing power device before calculating the minimum memory access bandwidth that does not affect the computing performance of each heterogeneous computing power device based on the computing power peak value of each heterogeneous computing power device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed; obtain the computing power deployment information of the distributed task; and determine the execution complexity, the amount of data to be processed, and the amount of parameters to be processed of each heterogeneous computing power device for the distributed task based on the computing power deployment information.

[0020] Optionally, the information processing module is used to: query the performance parameters of each heterogeneous computing power device to obtain the peak computing power of each heterogeneous computing power device; or, determine the historical peak computing power of each heterogeneous computing power device; or, use performance analysis tools to obtain the real-time peak computing power of each heterogeneous computing power device.

[0021] Optionally, the calculation module is used to: take each heterogeneous computing device as a target device, calculate the sum of the amount of data to be processed and the amount of parameters to be processed of the target device for the distributed task, and obtain the off-site memory allocation amount of the target device; calculate the product of the off-site memory allocation amount and the peak computing power of the target device; and take the quotient of the execution complexity of the target device for the distributed task and the product as the minimum memory access bandwidth of the target device.

[0022] Optionally, the allocation module is used to: arrange each heterogeneous computing power device in ascending order according to the size of the minimum memory access bandwidth to obtain a first device sequence; starting from the last device in the first device sequence, each heterogeneous computing power device in the first device sequence is used as the device to be allocated in turn; or, arrange each heterogeneous computing power device in descending order according to the size of the minimum memory access bandwidth to obtain a second device sequence; starting from the first device in the second device sequence, each heterogeneous computing power device in the first device sequence is used as the device to be allocated in turn.

[0023] Optionally, it further includes: a first recording module, used for marking the device to be allocated as an allocated device if the memory allocation to the device to be allocated has been completed; and recording the allocated device and the corresponding memory allocation result into a memory allocation list.

[0024] Optionally, it also includes: a second recording module, used to mark the device to be allocated as an allocated device if the memory allocation for the device to be allocated has been completed; record the allocated device and the corresponding memory allocation result into a memory allocation list, and confirm that the allocated device recorded in the first device sequence or the second device sequence has been deleted.

[0025] Optionally, it also includes: a sending module, which is used to send the memory allocation list to a preset management terminal if the number of devices recorded in the memory allocation list is equal to the number of the multiple heterogeneous computing power devices, so that the preset management terminal displays the memory allocation list.

[0026] Optionally, the allocation module is used to: determine, among the multiple available memory devices, at least one candidate memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated; and select, among the at least one candidate memory device, the candidate memory device whose memory access bandwidth with the device to be allocated is the smallest as the target memory device.

[0027] Optionally, the allocation module is used to: if there is no to-be-selected memory device among the multiple available memory devices whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated, then select from the multiple available memory devices an available memory device with the largest memory access bandwidth with the device to be allocated as the target memory device.

[0028] Optionally, the allocation module is used to: calculate the sum of the amount of data to be processed and the amount of parameters to be processed of the device to be allocated for the distributed task, and obtain the off-site memory allocation amount of the device to be allocated; if the available space of the target memory device is not less than the off-site memory allocation amount, divide the space of the off-site memory allocation amount in the target memory device as the off-site memory space, and confirm that the memory allocation of the device to be allocated has been completed; if the available space of the target memory device is less than the off-site memory allocation amount, allocate the available space of the target memory device to the device to be allocated, and calculate the difference between the off-site memory allocation amount and the available space of the target memory device, use the difference as the off-site memory allocation amount, and execute the steps of selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from multiple available memory devices included in the heterogeneous system, and allocating off-site memory space to the device to be allocated in the target memory device until it is confirmed that the memory allocation of the device to be allocated has been completed.

[0029] In a third aspect, the present invention provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the aforementioned method for distributed task memory allocation in a heterogeneous system.

[0030] In a fourth aspect, the present invention provides a non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned method for distributed task memory allocation in a heterogeneous system.

[0031] In a fifth aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned method for distributed task memory allocation in a heterogeneous system.

[0032] It can be seen from the above scheme that the present invention provides a method for memory allocation of distributed tasks in a heterogeneous system, including: obtaining a distributed task; determining multiple heterogeneous computing devices that execute the distributed task in the heterogeneous system; calculating the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device based on the computing power peak value of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed; taking each heterogeneous computing device as a device to be allocated, and selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from the multiple available memory devices included in the heterogeneous system, and allocating remote memory space for the device to be allocated in the target memory device.

[0033] It can be seen that the beneficial effects of the present invention are: for a specific distributed task; multiple heterogeneous computing devices that execute the distributed task are determined in a heterogeneous system; according to the computing power peak of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device is calculated respectively; then each heterogeneous computing device is used as a device to be allocated, and among the multiple available memory devices included in the heterogeneous system, a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated is selected, and a remote memory space is allocated to the device to be allocated in the target memory device. Thus, the scheme realizes memory allocation for multiple heterogeneous computing devices participating in the same distributed task under the constraint of the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device, which can not only ensure the computing performance of each heterogeneous computing device when executing the distributed task, but also complete the memory allocation, thereby realizing reasonable memory allocation under the premise of balancing the memory characteristics and the computing power of the computing device in the heterogeneous system, and can make full use of the computing performance of the heterogeneous computing device, and accelerate the processing efficiency of distributed tasks.

[0034] Correspondingly, a distributed task memory allocation device, medium and product in a heterogeneous system provided by the present invention also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0036] Figure 1 A flow chart of a distributed task memory allocation method in a heterogeneous system disclosed in the present invention;

[0037] Figure 2 A schematic diagram of a heterogeneous computing system disclosed in the present invention;

[0038] Figure 3 A schematic diagram of a memory allocation solution disclosed in the present invention;

[0039] Figure 4 A schematic diagram of a distributed task memory allocation device in a heterogeneous system disclosed in the present invention;

[0040] Figure 5 A schematic diagram of an electronic device disclosed in the present invention;

[0041] Figure 6 A server structure diagram provided by the present invention;

[0042] Figure 7 A terminal structure diagram provided by the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other examples obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] At present, in heterogeneous systems, the memory space of a single computing device is limited, and it is often necessary to use additional memory devices to store the data to be processed and the corresponding processing results. In this separated memory scenario, due to the differences in memory performance of different memory devices and the differences in computing power of different computing devices, when multiple computing devices participate in the execution of the same task, the memory allocation of each computing device is difficult to balance the memory characteristics and the computing power of the computing device, resulting in waste in memory allocation. To this end, the present invention provides a distributed task memory allocation scheme in a heterogeneous system, which can balance the memory characteristics and the computing power of computing devices in a heterogeneous system and achieve reasonable memory allocation.

[0045] See also Figure 1 As shown, an embodiment of the present invention discloses a distributed task memory allocation method in a heterogeneous system, comprising:

[0046] S101. Obtain distributed tasks.

[0047] In this embodiment, the distributed tasks may be distributed model training tasks, distributed model reasoning tasks, and distributed data storage tasks, etc. Accordingly, after obtaining the distributed tasks, the computing power deployment information of the distributed tasks will also be obtained accordingly, and the computing power deployment information records: which heterogeneous computing power devices are used to execute this distributed task, which may specifically include information such as the IP addresses of each heterogeneous computing power device participating in this distributed task; which part of the subtasks in this distributed task are specifically executed by each heterogeneous computing power device; whether the subtasks executed by each heterogeneous computing power device are parallel or serial. From this computing power deployment information, the execution complexity, the amount of data to be processed, and the amount of parameters to be processed of each heterogeneous computing power device can be determined.

[0048] In one embodiment, before respectively calculating the minimum memory access bandwidth that does not affect the computing performance of each heterogeneous computing device based on the computing power peak value of each heterogeneous computing device, the execution complexity of distributed tasks, the amount of data to be processed, and the amount of parameters to be processed, the method further includes: obtaining the computing power peak value of each heterogeneous computing device; obtaining the computing power deployment information of the distributed task; and determining the execution complexity of each heterogeneous computing device for the distributed task, the amount of data to be processed, and the amount of parameters to be processed according to the computing power deployment information. Among them, obtaining the computing power peak value of each heterogeneous computing device includes: querying the performance parameters of each heterogeneous computing device to obtain the computing power peak value of each heterogeneous computing device; or determining the historical computing power peak value of each heterogeneous computing device; or using a performance analysis tool to obtain the real-time computing power peak value of each heterogeneous computing device. Specifically, the user manual of each heterogeneous computing device can be queried to obtain the computing power peak value of each heterogeneous computing device; the historical computing power peak value of each heterogeneous computing device is: the computing power peak value of each heterogeneous computing device during the historical operation process.

[0049] S102. Determine multiple heterogeneous computing devices that execute distributed tasks in a heterogeneous system.

[0050] In one embodiment, determining multiple heterogeneous computing devices for executing distributed tasks in a heterogeneous system includes: obtaining computing power deployment information of the distributed task; and determining multiple heterogeneous computing devices based on the computing power deployment information.

[0051] S103. Calculate the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device based on the computing power peak of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed.

[0052] It should be noted that the execution complexity of distributed tasks, the amount of data to be processed, and the amount of parameters to be processed by heterogeneous computing devices can be determined based on which part of the distributed task the corresponding heterogeneous computing devices recorded in the computing power deployment information will specifically execute.

[0053] This embodiment does not consider the local memory of the heterogeneous computing power device, that is: the local memory of each heterogeneous computing power device is not used by default. Therefore, in one embodiment, according to the peak computing power of each heterogeneous computing power device, the execution complexity of distributed tasks, the amount of data to be processed, and the amount of parameters to be processed, the minimum memory access bandwidth that does not affect the computing performance of each heterogeneous computing power device is calculated separately, including: taking each heterogeneous computing power device as a target device, calculating the sum of the amount of data to be processed and the amount of parameters to be processed of the target device for distributed tasks, and obtaining the off-site memory allocation of the target device; calculating the product of the off-site memory allocation amount and the peak computing power of the target device; taking the quotient of the execution complexity of the target device for distributed tasks and the product as the minimum memory access bandwidth of the target device.

[0054] S103. Treat each heterogeneous computing device as a device to be allocated, and select a target memory device whose memory access bandwidth to the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from multiple available memory devices included in the heterogeneous system, and allocate off-site memory space for the device to be allocated in the target memory device.

[0055] In one embodiment, each heterogeneous computing power device is used as a device to be allocated, including: arranging each heterogeneous computing power device in ascending order according to the size of the minimum memory access bandwidth to obtain a first device sequence; starting from the last device in the first device sequence, each heterogeneous computing power device in the first device sequence is used as a device to be allocated in turn; or, arranging each heterogeneous computing power device in descending order according to the size of the minimum memory access bandwidth to obtain a second device sequence; starting from the first device in the second device sequence, each heterogeneous computing power device in the first device sequence is used as a device to be allocated in turn.

[0056] In order to facilitate the collection of memory allocation results of each heterogeneous computing power device, this embodiment records in the memory allocation list: the allocated devices that have completed memory allocation and the corresponding memory allocation results, and the memory allocation results include: the device identification information (such as IP address, device ID, etc.) of the target memory device and the allocated memory address segment. In one embodiment, after allocating off-site memory space to the device to be allocated in the target memory device, it also includes: if the device to be allocated has completed memory allocation, the device to be allocated is marked as an allocated device; the allocated device and the corresponding memory allocation result are recorded in the memory allocation list. In one embodiment, after allocating off-site memory space to the device to be allocated in the target memory device, it also includes: if the device to be allocated has completed memory allocation, the device to be allocated is marked as an allocated device; the allocated device and the corresponding memory allocation result are recorded in the memory allocation list, and it is confirmed that the allocated device recorded in the first device sequence or the second device sequence has been deleted. Correspondingly, if the number of devices recorded in the memory allocation list is equal to the number of multiple heterogeneous computing power devices, the memory allocation list is sent to the preset management end so that the preset management end displays the memory allocation list.

[0057] An available memory device is a memory device whose available space in a heterogeneous system is not zero. In one embodiment, among multiple available memory devices included in the heterogeneous system, a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated is selected, including: among multiple available memory devices, determining at least one memory device to be selected whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated; among at least one memory device to be selected, selecting the memory device to be selected with the smallest memory access bandwidth with the device to be allocated as the target memory device. If there is no memory device to be selected whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated among the multiple available memory devices, then among the multiple available memory devices, an available memory device with the largest memory access bandwidth with the device to be allocated is selected as the target memory device.

[0058] In one embodiment, allocating off-site memory space to a device to be allocated in a target memory device includes: calculating the sum of the amount of data to be processed and the amount of parameters to be processed of the device to be allocated for a distributed task, and obtaining the off-site memory allocation amount of the device to be allocated; if the available space of the target memory device is not less than the off-site memory allocation amount, dividing the space of the off-site memory allocation amount in the target memory device as the off-site memory space, and confirming that the memory allocation of the device to be allocated has been completed; if the available space of the target memory device is less than the off-site memory allocation amount, allocating the available space of the target memory device to the device to be allocated, and calculating the difference between the off-site memory allocation amount and the available space of the target memory device, taking the difference as the off-site memory allocation amount, and executing the steps of selecting a target memory device whose memory access bandwidth to the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from multiple available memory devices included in the heterogeneous system, and allocating off-site memory space to the device to be allocated in the target memory device, until confirming that the memory allocation of the device to be allocated has been completed.

[0059] It can be seen that this embodiment is aimed at a specific distributed task; multiple heterogeneous computing devices that execute the distributed task are determined in a heterogeneous system; according to the computing power peak of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device is calculated respectively; then each heterogeneous computing device is used as a device to be allocated, and among the multiple available memory devices included in the heterogeneous system, a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated is selected, and a remote memory space is allocated to the device to be allocated in the target memory device. Therefore, the scheme realizes memory allocation for multiple heterogeneous computing devices participating in the same distributed task under the constraint of the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device, which can not only ensure the computing performance of each heterogeneous computing device when executing the distributed task, but also complete the memory allocation, thereby realizing reasonable memory allocation in the heterogeneous system under the premise of balancing the memory characteristics and the computing power of the computing device, and can make full use of the computing performance of the heterogeneous computing device and accelerate the processing efficiency of the distributed task. Heterogeneous computing devices such as GPU, FPGA, etc.

[0060] See also Figure 2 In a heterogeneous system, multiple heterogeneous computing chips (i.e., heterogeneous computing devices) are included. Each heterogeneous computing chip is equipped with a chip local memory, but the chip local memory space is limited. Therefore, the heterogeneous system also includes multiple non-local memory devices. In this scenario, the local memory of the heterogeneous computing chip and the remote non-local memory will be connected to a separate memory interconnection network. All memories in the interconnection network can be directly accessed as the memory of any heterogeneous computing chip. In this case, the memory allocation plan for the corresponding task can be formulated without considering the local memory capacity of the computing chip and the access delay of the remote memory.

[0061] However, since heterogeneous computing chips access non-local memory devices, memory access performance will be reduced. Therefore, this embodiment will comprehensively consider the memory access performance and the computing performance of the computing power device to formulate a memory allocation plan. Figure 3 The memory allocation solution provided in this embodiment includes: a heterogeneous computing system information collection module, a separate memory deployment solution calculation module, and a separate memory deployment delivery module.

[0062] The heterogeneous computing system information collection module is used to collect information about distributed AI training tasks (i.e., a distributed task) after the distributed AI training tasks (i.e., a distributed task) are issued, and at the same time, collect information about the separated memory interconnection network; the collected information will be sent to the separated memory deployment solution calculation module, and the collected information will be used as input data for the separated memory deployment solution calculation module.

[0063] Specifically, the information of the distributed AI training task includes: the computing power deployment of the distributed AI training task, the information of the training subtasks to be performed by each computing power device, and the training data information of each training subtask. Through the computing power deployment of the distributed AI training task, it can be clarified: on which heterogeneous computing power devices the distributed AI training task will be performed. Generally, it is possible to specify which heterogeneous computing power devices in the heterogeneous computing system are used to perform the distributed AI training task through code and other means, so as to clarify which heterogeneous computing power device specifically performs which training subtask. Among them, the information of each training subtask includes: the computational complexity FLOPs and parameter quantity of each training subtask. These data can be estimated based on the neural network structure involved in the training subtask, and the estimation process can use mathematical methods. For example: based on the structure of a certain neural layer (such as a fully connected layer in a neural network), the parameter quantity of the corresponding neural layer is determined. Exemplarily, for a fully connected layer with a weight matrix of [x, x], its parameters include a weight matrix and a bias, and the data volume of the fully connected layer parameters is (x×x+x)×P×8 bytes; P represents the training precision, for example, P=1 for single precision and P=2 for double precision. Of course, you can also use some open source tools to calculate the complexity of FLOPs and the number of parameters.

[0064] The training data information of each training subtask is specifically: the amount of input data of each training subtask. If the input of a training subtask is training data, then the amount of input training data is the amount of input data of the subtask. For example, in data parallel training mode, the input of each training subtask on a heterogeneous computing device is a split training data set. For example, in pipeline training mode, the input of a training subtask is the output of other training subtasks. If the input of a training subtask is some intermediate activation value data, then the amount of input data can be estimated by multiplying the batch-size by the tensor size of the training subtask input. For example: if the batch-size is 10 and the input size defined by the training subtask is [10,10], then the amount of input data is 10×10×10×training accuracy.

[0065] The separated memory interconnection network information includes: the peak computing power of the heterogeneous computing power device, the real-time bandwidth between the heterogeneous computing power device and each memory device, and the real-time remaining memory space of each memory device. Among them, the peak computing power information of the heterogeneous computing power device can be obtained through the performance manual of the heterogeneous computing power device or related public information, and entered into the heterogeneous computing system information collection module in advance for storage, so that it can be retrieved when needed. The peak computing power information can also be obtained by recording the execution information of the historical heterogeneous computing power device, that is, using the profiler tool that comes with the deep learning framework to obtain the device peak computing power when the device is executed. Among them, the real-time bandwidth between the heterogeneous computing power device and each memory device can be obtained using public testing tools such as sysbench.

[0066] After the above information is collected, it will be sent to the separated memory deployment solution calculation module to calculate the optimal memory deployment solution.

[0067] The separate memory deployment solution calculation module can calculate the separate memory deployment solution based on the information collected by the heterogeneous computing system information collection module, and send the results of the deployment solution to the separate memory deployment delivery module. Specifically, the separate memory deployment solution calculation module will output the corresponding memory allocation location and memory allocation space size for each heterogeneous computing device that needs to deploy distributed training subtasks. A heterogeneous computing device may correspond to multiple memory allocation locations. The result of this deployment location will be sent to the separate memory deployment delivery module for actual deployment.

[0068] Separate memory deployment distribution module: deploys the received separate memory deployment plan into the heterogeneous computing system, reasonably expands the memory for the heterogeneous computing devices, and avoids excessive performance impact on the calculation. According to the calculation results of the separate memory deployment plan calculation module, this module will open up memory space on the separate memory device of the heterogeneous computing platform according to the calculated deployment plan, and allocate it to the corresponding heterogeneous computing devices to complete this memory deployment. Subsequently, the distributed AI training task can start the actual execution, so as to execute this distributed AI training task in the heterogeneous computing system with better execution efficiency.

[0069] For example, for each heterogeneous computing device , output memory allocation results: , where D represents the memory device identifier allocated to the XPU, and M represents the size of the memory space that needs to be allocated to the XPU on D. The above memory allocation result indicates that there are m memory devices and n XPUs Participate in this distributed training task.

[0070] Specifically, the definitions of the parameters involved in the calculation are:

[0071] The heterogeneous computing devices involved in distributed training are: .

[0072] The peak computing power corresponding to each heterogeneous computing device is: .

[0073] The computational complexity of the training subtasks to which each heterogeneous computing power is allocated is .

[0074] The amount of input data allocated to each training subtask of each heterogeneous computing device is , the corresponding model parameter is .

[0075] Memory devices in a separate memory interconnect network The real-time available capacity is: .

[0076] Among them, heterogeneous computing power is defined To memory device The real-time bandwidth is .

[0077] Based on the above information, the specific calculation process is as follows:

[0078] Step 1: For each heterogeneous computing power participating in the distributed training task , its memory access bandwidth performance (i.e., minimum memory access bandwidth) requirement is defined as: . Based on this, this embodiment sets: the memory bandwidth must be at least greater than or equal to , can ensure that the computing performance of heterogeneous computing power will not be affected.

[0079] Step 2: Corresponding Sort from large to small; that is, arrange the heterogeneous computing devices in descending order according to the size of the minimum memory access bandwidth to obtain the second device sequence.

[0080] Step 3: Select the unassigned device from the second device sequence The largest computing power equipment Allocate memory for it. The goal is to prioritize memory allocation for computing devices with large and difficult-to-satisfy computing bandwidth performance requirements. When allocating memory for the first time, the memory allocation result is recorded as an empty list. , the memory allocation requirement of the current computing power device is If it does not exist, it has not been allocated yet. , then go to step 7.

[0081] Step 4: Set the current computing power device To memory device The bandwidth is recorded as: , , ..., , and sort from largest to smallest.

[0082] Step 5: Select the one that satisfies Minimum And real-time available capacity Memory device that is not 0 , as a The target memory device for memory allocation. This choice can leave the memory device with better memory performance for other training subtasks without affecting computing performance.

[0083] If not present , then directly select The largest As a The target memory device for memory allocation. That is, if there is no target device to be allocated If there is a candidate memory device whose memory access bandwidth between the candidate memory device and the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated, then the available memory device with the largest memory access bandwidth between the candidate memory device and the device to be allocated is selected as the target memory device.

[0084] Step 6: Allocate remote memory space for the device to be allocated in the remote memory device, including: Real-time available capacity ,if , then directly Allocate memory for Z, denoted as , and added to , then update Real-time available capacity , and return to step 3. If , then directly distribute The memory is recorded as , and added to , then update Real-time available capacity is 0, and at the same time , and return to step 5.

[0085] Step 7: List the final memory allocation As the final memory allocation result, output is performed.

[0086] It can be seen that this embodiment allocates the memory deployment location and memory space for each relevant heterogeneous computing power according to the information collected by the heterogeneous computing system information collection module for the currently issued distributed AI training task, so as to fully utilize the XPU computing performance; it also fully utilizes the computing performance of heterogeneous computing power, accelerates the efficiency of distributed training, avoids the computing performance impact caused by unreasonable memory expansion, and ensures the computing efficiency of heterogeneous computing power devices with better performance.

[0087] A distributed task memory allocation device in a heterogeneous system provided by an embodiment of the present invention is introduced below. The distributed task memory allocation device in a heterogeneous system described below can be referenced to other embodiments described in this document.

[0088] See also Figure 4As shown, an embodiment of the present invention discloses a distributed task memory allocation device in a heterogeneous system, including: an acquisition module, used to acquire distributed tasks; a determination module, used to determine multiple heterogeneous computing devices that execute distributed tasks in the heterogeneous system; a calculation module, used to calculate the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device according to the computing power peak of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed; an allocation module, used to take each heterogeneous computing device as a device to be allocated, and select a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from the multiple available memory devices included in the heterogeneous system, and allocate remote memory space to the device to be allocated in the target memory device.

[0089] In one embodiment, the determination module is specifically used to: obtain computing power deployment information of a distributed task; and determine multiple heterogeneous computing power devices based on the computing power deployment information.

[0090] In one embodiment, it also includes: an information processing module, which is used to obtain the computing power peak value of each heterogeneous computing power device before calculating the minimum memory access bandwidth that does not affect the computing performance of each heterogeneous computing power device based on the computing power peak value of each heterogeneous computing power device, the execution complexity of distributed tasks, the amount of data to be processed, and the amount of parameters to be processed; obtain the computing power deployment information of the distributed tasks; and determine the execution complexity of each heterogeneous computing power device for distributed tasks, the amount of data to be processed, and the amount of parameters to be processed based on the computing power deployment information.

[0091] In one embodiment, the information processing module is used to: query the performance parameters of each heterogeneous computing power device to obtain the peak computing power of each heterogeneous computing power device; or, determine the historical peak computing power of each heterogeneous computing power device; or, use a performance analysis tool to obtain the real-time peak computing power of each heterogeneous computing power device.

[0092] In one embodiment, the calculation module is used to: take each heterogeneous computing device as a target device, calculate the sum of the amount of data to be processed and the amount of parameters to be processed of the target device for distributed tasks, and obtain the off-site memory allocation of the target device; calculate the product of the off-site memory allocation and the peak computing power of the target device; and use the quotient of the execution complexity of the target device for distributed tasks and the product as the minimum memory access bandwidth of the target device.

[0093] In one embodiment, the allocation module is used to: arrange each heterogeneous computing power device in ascending order according to the size of the minimum memory access bandwidth to obtain a first device sequence; starting from the last device in the first device sequence, each heterogeneous computing power device in the first device sequence is used as a device to be allocated in turn; or, arrange each heterogeneous computing power device in descending order according to the size of the minimum memory access bandwidth to obtain a second device sequence; starting from the first device in the second device sequence, each heterogeneous computing power device in the first device sequence is used as a device to be allocated in turn.

[0094] In one embodiment, the method further includes: a first recording module, configured to mark the device to be allocated as an allocated device if the device to be allocated has completed memory allocation; and record the allocated device and the corresponding memory allocation result into a memory allocation list.

[0095] In one embodiment, it also includes: a second recording module, which is used to mark the device to be allocated as an allocated device if the memory allocation has been completed for the device to be allocated; record the allocated device and the corresponding memory allocation result into the memory allocation list, and confirm that the allocated device recorded in the first device sequence or the second device sequence has been deleted.

[0096] In one embodiment, it also includes: a sending module, which is used to send the memory allocation list to the preset management end if the number of devices recorded in the memory allocation list is equal to the number of multiple heterogeneous computing power devices, so that the preset management end displays the memory allocation list.

[0097] In one embodiment, the allocation module is used to: determine, among multiple available memory devices, at least one candidate memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated; and select, among at least one candidate memory device, the candidate memory device whose memory access bandwidth with the device to be allocated is the smallest as the target memory device.

[0098] In one embodiment, the allocation module is used to: if there is no to-be-selected memory device among multiple available memory devices whose memory access bandwidth with the to-be-allocated device is not less than the minimum memory access bandwidth of the to-be-allocated device, then among the multiple available memory devices, select the available memory device with the largest memory access bandwidth with the to-be-allocated device as the target memory device.

[0099] In one embodiment, the allocation module is used to: calculate the sum of the amount of data to be processed and the amount of parameters to be processed for the distributed task of the device to be allocated, and obtain the off-site memory allocation amount of the device to be allocated; if the available space of the target memory device is not less than the off-site memory allocation amount, divide the space of the off-site memory allocation amount in the target memory device as the off-site memory space, and confirm that the memory allocation of the device to be allocated has been completed; if the available space of the target memory device is less than the off-site memory allocation amount, allocate the available space of the target memory device to the device to be allocated, and calculate the difference between the off-site memory allocation amount and the available space of the target memory device, use the difference as the off-site memory allocation amount, and execute the steps of selecting a target memory device whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from multiple available memory devices included in the heterogeneous system, and allocating off-site memory space to the device to be allocated in the target memory device until it is confirmed that the memory allocation of the device to be allocated has been completed.

[0100] Among them, for more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0101] It can be seen that this embodiment can balance the memory characteristics and the computing power of computing devices in a heterogeneous system and achieve reasonable memory allocation.

[0102] An electronic device provided by an embodiment of the present invention is introduced below. The electronic device described below can be referenced to other embodiments described in this document.

[0103] See also Figure 5 As shown, an embodiment of the present invention discloses an electronic device, including:

[0104] Memory 501, used for storing computer programs;

[0105] The processor 502 is used to execute the computer program to implement the method disclosed in any of the above embodiments.

[0106] Furthermore, an embodiment of the present invention also provides an electronic device. The electronic device can be Figure 6 The server shown can also be Figure 7 Terminal shown. Figure 6 and Figure 7 All of them are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the diagrams cannot be regarded as any limitation on the scope of application of the present invention.

[0107] Figure 6A schematic diagram of the structure of a server provided in an embodiment of the present invention. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the relevant steps in the distributed task memory allocation in the heterogeneous system disclosed in any of the aforementioned embodiments.

[0108] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present invention, and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0109] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.

[0110] The operating system is used to manage and control the hardware devices and computer programs on the server to realize the operation and processing of the data in the memory by the processor, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the distributed task memory allocation method in the heterogeneous system disclosed in any of the aforementioned embodiments, the computer program can further include computer programs that can be used to complete other specific tasks. In addition to data such as application update information, data can also include data such as application developer information.

[0111] Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of the present invention, wherein the terminal may specifically include but is not limited to a smart phone, a tablet computer, a laptop computer or a desktop computer.

[0112] Generally, the terminal in this embodiment includes: a processor and a memory.

[0113] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0114] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is used to store at least the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the distributed task memory allocation method in the heterogeneous system executed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be short-term storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, update information of the application.

[0115] In some embodiments, the terminal may also include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0116] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than those shown in the figure.

[0117] A nonvolatile storage medium provided in an embodiment of the present invention is introduced below. The nonvolatile storage medium described below and other embodiments described herein may be cross-referenced.

[0118] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the method for allocating memory for distributed tasks in a heterogeneous system disclosed in the above-mentioned embodiment. The non-volatile storage medium is a computer-readable non-volatile storage medium, which, as a carrier for storing resources, may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system, a computer program, and data, etc., and the storage method may be temporary storage or permanent storage.

[0119] A computer program product provided by an embodiment of the present invention is introduced below. The computer program product described below can be referenced to other embodiments described in this document.

[0120] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned method for allocating memory for distributed tasks in a heterogeneous system.

[0121] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0122] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.

[0123] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A distributed task memory allocation method in a heterogeneous system, characterized in that: include: Get distributed tasks; Determining multiple heterogeneous computing devices in a heterogeneous system to execute the distributed task; According to the peak computing power of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, respectively calculate the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device; Taking each heterogeneous computing device as a device to be allocated, selecting a target memory device whose memory access bandwidth to the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated from multiple available memory devices included in the heterogeneous system, and allocating a remote memory space for the device to be allocated in the target memory device; Among them, according to the peak computing power of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device is calculated respectively, including: Taking each heterogeneous computing device as a target device, calculating the sum of the amount of data to be processed and the amount of parameters to be processed by the target device for the distributed task, and obtaining the amount of off-site memory allocated to the target device; Calculate the product of the remote memory allocation amount and the peak computing power of the target device; The quotient of the execution complexity of the target device for the distributed task and the product is used as the minimum memory access bandwidth of the target device.

2. The method according to claim 1, characterized in that Determining multiple heterogeneous computing devices for executing the distributed task in a heterogeneous system includes: Obtaining computing power deployment information of the distributed task; The multiple heterogeneous computing devices are determined according to the computing power deployment information.

3. The method according to claim 1, characterized in that: Before respectively calculating the minimum memory access bandwidth that does not reduce the computing performance of each heterogeneous computing device according to the computing power peak of each heterogeneous computing device, the execution complexity of the distributed task, the amount of data to be processed, and the amount of parameters to be processed, the method further includes: Obtain the peak computing power of each heterogeneous computing device; Obtaining computing power deployment information of the distributed task; The execution complexity, amount of data to be processed and amount of parameters to be processed of each heterogeneous computing device for the distributed task are determined according to the computing power deployment information.

4. The method according to claim 3, characterized in that: Obtain the peak computing power of each heterogeneous computing device, including: Query the performance parameters of each heterogeneous computing device to obtain the peak computing power of each heterogeneous computing device; or, determine the historical peak computing power of each heterogeneous computing device; or, use performance analysis tools to obtain the real-time peak computing power of each heterogeneous computing device.

5. The method according to claim 1, characterized in that Each heterogeneous computing device is used as a device to be allocated, including: Arrange the heterogeneous computing devices in ascending order according to the size of the minimum memory access bandwidth to obtain a first device sequence; starting from the last device in the first device sequence, use the heterogeneous computing devices in the first device sequence as the devices to be allocated in turn; or, arrange the heterogeneous computing devices in descending order according to the size of the minimum memory access bandwidth to obtain a second device sequence; starting from the first device in the second device sequence, use the heterogeneous computing devices in the first device sequence as the devices to be allocated in turn.

6. The method according to claim 1, characterized in that After allocating the off-site memory space for the device to be allocated in the target memory device, the method further includes: If the device to be allocated has completed memory allocation, marking the device to be allocated as an allocated device; The allocated devices and corresponding memory allocation results are recorded in a memory allocation list.

7. The method according to claim 5, characterized in that After allocating the off-site memory space for the device to be allocated in the target memory device, the method further includes: If the device to be allocated has completed memory allocation, marking the device to be allocated as an allocated device; The allocated device and the corresponding memory allocation result are recorded in a memory allocation list, and it is confirmed that the allocated device recorded in the first device sequence or the second device sequence has been deleted.

8. The method according to claim 6 or 7, characterized in that: Also includes: If the number of devices recorded in the memory allocation list is equal to the number of the multiple heterogeneous computing devices, the memory allocation list is sent to the preset management end so that the preset management end displays the memory allocation list.

9. The method according to any one of claims 1 to 7, characterized in that: Selecting a target memory device whose memory access bandwidth to the device to be allocated is not less than a minimum memory access bandwidth of the device to be allocated from a plurality of available memory devices included in the heterogeneous system comprises: Determine, among the multiple available memory devices, at least one memory device to be selected whose memory access bandwidth to the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated; Among the at least one memory device to be selected, a memory device having the smallest memory access bandwidth to the device to be allocated is selected as the target memory device.

10. The method according to claim 9, characterized in that Also includes: If there is no memory device to be selected among the multiple available memory devices whose memory access bandwidth with the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated, then among the multiple available memory devices, an available memory device with the largest memory access bandwidth with the device to be allocated is selected as the target memory device.

11. The method according to any one of claims 1 to 7, characterized in that: Allocating off-site memory space for the device to be allocated in the target memory device includes: Calculating the sum of the amount of data to be processed and the amount of parameters to be processed of the to-be-allocated device for the distributed task, and obtaining the off-site memory allocation amount of the to-be-allocated device; If the available space of the target memory device is not less than the remote memory allocation amount, a space of the size of the remote memory allocation amount is divided in the target memory device as the remote memory space, and it is confirmed that the memory allocation has been completed for the device to be allocated; If the available space of the target memory device is less than the off-site memory allocation amount, the available space of the target memory device is allocated to the device to be allocated, and the difference between the off-site memory allocation amount and the available space of the target memory device is calculated, and the difference is used as the off-site memory allocation amount. Then, from the multiple available memory devices included in the heterogeneous system, a target memory device whose memory access bandwidth to the device to be allocated is not less than the minimum memory access bandwidth of the device to be allocated is selected, and off-site memory space is allocated to the device to be allocated in the target memory device until it is confirmed that the memory allocation of the device to be allocated has been completed.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 11.

13. A non-volatile storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Streaming computing system for distributed training, method and device thereof

    CN114048039A

  • Method and device for processing key value pair storage system

    CN118227337A