Memory allocation method, device, medium and product for models in heterogeneous systems

By decomposing the target task into multiple processing stages in a heterogeneous system and matching the memory allocation strategy based on the processing time and the amount of memory access data, the problem of unreasonable memory allocation in a heterogeneous system is solved and the task processing efficiency is improved.

CN119621356BActive Publication Date: 2025-05-06SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510163019.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-06
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In heterogeneous systems, due to the performance differences between different memory devices and the computing power differences in computing power devices, unreasonable memory allocation may affect task efficiency.

Method used

By obtaining the target task of the target model and decomposing it into multiple processing stages, the processing time and the amount of data accessed for each processing stage are obtained. Then, based on the processing time and the amount of memory access data match with the local memory access performance of the computing power device, it is determined whether to allocate memory space in local memory or remote memory.

Benefits of technology

Reasonable memory allocation in heterogeneous systems is realized, task processing efficiency is improved, and performance degradation caused by unreasonable allocation is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621356B_ABST
    Figure CN119621356B_ABST
Patent Text Reader

Abstract

The present invention discloses a memory allocation method, device, medium and product for a model in a heterogeneous system in the field of computer technology. The present invention innovatively regards the corresponding processing process of the target task executed by each network layer in the model as each processing stage, and then performs memory allocation for each processing stage; and under the premise of considering the total local memory access latency of the target computing power device for the specific task execution process, the memory access data volume of the specific task execution process and the local memory space of the target computing power device, reasonable memory allocation is achieved for each processing stage of the task, ensuring that the memory allocation does not affect the efficiency and performance of the currently executed task as much as possible, thereby accelerating the task processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a memory allocation method, device, medium and product for a model in a heterogeneous system. Background Art

[0002] In a heterogeneous system, the memory space of a single computing device is limited, and additional memory devices are often needed to store the data to be processed and the corresponding processing results. In this separated memory scenario, due to the differences in memory performance of different memory devices and the differences in computing power of different computing devices, unreasonable memory allocation may affect the efficiency of tasks executed on the computing devices.

[0003] Therefore, how to achieve reasonable memory allocation in a heterogeneous system is a problem that needs to be solved by those skilled in the art. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a memory allocation method, device, medium and product for models in a heterogeneous system, so as to achieve reasonable memory allocation in the heterogeneous system. The specific scheme is as follows:

[0005] In a first aspect, the present invention provides a memory allocation method for a model in a heterogeneous system, comprising:

[0006] Obtaining a target task of a target model; the target task includes: a model single iteration task or a model single reasoning task;

[0007] Utilize the target computing device in the heterogeneous system to run the target task, and sequentially determine the corresponding processing process of the target task executed by each network layer in the target model into multiple processing stages;

[0008] Get the processing time and amount of data accessed for each processing stage.

[0009] Each processing stage is taken as the target object respectively. If the processing time of the target object is less than the total local memory access latency of the target computing power device for the target object, and the amount of memory access data of the target object is not larger than the local memory space of the target computing power device, then memory space of the same size as the amount of memory access data is allocated to the target object in the target computing power device.

[0010] Optionally, the corresponding processing process of the target task performed by each network layer in the target model is sequentially determined into multiple processing stages, including:

[0011] Determining each network layer included in the target model;

[0012] According to the execution order of the corresponding processing procedures of the target task executed by each network layer in the target model, the corresponding processing procedures of the target task executed by each network layer are determined as each processing stage;

[0013] Among them, the same network layer corresponds to at least one processing stage.

[0014] Optionally, before obtaining the processing time and the amount of accessed data corresponding to each processing stage, the following is further included:

[0015] Arranging the processing stages into a target sequence according to the execution order of the corresponding processing processes of the target task performed by each network layer in the target model;

[0016] Accordingly, each processing stage is taken as a target object, including:

[0017] Starting from the first position of the target sequence, each processing stage in the target sequence is taken as the target object in turn.

[0018] Optionally, determining the amount of memory access data corresponding to each processing stage includes:

[0019] The amount of memory access data corresponding to each processing stage is calculated respectively according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage.

[0020] Optionally, according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage, the amount of memory access data corresponding to each processing stage is calculated respectively, including:

[0021] If any processing stage belongs to the forward propagation of the model, the sum of the amount of data to be processed and the amount of parameters to be processed in the current processing stage is determined as the amount of memory access data corresponding to the current processing stage;

[0022] If any processing stage belongs to model back propagation, the amount of memory access data corresponding to the current processing stage is calculated based on the amount of parameters to be processed in the current processing stage.

[0023] Optionally, the amount of memory access data corresponding to the current processing stage is calculated according to the amount of parameters to be processed in the current processing stage, including:

[0024] The sum of the data amounts of the calculation parameters, gradients, optimizer states, and activation values ​​in the parameters to be processed in the current processing stage is determined as the memory access data amount corresponding to the current processing stage.

[0025] Optionally, determining the processing time corresponding to each processing stage includes:

[0026] Obtaining the peak computing power of the target computing power device;

[0027] The ratio of the computational complexity of any processing stage to the computing power peak is taken as the processing duration corresponding to the current processing stage.

[0028] Optionally, obtaining a local memory access bandwidth and a local memory access latency of the target computing power device;

[0029] Calculating a ratio of the amount of memory access data of the target object to the local memory access bandwidth;

[0030] The sum of the ratio and the local memory access delay is taken as the total local memory access delay.

[0031] Optionally, it also includes:

[0032] If the processing time of the target object is not less than the total local memory access latency of the target computing power device for the target object, a memory space of the size of the memory access data is allocated to the target object in the target memory device and the target computing power device in the heterogeneous system.

[0033] Optionally, in the target memory device and the target computing power device in the heterogeneous system, allocating a memory space of the size of the access data amount for the target object includes:

[0034] Calculate a first memory space allocated for the target object in the target computing device;

[0035] Calculating a second memory space allocated for the target object in the target memory device;

[0036] The sum of the first memory space and the second memory space is equal to the memory space of the access data amount.

[0037] Optionally, calculating a first memory space allocated for the target object in the target computing power device includes:

[0038] The first memory space is obtained by calculating according to a first formula; the first formula is:

[0039] ;

[0040] in, is the first memory space, t is the processing time of the target object Ci, l local is the local memory access latency of the target computing device, l remote is the remote memory access latency of the target computing device to access the target memory device, B local is the local memory access bandwidth of the target computing device, B remote is the remote memory access bandwidth of the target computing device to the target memory device, and N is the amount of memory access data corresponding to the target object;

[0041] Accordingly, calculating the second memory space allocated for the target object in the target memory device includes:

[0042] The difference between the amount of accessed data corresponding to the target object and the first memory space is used as the second memory space.

[0043] Optionally, it also includes:

[0044] If the local memory space is not smaller than the first memory space, determining the first memory space and the second memory space as the allocation result of the target object, and updating the local memory space to the difference between the local memory space and the first memory space;

[0045] If the local memory space is smaller than the first memory space, the first memory space is changed to the local memory space; the second memory space is changed to the difference between the amount of access data corresponding to the target object and the local memory space; the changed first memory space and the changed second memory space are determined as the allocation results of the target object, and the local memory space is marked as empty.

[0046] In a second aspect, the present invention provides a memory allocation device for a model in a heterogeneous system, comprising:

[0047] An acquisition module is used to acquire a target task of a target model; the target task includes: a single iteration task of a model or a single reasoning task of a model;

[0048] A determination module, used for running the target task by using the target computing device in the heterogeneous system, and determining the corresponding processing process of the target task executed by each network layer in the target model into multiple processing stages in sequence;

[0049] A calculation module is used to obtain the processing time and amount of memory access data corresponding to each processing stage;

[0050] An allocation module is used to take each processing stage as a target object respectively. If the processing time of the target object is less than the total local memory access delay of the target computing power device for the target object, and the amount of memory access data of the target object is not larger than the local memory space of the target computing power device, then a memory space of the size of the memory access data is allocated to the target object in the target computing power device.

[0051] Optionally, the determination module is specifically used for:

[0052] Determining each network layer included in the target model;

[0053] According to the execution order of the corresponding processing procedures of the target task executed by each network layer in the target model, the corresponding processing procedures of the target task executed by each network layer are determined as each processing stage;

[0054] Among them, the same network layer corresponds to at least one processing stage.

[0055] Optionally, it also includes:

[0056] A sorting module is used to arrange the processing stages into a target sequence according to the execution order of the corresponding processing processes of the target tasks executed by each network layer in the target model before obtaining the processing time and the amount of memory access data corresponding to each processing stage;

[0057] Accordingly, the allocation module is specifically used for:

[0058] Starting from the first position of the target sequence, each processing stage in the target sequence is taken as the target object in turn.

[0059] Optionally, the computing module is specifically used for:

[0060] The amount of memory access data corresponding to each processing stage is calculated respectively according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage.

[0061] Optionally, the computing module is specifically used for:

[0062] If any processing stage belongs to the forward propagation of the model, the sum of the amount of data to be processed and the amount of parameters to be processed in the current processing stage is determined as the amount of memory access data corresponding to the current processing stage;

[0063] If any processing stage belongs to model back propagation, the amount of memory access data corresponding to the current processing stage is calculated based on the amount of parameters to be processed in the current processing stage.

[0064] Optionally, the computing module is specifically used for:

[0065] The sum of the data amounts of the calculation parameters, gradients, optimizer states, and activation values ​​in the parameters to be processed in the current processing stage is determined as the memory access data amount corresponding to the current processing stage.

[0066] Optionally, the computing module is specifically used for:

[0067] Obtaining the peak computing power of the target computing power device;

[0068] The ratio of the computational complexity of any processing stage to the computing power peak is taken as the processing duration corresponding to the current processing stage.

[0069] Optionally, it also includes:

[0070] The total delay calculation module is used to obtain the local memory access bandwidth and local memory access delay of the target computing power device; calculate the ratio of the memory access data volume of the target object to the local memory access bandwidth; and sum the ratio and the local memory access delay as the total local memory access delay.

[0071] Optionally, the allocation module is further configured to:

[0072] If the processing time of the target object is not less than the total local memory access latency of the target computing power device for the target object, a memory space of the size of the memory access data is allocated to the target object in the target memory device and the target computing power device in the heterogeneous system.

[0073] Optionally, the allocation module is further configured to:

[0074] Calculate a first memory space allocated for the target object in the target computing device;

[0075] Calculating a second memory space allocated for the target object in the target memory device;

[0076] The sum of the first memory space and the second memory space is equal to the memory space of the access data amount.

[0077] Optionally, the allocation module is further configured to:

[0078] The first memory space is obtained by calculating according to a first formula; the first formula is:

[0079] ;

[0080] in, is the first memory space, t is the processing time of the target object Ci, l local is the local memory access latency of the target computing device, l remote is the remote memory access latency of the target computing device to access the target memory device, B local is the local memory access bandwidth of the target computing device, B remote is the remote memory access bandwidth of the target computing device to the target memory device, and N is the amount of memory access data corresponding to the target object;

[0081] Accordingly, the allocation module is also used to:

[0082] The difference between the amount of accessed data corresponding to the target object and the first memory space is used as the second memory space.

[0083] Optionally, it also includes:

[0084] The allocation module is also used to, if the local memory space is not smaller than the first memory space, determine the first memory space and the second memory space as the allocation result of the target object, and update the local memory space to the difference between the local memory space and the first memory space; if the local memory space is smaller than the first memory space, change the first memory space to the local memory space; change the second memory space to the difference between the amount of memory access data corresponding to the target object and the local memory space; determine the changed first memory space and the changed second memory space as the allocation result of the target object, and mark the local memory space as empty.

[0085] In a third aspect, the present invention provides an electronic device, comprising:

[0086] Memory for storing computer programs;

[0087] A processor is used to execute the computer program to implement the memory allocation method for the model in the heterogeneous system disclosed above.

[0088] In a fourth aspect, the present invention provides a non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned disclosed method for allocating memory for a model in a heterogeneous system.

[0089] In a fifth aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned disclosed method for memory allocation for a model in a heterogeneous system.

[0090] It can be seen from the above scheme that the present invention provides a memory allocation method for a model in a heterogeneous system, including: obtaining a target task of a target model; the target task includes: a single iteration task of a model or a single inference task of a model; using a target computing device in a heterogeneous system to run the target task, and determining the corresponding processing process of the target task executed by each network layer in the target model in sequence as multiple processing stages; obtaining the processing time and memory access data volume corresponding to each processing stage; taking each processing stage as a target object, if the processing time of the target object is less than the total local memory access delay of the target computing device for the target object, and the memory access data volume of the target object is not greater than the local memory space of the target computing device, then allocating memory space of the size of the memory access data volume to the target object in the target computing device.

[0091] It can be seen that the beneficial effects of the present invention are: innovatively treating the corresponding processing process of the target task executed by each network layer in the model as each processing stage, and then performing memory allocation for each processing stage. That is: taking each processing stage as a target object, if the processing time of the target object is less than the total local memory access delay of the target computing power device for the target object, and the amount of memory access data of the target object is not greater than the local memory space of the target computing power device, then the memory space of the target object is allocated in the target computing power device with the amount of memory access data. This scheme thus realizes reasonable memory allocation for each processing stage of the task, under the premise of considering the total local memory access delay of the target computing power device for the execution process of a specific task, the amount of memory access data in the execution process of a specific task, and the local memory space of the target computing power device, to ensure that the memory allocation does not affect the efficiency and performance of the currently executed task as much as possible, thereby accelerating the task processing efficiency.

[0092] Correspondingly, a memory allocation device, medium and program product for a model in a heterogeneous system provided by the present invention also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0094] Figure 1 A flow chart of a memory allocation method for a model in a heterogeneous system disclosed in the present invention;

[0095] Figure 2 A schematic diagram of a heterogeneous system disclosed in the present invention;

[0096] Figure 3 A schematic diagram of a memory allocation scheme for a model in a heterogeneous system disclosed in the present invention;

[0097] Figure 4 A schematic diagram of a memory allocation device for a model in a heterogeneous system disclosed in the present invention;

[0098] Figure 5 A schematic diagram of an electronic device disclosed in the present invention;

[0099] Figure 6 A server structure diagram provided by the present invention;

[0100] Figure 7 A terminal structure diagram provided by the present invention. DETAILED DESCRIPTION

[0101] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other examples obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0102] At present, in heterogeneous systems, the memory space of a single computing device is limited, and it is often necessary to use additional memory devices to store the data to be processed and the corresponding processing results. In this separate memory scenario, due to the differences in memory performance of different memory devices and the differences in computing power of different computing devices, unreasonable memory allocation may affect the efficiency of tasks executed on the computing devices. To this end, the present invention provides a memory allocation scheme for models in a heterogeneous system, which can achieve reasonable memory allocation for each processing stage of the task under the premise of considering the total local memory access latency of the target computing device for the execution process of a specific task, the amount of memory access data in the execution process of a specific task, and the local memory space of the target computing device, to ensure that memory allocation does not affect the efficiency and performance of the currently executed task as much as possible, thereby accelerating the task processing efficiency.

[0103] See also Figure 1 As shown, an embodiment of the present invention discloses a memory allocation method for a model in a heterogeneous system, including:

[0104] S101. Obtain a target task of a target model; the target task includes: a single iteration task of a model or a single inference task of a model.

[0105] In this embodiment, the target model can be a model-related task model, a reinforcement learning model, a neural network model, etc. A single iteration task of the model includes: one forward propagation and one back propagation. A single inference task of the model includes: one forward propagation.

[0106] S102. Utilize the target computing device in the heterogeneous system to run the target task, and sequentially determine the corresponding processing process of the target task executed by each network layer in the target model into multiple processing stages.

[0107] It should be noted that the target computing power device is any device in a heterogeneous system that can run the target task. The execution order of the corresponding processing procedures of the target tasks performed by each network layer in the target model is: the arrangement order and determination order of multiple processing stages. In one embodiment, according to the execution order of the corresponding processing procedures of the target tasks performed by each network layer in the target model, the corresponding processing procedures of the target tasks performed by each network layer in the target model are determined as various processing stages, including: determining each network layer included in the target model; according to the execution order of the corresponding processing procedures of the target tasks performed by each network layer in the target model, the corresponding processing procedures of the target tasks performed by each network layer are determined as various processing stages; wherein, the same network layer corresponds to at least one processing stage. For example: if a model has 3 network layers, then for a single iteration task of the model, the corresponding processing of the target task performed by each network layer is arranged in the order of execution of the corresponding processing of the target task performed by each network layer in the target model, and there can be a target sequence: C1, C2, C3, C4, C5, C6, with a total of 6 processing stages; and C3 and C4 correspond to the last network layer, C2 and C5 correspond to the second network layer, and C1 and C6 correspond to the first network layer. For a single reasoning task of the model, the corresponding processing of the target task performed by each network layer is arranged in the order of execution of the corresponding processing of the target task performed by each network layer in the target model, and there can be a target sequence: C1, C2, C3, with a total of 3 processing stages; and C3 corresponds to the last network layer, C2 corresponds to the second network layer, and C1 corresponds to the first network layer.

[0108] In one example, before obtaining the processing time and memory access data volume corresponding to each processing stage, it also includes: arranging each processing stage into a target sequence according to the execution order of the corresponding processing process of the target task performed by each network layer in the target model; accordingly, taking each processing stage as a target object, including: starting from the first position of the target sequence, taking each processing stage in the target sequence as a target object in turn. For the target sequence: C1, C2, C3, C4, C5, C6, C1, C2, C3, C4, C5, C6 are taken as target objects respectively, and then the subsequent steps are executed.

[0109] S103: Obtain the processing duration and memory access data volume corresponding to each processing stage.

[0110] In one embodiment, the amount of memory access data corresponding to each processing stage is determined, including: calculating the amount of memory access data corresponding to each processing stage according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage. In one embodiment, the amount of memory access data corresponding to each processing stage is calculated according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage, including: if any processing stage belongs to the forward propagation of the model, the sum of the amount of data to be processed and the amount of parameters to be processed in the current processing stage is determined as the amount of memory access data corresponding to the current processing stage; if any processing stage belongs to the reverse propagation of the model, the amount of memory access data corresponding to the current processing stage is calculated according to the amount of parameters to be processed in the current processing stage. Among them, the amount of memory access data corresponding to the current processing stage is calculated according to the amount of parameters to be processed in the current processing stage, including: determining the sum of the amount of data of the calculation parameters, gradients, optimizer states and activation values ​​in the parameters to be processed in the current processing stage as the amount of memory access data corresponding to the current processing stage.

[0111] In one embodiment, determining the processing time corresponding to each processing stage includes: obtaining the peak computing power of the target computing power device; taking the ratio of the computational complexity of any processing stage to the peak computing power as the processing time corresponding to the current processing stage. Among them, obtaining the peak computing power of the target computing power device includes: querying the performance parameters of the target computing power device to obtain the peak computing power of the target computing power device; or determining the historical peak computing power of the target computing power device; or using a performance analysis tool to obtain the real-time peak computing power of the target computing power device. Specifically, the user manual of the target computing power device can be queried to obtain the peak computing power of the target computing power device; the historical peak computing power of the target computing power device is: the peak computing power of the target computing power device during historical operation. The computational complexity of any processing stage can be determined based on the structure and parameter amount of the network layer corresponding to the processing stage, and the amount of data input to the network layer. The more complex the structure of the network layer, the more parameters, and the more data input to the network layer, the higher the computational complexity of the processing stage corresponding to the network layer.

[0112] S104. Take each processing stage as the target object respectively. If the processing time of the target object is less than the total local memory access latency of the target computing device for the target object, and the amount of memory access data of the target object is not larger than the local memory space of the target computing device, then allocate memory space of the same size as the amount of memory access data for the target object in the target computing device.

[0113] In one embodiment, the local memory access bandwidth and local memory access latency of the target computing device are obtained; the ratio of the target object's memory access data volume to the local memory access bandwidth is calculated; and the sum of the ratio and the local memory access latency is taken as the total local memory access latency.

[0114] In one embodiment, if the processing time of the target object is not less than the total local memory access latency of the target computing device for the target object, a memory space of the size of the memory access data is allocated to the target object in the target memory device and the target computing device in the heterogeneous system. There is a one-to-one correspondence between the target computing device and the target memory device.

[0115] Among them, in the target memory device and the target computing power device in the heterogeneous system, a memory space of the size of the memory access data is allocated to the target object, including: calculating the first memory space allocated to the target object in the target computing power device; calculating the second memory space allocated to the target object in the target memory device; wherein the sum of the first memory space and the second memory space is equal to the memory space of the size of the memory access data.

[0116] In one example, calculating a first memory space allocated to a target object in a target computing device includes: obtaining the first memory space by calculating according to a first formula; the first formula is: ;in, is the first memory space, t is the processing time of the target object Ci, l local is the local memory access latency of the target computing device, l remote B is the remote memory access latency of the target computing device to access the target memory device. local is the local memory access bandwidth of the target computing device, B remote is the remote memory access bandwidth of the target computing device to the target memory device, and N is the amount of memory access data corresponding to the target object.

[0117] Correspondingly, calculating the second memory space allocated for the target object in the target memory device includes: taking the difference between the amount of memory access data corresponding to the target object and the first memory space as the second memory space.

[0118] In one embodiment, if the local memory space is not smaller than the first memory space, the first memory space and the second memory space are determined as allocation results of the target object, and the local memory space is updated to the difference between the local memory space and the first memory space; if the local memory space is smaller than the first memory space, the first memory space is changed to the local memory space; the second memory space is changed to the difference between the amount of memory access data corresponding to the target object and the local memory space; the changed first memory space and the changed second memory space are determined as allocation results of the target object, and the local memory space is marked as empty.

[0119] It can be seen that this embodiment innovatively regards the corresponding processing process of the target task executed by each network layer in the model as each processing stage, and then allocates memory for each processing stage. That is: each processing stage is taken as a target object. If the processing time of the target object is less than the total local memory access delay of the target computing power device for the target object, and the amount of memory access data of the target object is not greater than the local memory space of the target computing power device, then the memory space of the target object is allocated in the target computing power device with the amount of memory access data. This solution thus realizes reasonable memory allocation for each processing stage of the task under the premise of considering the total local memory access delay of the target computing power device for the execution process of a specific task, the amount of memory access data in the execution process of a specific task, and the local memory space of the target computing power device, so as to ensure that the memory allocation does not affect the efficiency and performance of the currently executed task as much as possible, thereby accelerating the task processing efficiency.

[0120] See also Figure 2 In a heterogeneous system, there are multiple heterogeneous computing chips (i.e., computing devices). Each heterogeneous computing chip is equipped with a local memory. However, the local memory space of the chip is limited. Therefore, the heterogeneous system also includes multiple non-local memory devices. In this scenario, each heterogeneous computing chip can access its main memory and its extended memory in the separated memory network, i.e. Figure 2 In this case, when a heterogeneous computing chip is training a neural network, if the memory required for training exceeds the local memory of the heterogeneous computing chip, the memory space needs to be divided to determine which data of the task is placed in the chip's local memory and which data is placed in the extended memory in the separated memory network. Heterogeneous computing chips include: GPU, FPGA, etc.

[0121] See also Figure 3 , this embodiment provides a memory allocation solution, including:

[0122] Heterogeneous computing system information collection module: used to regularly collect computing power device information and memory device information in heterogeneous computing systems. Once a neural network is issued, the information will be sent to the neural network memory allocation calculation module.

[0123] Neural network information collection module: When a neural network is issued, this module will collect relevant information of each processing stage of the target task performed by each network layer of the neural network, and send it to the neural network memory allocation calculation module.

[0124] Neural network memory allocation calculation module: Based on the collected heterogeneous computing system information and neural network information, this module will calculate the memory allocation plan for each processing stage and send it to the neural network deployment and distribution module.

[0125] Neural network deployment and delivery module: Determine the final allocation result based on the memory allocation plan of each processing stage, and deliver the allocation result and the target task executed by the related neural network to the heterogeneous computing system, so that the heterogeneous computing system can allocate memory for the target task executed by the neural network according to the allocation result and process related tasks.

[0126] Furthermore, the computing power information in the heterogeneous computing system collected by the heterogeneous computing system information collection module at regular intervals includes: the computing power peak information FLOPS that the heterogeneous computing power device can achieve. The computing power peak information can be obtained through the performance manual of the heterogeneous computing power device or related public information, and is entered into the heterogeneous computing system information collection module in advance for storage, so that it can be retrieved when needed. The computing power peak information can also be: the historical peak information during the operation of the heterogeneous computing power device, and specifically, the tool profiler provided by the deep learning framework can be run to obtain the historical peak information. Correspondingly, the memory device information collected by the heterogeneous computing system information collection module at regular intervals includes: the local main memory capacity of the heterogeneous computing power device, the bandwidth and latency information between the heterogeneous computing power device and the corresponding remote memory device (assuming that any heterogeneous computing power device and the corresponding remote memory device already have a one-to-one correspondence physically). The computing power device reads data from the local main memory, and the latency and bandwidth of the data read from the remote memory device are different. The corresponding latency and bandwidth information can be obtained using public testing tools such as sysbench.

[0127] Furthermore, when a model-related computing task is issued by the neural network information collection module, the collected neural network task information includes: the computational complexity of the computational process performed by each network layer in the model-related computing task, such as the model training task or the model reasoning task. The model training task includes forward calculation and reverse calculation; the model reasoning task only has forward calculation. The computational complexity can be estimated using mathematical methods. For example, for the forward calculation of the fully connected layer, assuming that the input data dimension is (N, D), the hidden layer weight dimension is (D, out), and the output is (N, out), then the computational complexity is FLOPs=Nx(2×D-1)×out. The reverse calculation can be estimated by multiplying the FLOPs of the forward calculation by 2, or by using some existing open source tools such as torchstat to count the FLOPs.

[0128] The memory access information of each network layer refers to the size of memory data that needs to be read when any network layer in the model-related tasks performs calculations. For the forward transmission calculation process of a certain layer, the memory data size is the input data volume + parameter volume. Among them, for the first network layer of the model, the input data volume can be the original input data; for the non-first network layer of the model, the input data volume can be the intermediate activation value data. The input data volume can be estimated by multiplying the batch-size by the tensor size of the input of the layer. For example, if the batch-size is 10 and the input size of the neural network layer is [10,10], then the input data volume is 10×10×10×data precision; the data precision can be 4 bytes when the data type is fp32. The above information can also be counted using existing open source tools such as torchstat.

[0129] For the reverse transmission calculation process performed by a certain network layer, the data size can be estimated using the parameters of the layer. For example, given a model parameter size of X, in PyTorch automatic mixed precision training, the data size includes calculation parameters, gradients, optimizer states, and activation values, where the data size of calculation parameters, gradients, and optimizer states is 16X bytes. The activation value can be obtained from the forward calculation process of the same layer, or it can be estimated by multiplying the batch-size by the size of the tensor output by the layer. This process can also be statistically analyzed using existing open source tools such as torchstat.

[0130] Based on the collected information, the neural network memory allocation calculation module first determines whether the main memory of the heterogeneous computing device can support the execution of the current neural network task, that is, whether the main memory space is greater than the sum of the memory access information of each network layer mentioned above. If it is greater, the main memory can be used directly to execute the neural network task. If it is less, the module will call the memory allocation algorithm to allocate appropriate memory for the calculation process of each network layer in the neural network task to ensure that the execution efficiency of the neural network task is as high as possible.

[0131] Neural network deployment and distribution module: Based on the memory allocation results, the neural network tasks are deployed according to the corresponding memory allocation method.

[0132] For example, if a neural network task is divided into various processing stages C1, C2, ..., Cn according to its computational logic and network layer, then the corresponding memory allocation result can be obtained based on the memory allocation algorithm: ,in The memory space allocated for a computing process in the local main memory of the corresponding computing device. The memory space allocated in the corresponding remote memory for a certain computing process; any computing device corresponds to its own remote memory device, that is, heterogeneous computing devices and remote memory devices correspond one to one.

[0133] Define the remaining allocatable main memory capacity of the computing device that executes the current task as Z, the peak computing power of the computing device that executes the task as F, and the local memory access bandwidth as , the local memory access latency is , the remote memory access bandwidth for accessing remote separate memory devices is , the remote memory access latency is .

[0134] Based on the above information, the specific memory allocation process includes:

[0135] 1. Iterate the network layer calculation process C1, C2, ..., Cn from beginning to end. If the iteration is complete, go to step 6.

[0136] 2. For the current iteration Ci, its computational complexity is obtained as M based on the collected information, and the size of the memory data that needs to be read is obtained based on the memory access information, that is, the memory access data amount N.

[0137] 3. Estimate the processing time of Ci, using the calculation formula t=M / F.

[0138] 4. If , Represents the total latency of local memory access, which means that the computing time is mainly restricted by memory access. The given memory allocation result is =N, =0.

[0139] if , which means that the computation time is mainly limited by the computation, and the memory access task can be appropriately offloaded to the remote separate memory without affecting the computation efficiency. By solving the following equation: , we can get: the first memory space , the second memory space =N- .

[0140] 5. According to the above results, if ,but , As the calculation result of this calculation process Ci, the remaining capacity of the main memory is updated to , and return to step 1 to continue iterating.

[0141] if , then the main memory available space is insufficient, and the , , and given Ci, all the distribution results are =0, The value is the memory data size N that needs to be read in the corresponding calculation process, and then go to step 6.

[0142] 6. Get the final ( , ),( , ),...,( , ) and returns it to the neural network deployment and distribution module for actually executing the neural network inference task.

[0143] This embodiment can allocate appropriate memory to each calculation process of each network layer in the neural network reasoning task, ensure that the execution efficiency of the neural network reasoning task is as high as possible, ensure that the memory allocation does not affect the calculation task as much as possible, and ensure that the execution speed of the neural network reasoning task is as fast as possible after the memory is expanded.

[0144] A memory allocation device for a model in a heterogeneous system provided by an embodiment of the present invention is introduced below. The memory allocation device for a model in a heterogeneous system described below can be referenced to other embodiments described in this document.

[0145] See also Figure 4 As shown, an embodiment of the present invention discloses a memory allocation device for a model in a heterogeneous system, including:

[0146] An acquisition module is used to acquire a target task of a target model; the target task includes: a single iteration task of a model or a single inference task of a model;

[0147] A determination module, used for running a target task using a target computing device in a heterogeneous system, and determining a corresponding processing process of the target task executed by each network layer in the target model into a plurality of processing stages in sequence;

[0148] A calculation module is used to obtain the processing time and amount of memory access data corresponding to each processing stage;

[0149] The allocation module is used to take each processing stage as the target object respectively. If the processing time of the target object is less than the total local memory access delay of the target computing power device for the target object, and the amount of memory access data of the target object is not larger than the local memory space of the target computing power device, then the memory space of the same size as the amount of memory access data is allocated to the target object in the target computing power device.

[0150] In one implementation, the determination module is specifically configured to:

[0151] Determine the various network layers included in the target model;

[0152] According to the execution order of the corresponding processing process of the target task executed by each network layer in the target model, the corresponding processing process of the target task executed by each network layer is determined as each processing stage;

[0153] Among them, the same network layer corresponds to at least one processing stage.

[0154] In one embodiment, it further includes:

[0155] A sorting module is used to arrange each processing stage into a target sequence according to the execution order of the corresponding processing process of the target task executed by each network layer in the target model before obtaining the processing time and the amount of memory access data corresponding to each processing stage;

[0156] Accordingly, the allocation module is specifically used for:

[0157] Starting from the first position of the target sequence, each processing stage in the target sequence is taken as the target object.

[0158] In one embodiment, the computing module is specifically used for:

[0159] The amount of memory access data corresponding to each processing stage is calculated respectively according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage.

[0160] In one embodiment, the computing module is specifically configured to:

[0161] If any processing stage belongs to the forward propagation of the model, the sum of the amount of data to be processed and the amount of parameters to be processed in the current processing stage is determined as the amount of memory access data corresponding to the current processing stage;

[0162] If any processing stage belongs to model back propagation, the amount of memory access data corresponding to the current processing stage is calculated based on the amount of parameters to be processed in the current processing stage.

[0163] In one embodiment, the computing module is specifically used for:

[0164] The sum of the data amounts of the calculation parameters, gradients, optimizer states, and activation values ​​in the parameters to be processed in the current processing stage is determined as the memory access data amount corresponding to the current processing stage.

[0165] In one embodiment, the computing module is specifically used for:

[0166] Get the peak computing power of the target computing device;

[0167] The ratio of the computational complexity of any processing stage to the peak computing power is taken as the processing time corresponding to the current processing stage.

[0168] In one embodiment, it further includes:

[0169] The total latency calculation module is used to obtain the local memory access bandwidth and local memory access latency of the target computing device; calculate the ratio of the target object's memory access data volume to the local memory access bandwidth; and sum the ratio and the local memory access latency as the total local memory access latency.

[0170] In one embodiment, the allocation module is further configured to:

[0171] If the processing time of the target object is not less than the total local memory access latency of the target computing device for the target object, a memory space of the same size as the memory access data volume is allocated to the target object in the target memory device and the target computing device in the heterogeneous system.

[0172] In one embodiment, the allocation module is further configured to:

[0173] Calculate the first memory space allocated for the target object in the target computing device;

[0174] Calculate a second memory space allocated for the target object in the target memory device;

[0175] The sum of the first memory space and the second memory space is equal to the memory space of the accessed data amount.

[0176] In one embodiment, the allocation module is further configured to:

[0177] The first memory space is calculated according to the first formula; the first formula is:

[0178] ;

[0179] in, is the first memory space, t is the processing time of the target object Ci, l local is the local memory access latency of the target computing device, l remote The remote memory access latency of the target computing device to the target memory device, B local is the local memory access bandwidth of the target computing device, B remote is the remote memory access bandwidth of the target computing device to the target memory device, and N is the amount of memory access data corresponding to the target object;

[0180] Accordingly, the allocation module is also used to:

[0181] The difference between the amount of accessed data corresponding to the target object and the first memory space is used as the second memory space.

[0182] In one embodiment, it further includes:

[0183] The allocation module is also used to determine the first memory space and the second memory space as the allocation results of the target object if the local memory space is not smaller than the first memory space, and update the local memory space to the difference between the local memory space and the first memory space; if the local memory space is smaller than the first memory space, change the first memory space to the local memory space; change the second memory space to the difference between the amount of memory access data corresponding to the target object and the local memory space; determine the changed first memory space and the changed second memory space as the allocation results of the target object, and mark the local memory space as empty.

[0184] Among them, for more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0185] It can be seen that this embodiment provides a memory allocation device for models in a heterogeneous system, which can realize reasonable memory allocation for each processing stage of the task under the premise of considering the total local memory access latency of the target computing power device for the execution process of a specific task, the amount of memory access data in the execution process of the specific task, and the local memory space of the target computing power device, thereby ensuring that the memory allocation does not affect the efficiency and performance of the currently executed task as much as possible, thereby accelerating the task processing efficiency.

[0186] An electronic device provided by an embodiment of the present invention is introduced below. The electronic device described below can be referenced to other embodiments described in this document.

[0187] See also Figure 5 As shown, an embodiment of the present invention discloses an electronic device, including:

[0188] Memory 501, used for storing computer programs;

[0189] The processor 502 is used to execute the computer program to implement the method disclosed in any of the above embodiments.

[0190] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: obtaining the target task of the target model; the target task includes: a single iteration task of the model or a single inference task of the model; using the target computing device in the heterogeneous system to run the target task, and determining the corresponding processing process of the target task executed by each network layer in the target model in sequence as multiple processing stages; obtaining the processing time and memory access data volume corresponding to each processing stage; taking each processing stage as a target object, if the processing time of the target object is less than the total local memory access delay of the target computing device for the target object, and the memory access data volume of the target object is not greater than the local memory space of the target computing device, then allocating memory space of the same size as the memory access data volume to the target object in the target computing device.

[0191] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: determining each network layer included in the target model; determining the corresponding processing process of the target task performed by each network layer as each processing stage according to the execution order of the corresponding processing process of the target task performed by each network layer in the target model; wherein the same network layer corresponds to at least one processing stage.

[0192] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: the various processing stages are arranged into a target sequence according to the execution order of the corresponding processing processes of the target tasks performed by each network layer in the target model.

[0193] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: starting from the first position of the target sequence, each processing stage in the target sequence is sequentially used as a target object.

[0194] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage, the amount of memory access data corresponding to each processing stage is calculated respectively.

[0195] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: if any processing stage belongs to the forward propagation of the model, the sum of the amount of data to be processed and the amount of parameters to be processed in the current processing stage is determined as the amount of memory access data corresponding to the current processing stage; if any processing stage belongs to the reverse propagation of the model, the amount of memory access data corresponding to the current processing stage is calculated based on the amount of parameters to be processed in the current processing stage.

[0196] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: the sum of the data amounts of the calculation parameters, gradients, optimizer states, and activation values ​​in the parameters to be processed in the current processing stage is determined as the memory access data amount corresponding to the current processing stage.

[0197] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: obtaining the computing power peak value of the target computing power device; taking the ratio of the computational complexity of any processing stage to the computing power peak value as the processing time corresponding to the current processing stage.

[0198] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: obtaining the local memory access bandwidth and local memory access latency of the target computing device; calculating the ratio of the target object's memory access data volume to the local memory access bandwidth; and taking the sum of the ratio and the local memory access latency as the total local memory access latency.

[0199] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: if the processing time of the target object is not less than the total local memory access latency of the target computing power device for the target object, then in the target memory device and the target computing power device in the heterogeneous system, a memory space of the same size as the memory access data volume is allocated to the target object.

[0200] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: calculating the first memory space allocated for the target object in the target computing device; calculating the second memory space allocated for the target object in the target memory device; wherein the sum of the first memory space and the second memory space is equal to the memory space of the size of the accessed data.

[0201] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: if the local memory space is not smaller than the first memory space, the first memory space and the second memory space are determined as the allocation result of the target object, and the local memory space is updated to the difference between the local memory space and the first memory space; if the local memory space is smaller than the first memory space, the first memory space is changed to the local memory space; the second memory space is changed to the difference between the amount of memory access data corresponding to the target object and the local memory space; the changed first memory space and the changed second memory space are determined as the allocation result of the target object, and the local memory space is marked as empty.

[0202] Furthermore, an embodiment of the present invention also provides an electronic device. The electronic device can be Figure 6 The server shown can also be Figure 7 Terminal shown. Figure 6 and Figure 7 All of them are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the diagrams cannot be regarded as any limitation on the scope of application of the present invention.

[0203] Figure 6A schematic diagram of the structure of a server provided in an embodiment of the present invention. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the relevant steps in the memory allocation for the model in the heterogeneous system disclosed in any of the aforementioned embodiments.

[0204] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present invention, and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0205] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.

[0206] The operating system is used to manage and control the hardware devices and computer programs on the server to realize the operation and processing of the data in the memory by the processor, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the memory allocation method for the model in the heterogeneous system disclosed in any of the aforementioned embodiments, the computer program can further include computer programs that can be used to complete other specific tasks. In addition to data such as application update information, data can also include data such as application developer information.

[0207] Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of the present invention, wherein the terminal may specifically include but is not limited to a smart phone, a tablet computer, a laptop computer or a desktop computer.

[0208] Generally, the terminal in this embodiment includes: a processor and a memory.

[0209] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0210] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is used to store at least the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the memory allocation method for the model in the heterogeneous system executed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, update information of the application.

[0211] In some embodiments, the terminal may also include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0212] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than those shown in the figure.

[0213] A nonvolatile storage medium provided in an embodiment of the present invention is introduced below. The nonvolatile storage medium described below and other embodiments described herein may be cross-referenced.

[0214] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the memory allocation method for a model in a heterogeneous system disclosed in the above-mentioned embodiment. The non-volatile storage medium is a computer-readable non-volatile storage medium, which, as a carrier for resource storage, may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system, a computer program, and data, etc., and the storage method may be temporary storage or permanent storage.

[0215] A computer program product provided by an embodiment of the present invention is introduced below. The computer program product described below can be referenced to other embodiments described in this document.

[0216] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned disclosed method for allocating memory for a model in a heterogeneous system.

[0217] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0218] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.

[0219] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A memory allocation method for a model in a heterogeneous system, characterized in that: include: Get the target task of the target model; The target tasks include: a single model iteration task or a single model reasoning task; Utilize the target computing device in the heterogeneous system to run the target task, and sequentially determine the corresponding processing process of the target task executed by each network layer in the target model into multiple processing stages; Get the processing time and amount of data accessed for each processing stage. Each processing stage is respectively taken as a target object. If the processing time of the target object is less than the total local memory access delay of the target computing power device for the target object, and the memory access data volume of the target object is not greater than the local memory space of the target computing power device, a memory space of the size of the memory access data volume is allocated to the target object in the target computing power device; wherein, the ratio of the memory access data volume of the target object to the local memory access bandwidth of the target computing power device is calculated; the sum of the ratio and the local memory access delay of the target computing power device is taken as the total local memory access delay; If the processing time of the target object is not less than the total local memory access delay, a memory space of the size of the memory access data is allocated to the target object in the target memory device and the target computing power device in the heterogeneous system.

2. The method according to claim 1, characterized in that The corresponding processing process of the target task performed by each network layer in the target model is sequentially determined into a plurality of processing stages, including: Determining each network layer included in the target model; According to the execution order of the corresponding processing procedures of the target task executed by each network layer in the target model, the corresponding processing procedures of the target task executed by each network layer are determined as each processing stage; Among them, the same network layer corresponds to at least one processing stage.

3. The method according to claim 1, characterized in that: Before obtaining the processing time and amount of memory access data corresponding to each processing stage, it also includes: Arranging the processing stages into a target sequence according to the execution order of the corresponding processing processes of the target task performed by each network layer in the target model; Accordingly, each processing stage is taken as a target object, including: Starting from the first position of the target sequence, each processing stage in the target sequence is taken as the target object in turn.

4. The method according to claim 1, characterized in that Determine the amount of memory access data corresponding to each processing stage, including: The amount of memory access data corresponding to each processing stage is calculated respectively according to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage.

5. The method according to claim 4, characterized in that According to the amount of data to be processed and / or the amount of parameters to be processed in each processing stage, the amount of memory access data corresponding to each processing stage is calculated respectively, including: If any processing stage belongs to the forward propagation of the model, the sum of the amount of data to be processed and the amount of parameters to be processed in the current processing stage is determined as the amount of memory access data corresponding to the current processing stage; If any processing stage belongs to model back propagation, the amount of memory access data corresponding to the current processing stage is calculated based on the amount of parameters to be processed in the current processing stage.

6. The method according to claim 5, characterized in that According to the amount of parameters to be processed in the current processing stage, the amount of memory access data corresponding to the current processing stage is calculated, including: The sum of the data amounts of the calculation parameters, gradients, optimizer states, and activation values ​​in the parameters to be processed in the current processing stage is determined as the memory access data amount corresponding to the current processing stage.

7. The method according to claim 1, characterized in that Determine the processing time corresponding to each processing stage, including: Obtaining the peak computing power of the target computing power device; The ratio of the computational complexity of any processing stage to the computing power peak is taken as the processing duration corresponding to the current processing stage.

8. The method according to claim 1, characterized in that: The local memory access bandwidth and the local memory access latency are obtained.

9. The method according to any one of claims 1 to 8, characterized in that: In the target memory device and the target computing power device in the heterogeneous system, allocating a memory space of the size of the access data amount for the target object includes: Calculate a first memory space allocated for the target object in the target computing device; Calculating a second memory space allocated for the target object in the target memory device; The sum of the first memory space and the second memory space is equal to the memory space of the access data amount.

10. The method according to claim 9, characterized in that Calculating a first memory space allocated for the target object in the target computing power device includes: The first memory space is obtained by calculating according to a first formula; the first formula is: ; in, is the first memory space, t For the target object Ci The processing time, l local is the local memory access latency of the target computing device, l remote The remote memory access latency of the target computing device to access the target memory device, B local is the local memory access bandwidth of the target computing device, B remote The remote memory access bandwidth of the target computing device to the target memory device, N The amount of accessed data corresponding to the target object; Accordingly, calculating the second memory space allocated for the target object in the target memory device includes: The difference between the amount of accessed data corresponding to the target object and the first memory space is used as the second memory space.

11. The method according to claim 9, characterized in that Also includes: If the local memory space is not smaller than the first memory space, determining the first memory space and the second memory space as the allocation result of the target object, and updating the local memory space to the difference between the local memory space and the first memory space; If the local memory space is smaller than the first memory space, changing the first memory space to the local memory space; The second memory space is changed to the difference between the amount of accessed data corresponding to the target object and the amount of the local memory space; The modified first memory space and the modified second memory space are determined as allocation results of the target object, and the local memory space is marked as empty.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 11.

13. A non-volatile storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Automatic test case generation device based on large language model

    CN117806980A

  • Memory scheduling method of heterogeneous computing system, heterogeneous computing system and device

    CN118260053A