Tensor Management Method, Electronic Device, Storage Medium and Program Product

By introducing tensor pools and intelligent tensor data structures in deep learning model training, the problem of inconsistent memory addresses allocated by Allocator is solved, and tensors with the same memory address are obtained in each iteration, improving training performance and communication efficiency.

CN119718693BActive Publication Date: 2025-05-30SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510239452.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-30
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

During the deep learning model training process, the memory address allocated by Allocator to the tensor each time is different, resulting in the inability to meet the performance optimization requirements. Especially in distributed training scenarios, communication efficiency decreases and performance acceleration cannot be achieved.

Method used

By introducing tensor pools and intelligent tensor data structures, based on the attribute information of the tensor required by the current operator, we look up whether the corresponding tensor queue exists from the tensor pool. If it exists, the target tensor will be taken out and automatically put back into the tensor queue after use is completed.

Benefits of technology

It realizes tensors that obtain the same memory address in each iteration of model training, reducing the overhead of memory allocation and release, and improving communication efficiency and overall training performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718693B_ABST
    Figure CN119718693B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and provides a tensor management method, an electronic device, a storage medium, and a program product. The method includes: based on the attribute information of the tensors required by the current operator, searching in the tensor pool to find whether there is a corresponding tensor queue, and the tensor queue stores instance objects of each tensor; in the case where a corresponding tensor queue is found in the tensor pool, based on the storage order of the tensors in the tensor queue, taking out the target tensor from the tensor queue; based on the current operator, using the target tensor to participate in model training, and after use, automatically putting the target tensor back into the tensor queue. By introducing a tensor pool and a tensor queue, the present invention can realize taking out a tensor from the queue when using it and automatically putting it back into the queue after use, thereby enabling the operator to obtain tensors with the same memory address in each iteration of model training, thus improving the performance of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a tensor management method, an electronic device, a storage medium, and a program product. Background Art

[0002] During the training process of a deep learning model, as the basic unit of data, tensors (Tensors) are crucial for improving training performance and resource utilization. Deep learning frameworks (such as PyTorch) provide a resource management pool called Allocator, which is responsible for the memory allocation and recycling of tensors. During the model training process, the allocation of each tensor needs to be applied from the Allocator. If there is sufficient memory resources inside the Allocator, it will directly allocate the tensor; if the resources are insufficient, the Allocator will apply to the driver for memory allocation and then allocate the tensor.

[0003] During the model training process, the Allocator manages most of the memory resources. This management mechanism can effectively reduce the number of times of directly applying for memory from the driver, avoid memory application blocking calculations during the training process, and significantly save time overhead. However, the memory address allocated by the Allocator for each tensor is not guaranteed to be the same, which has become a major obstacle to performance optimization during the model training process. In scenarios that require high-performance communication (such as distributed training), if the communication operator cannot obtain tensor objects with the same memory address during each iteration of training, it will lead to a decrease in communication efficiency and cannot play a role in accelerating performance. Summary of the Invention

[0004] The present invention provides a tensor management method, an electronic device, a storage medium, and a program product to solve the defect that the memory addresses allocated for tensors each time during the model training process are not the same, resulting in the inability to meet the performance optimization requirements.

[0005] The present invention provides a tensor management method, including:

[0006] Based on the attribute information of the tensors required by the current operator, check whether there is a corresponding tensor queue in the tensor pool. The attribute information of each tensor in the tensor queue is the same as the attribute information of the tensors required by the current operator, and instance objects of the tensors are stored in the tensor queue;

[0007] When it is found that there is a corresponding tensor queue in the tensor pool, based on the storage order of each tensor in the tensor queue, take out the target tensor from the tensor queue;

[0008] Based on the current operator, use the target tensor to participate in the model training, and after use, automatically put the target tensor back into the tensor queue.

[0009] According to a tensor management method provided by the present invention, before using the target tensor to participate in model training based on the current operator, it further includes:

[0010] In the case where the corresponding tensor queue is not found in the tensor pool, a target tensor is created based on the attribute information of the tensors required by the current operator.

[0011] According to a tensor management method provided by the present invention, using the target tensor to participate in model training based on the current operator includes:

[0012] Create a weak reference object for the target tensor, register the weak reference object in the tensor pool, and set the target tensor as the original tensor attribute of the weak reference object;

[0013] Based on the current operator, use the weak reference object to replace the target tensor to participate in model training.

[0014] According to a tensor management method provided by the present invention, automatically putting the target tensor back into the tensor queue after use includes:

[0015] In the case where it is detected that the life cycle of the weak reference object ends, automatically put the target tensor back into the tensor queue and recycle the weak reference object.

[0016] According to a tensor management method provided by the present invention, automatically putting the target tensor back into the tensor queue includes:

[0017] Obtain the target tensor based on the original tensor attribute of the weak reference object;

[0018] Based on the attribute information of the target tensor, check whether there is a corresponding tensor queue in the tensor pool;

[0019] In the case where a corresponding tensor queue is found in the tensor pool, add the target tensor to the end of the tensor queue;

[0020] In the case where no corresponding tensor queue is found in the tensor pool, create a new tensor queue in the tensor pool based on the attribute information of the target tensor and add the target tensor to the new tensor queue.

[0021] According to a tensor management method provided by the present invention, detecting that the life cycle of the weak reference object ends includes:

[0022] When it is detected that the reference count of the weak reference object reaches zero, it is determined that the life cycle of the weak reference object ends.

[0023] According to a tensor management method provided by the present invention, before automatically putting the target tensor back into the tensor queue, it further includes:

[0024] Checking whether the weak reference object has been registered in the tensor pool;

[0025] If the weak reference object has been registered in the tensor pool, unregister the weak reference object from the tensor pool.

[0026] According to a tensor management method provided by the present invention, the tensor pool is a data structure for storing and recycling tensors that need to be reused during model training. Each element in the tensor pool is stored in the form of a key-value pair. The key of any element is determined based on the attribute information of the tensor, and the value of any element is a first-in-first-out queue, and at least one instance object of the tensor is stored in the queue.

[0027] The present invention also provides a tensor management device, including:

[0028] A tensor lookup unit, configured to look up in the tensor pool whether there is a corresponding tensor queue based on the attribute information of the tensor required by the current operator, where the attribute information of each tensor in the tensor queue is the same as the attribute information of the tensor required by the current operator, and instance objects of the tensors are stored in the tensor queue;

[0029] A tensor extraction unit, configured to, when it is found that there is a corresponding tensor queue in the tensor pool, extract a target tensor from the tensor queue based on the storage order of each tensor in the tensor queue;

[0030] A tensor usage unit, configured to use the target tensor to participate in model training based on the current operator, and automatically put the target tensor back into the tensor queue after use.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it implements the tensor management method as described in any one of the above.

[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the tensor management method as described in any one of the above.

[0033] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the tensor management method as described in any one of the above.

[0034] The tensor management method, electronic device, storage medium, and program product provided by the present invention, based on the attribute information of the tensors required by the current operator, search in the tensor pool to find whether there is a corresponding tensor queue. When the corresponding tensor queue is found, the target tensor can be directly taken out from the tensor queue according to the storage order of the tensors in the queue, without creating a new tensor, which can reduce the overhead of memory allocation and release. After taking out the target tensor, the current operator can use the target tensor to participate in model training, and automatically put the target tensor back into the tensor queue after use. Since the tensors in the tensor queue are arranged in a certain storage order, such as in the first-in-first-out order, and the instance objects of the tensors are stored in the tensor queue, the memory address of the tensor can be obtained through the instance object. Therefore, for the current operator, the position order of the target tensor taken out in each iteration of model training in the tensor queue is the same. Thus, the operator using the target tensor will obtain tensors with the same memory address in each iteration, which not only realizes the reuse of tensor memory resources, but also can reduce the complexity of data transmission and synchronization, and reduce memory access latency, thereby playing a role in performance acceleration. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 is a schematic flowchart of the tensor management method provided by the present invention;

[0037] Figure 2 is a schematic flowchart of the tensor data structure and its resource management method provided by the present invention;

[0038] Figure 3 is a schematic structural diagram of the tensor management device provided by the present invention;

[0039] Figure 4 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.

[0041] Performance optimization is crucial in model training as it directly relates to the training speed and resource utilization efficiency. In distributed training or model parallelization scenarios, communication operations (such as gradient synchronization, parameter updates, etc.) account for a large portion of the training time. To ensure the efficient execution of these communication operations, it is generally necessary to ensure that the tensor objects operated on by the communication operator have the same memory address during each iteration of the model training. Here, the communication operator is a set of global communication operations responsible for synchronizing data and parameters between different computing nodes to ensure the consistency and accuracy of model training. Among them, global communication operations refer to those operations involving data exchange between multiple computing nodes or processors, which can be used for data synchronization, result aggregation, or information distribution, etc. It should be understood that in distributed training or model parallelization scenarios, the communication operator needs to transfer data between different computing nodes, and these data usually exist in the form of tensors. To ensure the efficiency of communication, the communication operator needs to be able to directly access and operate on these tensor objects, and to access and operate on these tensor objects, the communication operator needs to know the location of these tensor objects in memory, that is, their memory addresses.

[0042] In model training, especially in distributed training scenarios, if the communication operator can obtain tensors with the same memory address in each iteration, it can enable the operator to efficiently perform data transmission and synchronization during training without having to frequently remap the memory address in each iteration, reducing unnecessary memory mapping overhead. Moreover, by reducing the variation of memory addresses, the communication efficiency can be significantly improved, thereby accelerating the overall training process. In addition, tensors with the same memory address also mean a better cache hit rate, which can further reduce the memory access latency and thus improve the data access efficiency.

[0043] During the model training process, tensors serve as the basic units of data, and their efficient management is crucial for improving training performance and resource utilization. The PyTorch framework provides a resource management pool (i.e., Allocator) responsible for the memory allocation and recycling of tensors. During the model training process, whenever a new tensor needs to be created, PyTorch sends a memory resource request to the Allocator. The Allocator first checks whether the managed memory resource pool has sufficient resources to meet the request. If there are sufficient resources, the Allocator directly allocates memory from the resource pool to the new tensor; if the resources are insufficient, the Allocator requests more memory resources from the graphics card driver and allocates them to the new tensor after obtaining the resources. The allocated memory resources are used to store the data of the tensor. When the lifecycle of a tensor ends (e.g., the tensor is deleted or goes out of scope), the Allocator automatically reclaims the occupied memory resources and returns them to the resource pool for subsequent use.

[0044] The Allocator manages most of the memory resources, reducing the number of direct requests for memory from the graphics card driver, thus avoiding blocking calculations when applying for video memory during the model training process and helping to maintain the stability and efficiency of training. Moreover, by pre-managing and recycling memory resources, the Allocator reduces the number of interactions with the graphics card driver, thus saving a large amount of time for driver requests.

[0045] Although the Allocator of PyTorch performs well in memory resource management, the memory addresses it allocates to tensors each time are not guaranteed to be the same, mainly for considerations of efficiency and flexibility. When managing memory resources, the Allocator dynamically allocates addresses according to the current memory usage and application requirements to maximize memory utilization and reduce fragmentation. However, this feature of the Allocator has become a major obstacle to performance optimization. Each time the Allocator allocates memory to a tensor, it may return different addresses, which violates the requirement of performance optimization for memory address stability. Moreover, due to the instability of memory addresses, communication operators have to remap memory frequently, increasing communication latency and overhead.

[0046] To enable communication operators to obtain tensors with the same memory address in each iteration, related techniques propose manually creating a tensor object before the communication operator and performing communication operator operations based on this object. This is because the manually created tensor object is persistent during its lifecycle, and its memory address is fixed after creation. Therefore, in each iteration, as long as the same tensor object is used for communication operations, it can be ensured that tensors with the same address are obtained in each iteration.

[0047] In PyTorch, the management of memory resources is usually automatically completed by the Allocator. When the lifecycle of a tensor ends, the Allocator will automatically reclaim the memory resources it occupies for other tensors to use. However, in the case of manually creating tensor objects, this automatic recycling mechanism no longer applies, and instead, users need to manually reclaim and manage tensor resources. To achieve the manual recycling and management of tensor resources, users need to carefully read the code, judge based on professional experience, find the places where tensors are no longer in use, and manually modify the code for recycling. This operation not only increases the difficulty and complexity of code writing, but also, if the recycling position is incorrect or omitted, it will cause the program to crash and be difficult to troubleshoot and locate.

[0048] In response to this, the present invention proposes a tensor management method. By introducing a tensor pool (i.e., TensorPool) and an intelligent tensor data structure (i.e., SmartTensor), the above problems can be effectively solved. Among them, the tensor pool is used to solve the problem that the memory addresses allocated by the Allocator of PyTorch are different each time, and the tensor data structure is used to solve the problem of being unable to automatically recycle and manage tensor resources. The technical solutions provided by the present invention will be introduced in detail below.

[0049] Figure 1 is a schematic flowchart of the tensor management method provided by the present invention. As Figure 1 shown, the method includes:

[0050] Step 110, based on the attribute information of the tensor required by the current operator, check whether there is a corresponding tensor queue in the tensor pool. The attribute information of each tensor in the tensor queue is the same as the attribute information of the tensor required by the current operator, and instance objects of the tensors are stored in the tensor queue.

[0051] Specifically, the current operator refers to the operation or function that is currently being called or executed during model training. In the embodiments of the present invention, the current operator mainly refers to a communication operator, that is, an operation involving data transmission between different devices. The tensor required by the current operator refers to the tensor that needs to be used when the current operator is executed. Here, a tensor is a commonly used data structure in deep learning models, which can be regarded as a multi-dimensional array or matrix for storing the input, output, or intermediate results of an operator.

[0052] The attribute information of the tensors required by the current operator may include the data type of the tensors (denoted as type), the device type (denoted as device), the data shape (denoted as shape), etc. Among them, the data type refers to the data type of the elements in the tensor, such as float32, float64, int64, etc.; the device type refers to the hardware device type where the tensor is located, such as CPU (Central Processing Unit, central processor), GPU (Graphics Processing Unit, graphics processor), etc.; the data shape refers to the size of the tensor in each dimension.

[0053] For the current operator, according to the attribute information of the tensors it requires, it can first check whether there is a corresponding tensor queue in the tensor pool. Here, the tensor pool is a data structure pre-created for storing, managing, and recycling the tensors that all communication operators need to reuse during the model training process. Specifically, in order to check whether there is a corresponding tensor queue, the attribute information of the tensors required by the current operator can be used as the query condition to search for a matching tensor queue in the tensor pool.

[0054] It can be understood that during the model training process, there may be multiple different communication operators that call tensors with the same attribute information. Therefore, many tensors with the same attribute information may be created during the model training process. To manage these tensors, tensor queues are introduced in the tensor pool. A tensor queue refers to a set of tensors with the same attribute information stored in the tensor pool. All tensors in each tensor queue have the same data type, device type, and data shape, but the instance objects corresponding to each tensor are different, and the data elements they store are also different.

[0055] It should be noted that in actual operations, the memory address of the tensor is usually not directly stored in the tensor queue, but the memory address is encapsulated into the instance object of the tensor, and the instance object is stored in the tensor queue. After obtaining the instance object of a certain tensor from the tensor queue, the corresponding memory address can be obtained through the instance object. In other words, obtaining the instance object of the tensor is equivalent to obtaining the memory address of the tensor, so that the corresponding tensor data can be accessed and operated on.

[0056] Step 120, when it is found that there is a corresponding tensor queue in the tensor pool, based on the storage order of the tensors in the tensor queue, take out the target tensor from the tensor queue.

[0057] Specifically, if a tensor queue matching the attribute information of the tensors required by the current operator is found in the tensor pool, it indicates that during the previous iteration of model training, corresponding tensors have been created for the current operator and stored in the tensor queue. In this case, the tensors previously created for this operator can be directly retrieved according to the storage order of each tensor in the tensor queue, that is, the target tensors required by the current operator can be retrieved. This enables the current operator to reuse the previously created tensors without having to create new tensors again, thus avoiding repeated application and destruction of memory resources. It should be understood that the current operator reusing the previously created tensors means reusing the data structure and memory resources of the tensors, thereby reducing the overhead of memory allocation and release.

[0058] The storage order of each tensor in the tensor queue refers to the arrangement method of tensors in the queue. In the tensor queue, each tensor is arranged in the order in which they are added to the queue. In other words, the tensors in the tensor queue can be arranged according to an order-preserving method to ensure that the communication operator can obtain tensors with the same memory address in each iteration of model training. For example, the tensor queue can be a first-in-first-out queue, which means that the first tensor in the queue (i.e., the head of the queue) is the earliest added, and the last tensor (i.e., the tail of the queue) is the most recently added.

[0059] Specifically, for each communication operator, when a corresponding tensor queue is matched from the tensor pool according to the attribute information of the tensors required by the operator, the target tensor of the operator can be retrieved from the head of the tensor queue according to the storage order of each tensor in the queue (such as the first-in-first-out order). After the operator uses the target tensor, it is automatically placed back at the tail of the queue for the operator to reuse in the next iteration. It should be understood that since the tensor pool is searched based on the attribute information of the tensors, during multiple iterations of model training, as long as the attribute information of the tensors required by the current operator remains unchanged, it will always search for tensors from the same tensor queue. The tensors in the tensor queue are arranged in a certain storage order (such as arranged in the first-in-first-out order). In each iteration, when the operator retrieves the target tensor from the tensor queue, it is based on this stable storage order. Due to the stability of the tensor queue and the storage order, the position of the target tensor retrieved by the operator from the tensor queue in each iteration is the same. This means that in multiple iterations, the operator will reuse the tensor at the same position in the queue. That is, for a certain operator, it will reuse the same tensor in each iteration. Since the instance object of the tensor obtained by the operator in each iteration remains unchanged, it can ensure that the operator can obtain tensors with the same memory address in each iteration, thus playing a role in accelerating performance.

[0060] It can be understood that the target tensor refers to the tensor taken out from the tensor queue for the current operator operation, which is determined by matching the requirements of the current operator with the attribute information of the tensors in the tensor queue. In each iteration, the target tensor is taken out from the queue for model training and put back into the queue after use for subsequent reuse.

[0061] Step 130, based on the current operator, use the target tensor to participate in model training, and automatically put the target tensor back into the tensor queue after use.

[0062] Specifically, for the current operator, when the target tensor is taken out and passed to the current operator, the operator processes the tensor according to its specific operation logic. For example, the current operator can access the memory address of the target tensor, read the tensor data from it, and perform operations such as calculation, transformation, or transmission on the tensor data. During model training, these operations are continuously iterated to gradually optimize the parameters and performance of the model. When the target tensor is used up, it is automatically put back to the end of the tensor queue to ensure the recycling and orderly management of the tensors.

[0063] In addition, if no tensor queue matching the attribute information of the tensor required by the current operator is found in the tensor pool, it indicates that this may be the first iteration of model training, and the tensor required by the current operator is used for the first time, and no corresponding tensor has been created for the current operator before. In this case, a target tensor can be created according to the attribute information of the tensor required by the current operator so that the current operator can use this target tensor to participate in model training. Here, creating a target tensor means allocating corresponding memory resources for the tensor according to the attribute information of the tensor required by the current operator (such as data type, device type, and data shape, etc.). Subsequently, after the target tensor is used up, the target tensor can be put into the tensor pool. Here, when putting the target tensor into the tensor pool, a new tensor queue can be created in the tensor pool according to the attribute information of the target tensor, and the target tensor is added to this new tensor queue for subsequent reuse.

[0064] The method provided by the embodiment of the present invention, based on the attribute information of the tensor required by the current operator, looks up whether there is a corresponding tensor queue in the tensor pool. When the corresponding tensor queue is found, the target tensor can be directly taken out from the tensor queue according to the storage order of each tensor in the tensor queue, without creating a new tensor, which can reduce the overhead of memory allocation and release. After taking out the target tensor, the current operator can use the target tensor to participate in model training and automatically put the target tensor back into the tensor queue after use. Since the tensors in the tensor queue are arranged in a certain storage order, for example, in the first-in-first-out order, and the instance objects of each tensor are stored in the tensor queue, and the memory address of the tensor can be obtained through the instance object, for the current operator, the position order of the target tensor taken out in each iteration of model training in the tensor queue is the same. Therefore, the operator using the target tensor will obtain tensors with the same memory address in each iteration, which not only realizes the reuse of tensor memory resources, but also can reduce the complexity of data transmission and synchronization and reduce the memory access latency, thus playing a role in performance acceleration.

[0065] Based on the above embodiment, the tensor pool is a data structure for storing and recycling tensors that need to be reused during model training. Each element in the tensor pool is stored in the form of a key-value pair. The key of any element is determined based on the attribute information of the tensor, and the value of any element is a first-in-first-out queue, and at least one instance object of the tensor is stored in the queue.

[0066] Specifically, in order to quickly find out whether there is a tensor queue in the tensor pool that matches the attribute information of the tensor required by the operator, the tensor pool can be implemented using a hash table, that is, by using a specific key-value pair (key-value) to store and retrieve the tensor queue. Each element in the tensor pool corresponds to a unique key, and this key is a unique identifier composed of the attribute information of the tensor, that is, the type, device, and shape of the tensor. The value can be a first-in-first-out queue, and one or more instance objects of the tensor can be stored in the queue. Through the instance object, the memory address of the tensor can be obtained. It should be understood that the queue structure makes the allocation and recycling of tensors more orderly, and can provide tensors with the same memory address for the communication operator in each iteration, thereby improving the performance of model training.

[0067] For each operator, when it is necessary to find the corresponding target tensor, first, the corresponding key can be generated according to the attribute information of the tensors required by the operator, and then it is checked whether the key exists in the hash table. If the key to be searched does not exist in the tensor pool, it indicates that the tensor required by the operator is used for the first time. In this case, a new tensor can be created according to the attribute information of the tensors required by the operator, and after the tensor is used up, it can be put into the tensor pool. If the key to be searched exists in the tensor pool, it indicates that there is a tensor queue currently being searched in the tensor pool. At this time, the required target tensor can be directly taken out from the head of the tensor queue, and after the target tensor is used up, it can be automatically put back to the tail of the tensor queue.

[0068] It can be understood that during the entire model training process, the tensors in the tensor pool are dynamically transformed because the currently required tensors are taken out of the queue during use and then put back into the queue after use. Since model training is an iterative process, in each iteration, the order in which the operator obtains the target tensors from the tensor queue is the same, so that the operators that call the target tensors will obtain tensors with the same memory address, thus playing a role in performance acceleration.

[0069] Based on any of the above embodiments, before step 130, the method further includes:

[0070] In the case where the corresponding tensor queue is not found in the tensor pool, a target tensor is created based on the attribute information of the tensors required by the current operator.

[0071] Specifically, according to the attribute information of the tensors required by the current operator, if a matching tensor queue is not found in the tensor pool, it indicates that the tensor required by the current operator is used for the first time and no corresponding tensor has been created for this operator and stored in the tensor pool before. In this case, a target tensor can be created for it according to the attribute information of the tensors required by the current operator. Here, when creating the target tensor, the torch.empty interface provided by the PyTorch framework can be called for creation. Subsequently, the current operator can use the target tensor to participate in model training, and after the target tensor is used up, it can be put into the tensor pool for subsequent reuse.

[0072] Exemplarily, assume that in each iteration of model training, operations of communication operator A, communication operator B, and communication operator C are involved. Before model training, an empty tensor pool can be initialized and created. In the first iteration of model training, for communication operator A, assume that the attribute information of the required tensor is type1 + device1 + shape1. Since the tensor pool is empty at this time, a target tensor TensorA can be created according to the attribute information of the tensor required by communication operator A, so that communication operator A can use TensorA to participate in model training. Similarly, assume that the attribute information of the tensor required by communication operator B is also type1 + device1 + shape1, while the attribute information of the tensor required by communication operator C is type2 + device2 + shape2. Corresponding target tensors TensorB and TensorC can be created for communication operator B and communication operator C respectively, so that they can use their respective target tensors to participate in model training.

[0073] After the target tensor TensorA is used up, it is automatically put into the tensor pool. At this time, according to the attribute information of TensorA (i.e., type1 + device1 + shape1), a tensor queue Q1 can be created in the tensor pool, and TensorA is added to the tensor queue Q1. After the target tensor TensorB is used up, a key can be generated according to its attribute information (i.e., type1 + device1 + shape1), and the tensor pool can be searched according to this key. At this time, the tensor queue Q1 can be retrieved. Therefore, TensorB can be directly added to the end of the tensor queue Q1. Similarly, after the target tensor TensorC is used up, a key can be generated according to the attribute information of TensorC (i.e., type2 + device2 + shape2), and the tensor pool can be searched according to this key. At this time, no matching tensor queue can be retrieved. Therefore, a new tensor queue Q2 can be created in the tensor pool according to the attribute information of TensorC, and TensorC is added to the tensor queue Q2.

[0074] In the second iteration of model training, when the communication operator A is executed, a corresponding key can be generated according to the attribute information of the required tensor (i.e., type1 + device1 + shape1), and the key is used to retrieve from the tensor pool. At this time, a matching tensor queue Q1 can be retrieved. Therefore, the target tensor TensorA can be directly taken out from the head of the tensor queue Q1. Subsequently, the communication operator A can use TensorA to participate in model training. After TensorA is used, it can be automatically put back to the tail of the tensor queue Q1. Similarly, when the communication operator B is executed, according to the attribute information of the required tensor, the tensor queue Q1 can be retrieved from the tensor pool, and the target tensor TensorB can be taken out from the head of the tensor queue Q1 to participate in model training.

[0075] It should be understood that for the tensors required by each communication operator, usually only tensor creation is performed when the tensor is used for the first time. After one iteration of model training is completed, the tensor pool already contains all the tensors required by the communication operators. Therefore, in subsequent iteration processes, there is no need to recreate new tensors. Each communication operator can take out the corresponding target tensors from the tensor pool for reuse, thereby achieving the purpose of performance acceleration.

[0076] Based on any of the above embodiments, in step 130, the using the target tensor to participate in model training based on the current operator includes:

[0077] Step 131, creating a weak reference object for the target tensor, registering the weak reference object in the tensor pool, and setting the target tensor as the original tensor attribute of the weak reference object.

[0078] It should be noted that in order to achieve automatic recycling of tensor resources, an intelligent tensor data structure SmartTensor is introduced in the embodiments of the present invention. This data structure is used to construct a weak reference object of the tensor. This data structure inherits from torch.Tensor. Therefore, SmartTensor contains all the attributes and methods of torch.Tensor, and its usage method is the same as that of torch.Tensor. Here, torch.Tensor is a tensor data structure defined in the PyTorch framework.

[0079] It can be understood that a weak reference refers to a reference that does not increase the reference count of a tensor. Here, the reference count refers to the number of times an object (such as a tensor) is referenced. In programming languages such as Python, a normal reference (i.e., a strong reference) increases the reference count of an object, thus preventing the object from being reclaimed by the garbage collector. A weak reference, however, does not increase the reference count of an object and cannot guarantee the survival of the object. Therefore, when the only remaining reference to a certain tensor is a weak reference, the garbage collector can reclaim the tensor and reuse its memory resources for other purposes. It should be understood that the garbage collector is part of the programming language runtime environment and is responsible for automatically managing memory, especially for automatically reclaiming memory spaces that are no longer used during program execution.

[0080] Specifically, for the current operator, after obtaining its corresponding target tensor, a corresponding function can be called to create a weak reference object for the target tensor. Exemplarily, in the Python programming language, the weakref module provides the function to create weak references, and the weakref.ref() function can be used to create a weak reference object. For example, assuming tensor is the target tensor, then smart_tensor = weakref.ref(tensor) will create a weak reference object smart_tensor pointing to tensor. Here, tensor is the tensor object created by torch.Tensor, and smart_tensor is the created SmartTensor object.

[0081] After creating the weak reference object for the target tensor, it can be registered in the tensor pool. For example, the id of the weak reference object smart_tensor can be registered in the tensor pool so that the tensor pool can manage and track the lifecycle of the target tensor. Since weak references do not increase the reference count of tensors, the tensor pool can be notified when there are no other strong references to the tensors, enabling them to be safely reclaimed.

[0082] Furthermore, since weak references cannot guarantee the survival of tensor objects, to prevent the target tensor from being reclaimed by PyTorch's Allocator due to scope detachment, the target tensor can be set as the original tensor attribute of the weak reference object, which means maintaining a strong reference to the target tensor within the weak reference object. For example, the target tensor `tensor` can be set as an original tensor attribute `original_tensor` of the weak reference object `smart_tensor`, i.e., `smart_tensor.original_tensor = tensor`. It should be understood that since weak references do not increase the reference count of the target tensor and cannot guarantee the survival of the tensor object, if the target tensor has no other strong references, it may be reclaimed by the garbage collector. By setting the target tensor as an attribute of the weak reference object, it can be ensured that the target tensor has at least one strong reference, thus preventing it from being accidentally reclaimed.

[0083] In the embodiments of the present invention, by creating a weak reference object for the target tensor, registering the weak reference object in the tensor pool, and setting the target tensor as the original tensor attribute of the weak reference object, the constructed weak reference object not only includes all the functions of the target tensor, but also isolates the management of the target tensor by the Allocator, so that it will not be reclaimed by the Allocator, but will be reclaimed by the tensor pool after the end of its life cycle, in order to create a new weak reference during the next reuse.

[0084] Step 132: Based on the current operator, use the weak reference object to participate in model training instead of the target tensor.

[0085] Specifically, since the weak reference object is only a reference to the target tensor and includes all the functions of the target tensor, the current operator can directly use the weak reference object to participate in model training instead of the target tensor. In other words, the weak reference object can directly participate in model training in the form of the target tensor. For example, the weak reference object can be used to participate in the calculations of relevant operators, or data can be transmitted and exchanged between communication operators. The embodiments of the present invention do not make specific limitations in this regard.

[0086] Based on any of the above embodiments, in step 130, the automatically putting the target tensor back into the tensor queue after use includes:

[0087] Step 133: When it is detected that the life cycle of the weak reference object ends, automatically put the target tensor back into the tensor queue and recycle the weak reference object.

[0088] Specifically, when it is detected that the life cycle of a weakly-referenced object ends, it means that the weakly-referenced object is no longer held by any strong reference. At this time, the garbage collector can freely recycle the object. Here, detecting whether the life cycle of the weakly-referenced object ends usually depends on the behavior of the garbage collector. For example, in the Python programming language, the garbage collector will automatically track the reference situation of the weakly-referenced object. When it detects that the reference count of the weakly-referenced object reaches zero, it can determine that the life cycle of the weakly-referenced object ends.

[0089] Further, in step 133, before automatically putting the target tensor back into the tensor queue, it further includes:

[0090] Step 1331, check whether the weakly-referenced object has been registered in the tensor pool;

[0091] Step 1332, if the weakly-referenced object has been registered in the tensor pool, unregister the weakly-referenced object from the tensor pool.

[0092] Specifically, when the life cycle of a weakly-referenced object ends (for example, when no variable references it anymore), the Python garbage collector will trigger the destruction of the weakly-referenced object. Here, destruction refers to the process of object destruction. Before destroying the weakly-referenced object, it will first check whether the weakly-referenced object has been registered in the tensor pool. If it has been registered, an unregistration operation will be performed; otherwise, it can be skipped.

[0093] It can be understood that a list or hash table for registering weakly-referenced objects can be maintained in the tensor pool to track which objects have been registered. When creating each weakly-referenced object, the id of the weakly-referenced object can be registered in the corresponding list or hash table. When it is detected that the life cycle of a certain weakly-referenced object ends, it can be determined whether the weakly-referenced object has been registered in the tensor pool by querying this list or hash table.

[0094] When it is queried that the weakly-referenced object has been registered in the tensor pool, the weakly-referenced object can be unregistered from the tensor pool. Here, unregistration means removing the weakly-referenced object from the registration list maintained by the tensor pool. It should be understood that in the tensor pool, each weakly-referenced object will be registered so that the tensor pool can track and manage its life cycle; when the weakly-referenced object is no longer needed or has been used up, unregistration is required. After unregistration, the tensor pool can recycle the memory resources occupied by the weakly-referenced object to allocate these resources to other tensors in need, thereby improving the utilization rate and performance of memory resources.

[0095] After unregistering the weakly-referenced object, the corresponding target tensor can be obtained according to the original tensor attributes of the weakly-referenced object, and the target tensor can be detached from the computational graph, and then the tensor pool can recycle the target tensor. It should be understood that during the model training process based on the PyTorch framework, tensors are all within its entire computational graph. If a certain tensor needs to be recycled, it must first be ensured that the tensor is detached from the computational graph before it can be recycled, otherwise an exception will be introduced.

[0096] Finally, after completing the recycling of the target tensor, the weakly-referenced object can be recycled. Here, the recycling of the weakly-referenced object is automatically completed by the garbage collector. When the garbage collector detects that an object has no strong references, it will mark the object as recyclable and recycle its memory space at an appropriate time. For weakly-referenced objects, since they do not increase the reference count of the object, when all strong references are removed, the garbage collector can immediately recycle the object.

[0097] Based on any of the above embodiments, in step 133, the automatically putting the target tensor back into the tensor queue includes:

[0098] Step 1333, obtaining the target tensor based on the original tensor attributes of the weakly-referenced object;

[0099] Step 1334, based on the attribute information of the target tensor, checking whether there is a corresponding tensor queue in the tensor pool.

[0100] Specifically, after unregistering the weakly-referenced object, the recycling of the target tensor can be started. Since the target tensor has been previously set as the original tensor attributes of the weakly-referenced object, therefore, by obtaining the original tensor attributes of the weakly-referenced object, the target tensor can be obtained, and then the target tensor is automatically put back into the tensor pool for subsequent reuse. Here, when automatically putting the target tensor back into the tensor pool, a corresponding key can be generated according to the attribute information of the target tensor, and this key can be used to retrieve in the tensor pool to see if a matching tensor queue can be found.

[0101] Step 1335, when it is found that there is a corresponding tensor queue in the tensor pool, adding the target tensor to the tail of the tensor queue.

[0102] Specifically, if it is found that there is a corresponding tensor queue in the tensor pool, it indicates that tensors with the same attributes have been added to this queue before. In this case, the target tensor can be automatically added to the tail of the tensor queue so that in the next iteration, the communication operator that calls the same position of the tensor queue will obtain the target tensor with the same memory address, thus playing a role in accelerating performance.

[0103] Step 1336, in the case that no corresponding tensor queue is found in the tensor pool, based on the attribute information of the target tensor, create a new tensor queue in the tensor pool, and add the target tensor to the new tensor queue.

[0104] Specifically, if no corresponding tensor queue exists in the tensor pool, it indicates that the target tensor is used for the first time. Therefore, a new tensor queue needs to be created to store the target tensor. Here, a new tensor queue can be created and added to the tensor pool according to the attribute information of the target tensor (i.e., data type, device type, data shape, etc.). Specifically, a unique key can be assigned to the new tensor queue in the tensor pool (calculated based on the attribute information of the target tensor), and the key is associated with the new tensor queue.

[0105] After creating the new tensor queue, the target tensor can be automatically added to the tensor queue so that the target tensor can be reused subsequently.

[0106] Based on any of the above embodiments, Figure 2 is a schematic flowchart of the tensor data structure and its resource management method provided by the present invention. As Figure 2 shown, the embodiments of the present invention introduce a new tensor data structure SmartTensor and a tensor pool TensorPool. Among them, TensorPool solves the problem that the memory addresses allocated by PyTorch's Allocator are different each time, and SmartTensor solves the problem that tensor resources cannot be automatically recycled.

[0107] It should be noted that the tensor data structure SmartTensor inherits from torch.Tensor. Therefore, SmartTensor contains all the attributes and methods of torch.Tensor, and its usage method is the same as that of torch.Tensor. When creating SmartTensor, first create a tensor object, and then return a weak reference to this tensor object as the SmartTensor object, and use this tensor as an attribute of the SmartTensor object, so as to ensure that the tensor will not be recycled by PyTorch's Allocator. In addition, this tensor object is not provided for external use, and the external calculation uses the SmartTensor object. When the SmartTensor object is used up, it will be destructed. When destructing, its corresponding tensor object will be recycled by TensorPool for creating a new SmartTensor object for the next reuse. SmartTensor isolates the management of tensors by the Allocator and can realize the self-management of tensors.

[0108] TensorPool is a data structure that inherits from Dict (i.e., dictionary) and is a container for storing tensors that need to be reused. Its key is composed of the tensor's type + device + shape (type refers to the data type, device refers to the device type, and shape refers to the data shape), and the value corresponds to a first-in-first-out queue data structure to ensure that the tensors reused in each iteration are in order.

[0109] The most important thing for TensorPool is to implement the pop and append functions. Among them, the pop function means finding the corresponding queue according to the key and returning the head tensor, and the append function means generating a key according to the input tensor attributes (i.e., type + device + shape), and then appending the tensor to the tail of the queue corresponding to the key. During the entire training process, the tensors in TensorPool are dynamically changing. They are popped out from the queue during use and put back into the queue after use. Since model training is an iterative process, the order of the tensors obtained from the tensor pool in each iteration is the same, which will make the communication operator obtain tensors with the same memory address in each iteration, thus playing a role in accelerating performance.

[0110] The process of this method is as follows:

[0111] Step S1, find the position of the communication operator in the code, and create a tensor by calling the SmartTensor.empty interface before the communication operator.

[0112] Specifically, at the required position before the communication operator, developers can manually call the SmartTensor.empty interface to create a tensor according to the needs of model training. Specifically, when calling the SmartTensor.empty interface to create a tensor, first, the type, device, and shape of the tensor required by the communication operator are transmitted as parameters to this interface. This interface will generate a corresponding key according to these attribute information of the tensor and query whether there is a matching tensor queue in the tensor pool. If it exists, the pop() function can be called to take out the target tensor from the head of the matching tensor queue; if it does not exist, the torch.empty interface can be called to create a new tensor as the target tensor.

[0113] It can be understood that by providing the SmartTensor.empty interface, the management of tensor resources can be achieved. During model training, communication operators at the same code location can obtain the target tensor from the tensor pool, so that the tensor can be reused and ensure that for the same communication operator, the memory address of the tensor obtained in each iteration is the same.

[0114] Step S2: For the target tensor obtained in Step S1, call the corresponding constructor to create a weak reference object smart_tensor of the target tensor, and at the same time register the id of smart_tensor into the tensor pool. The purpose of this step is not to increase the reference count of the original target tensor and ensure that the tensor pool manages the life cycle of the target tensor.

[0115] Step S3: Set the target tensor as an attribute of smart_tensor to ensure that the target tensor has a strong reference, thus avoiding the target tensor being reclaimed by PyTorch's Allocator.

[0116] Step S4: The weak reference object smart_tensor can participate in model training in the form of the target tensor. For example, smart_tensor can participate in the calculation of computational operators and can also perform data transmission and exchange between communication operators.

[0117] Step S5: When the weak reference object smart_tensor is no longer used, that is, when its life cycle ends (the reference count becomes zero), Python's garbage collector will actively start the recycling work, and at this time smart_tensor is destructed. Before destruction, it will first be judged whether smart_tensor has been registered in the tensor pool. If it has been registered, unregister smart_tensor, otherwise skip. After unregistration, obtain the attribute of smart_tensor to get the target tensor, and after detaching the target tensor from the computational graph, it is recycled by the tensor pool. Here, when the tensor pool recycles the target tensor, it can automatically put the target tensor into the tensor pool through the append function of the tensor pool. Finally, the relevant resources of the weak reference object smart_tensor are automatically recycled by Python's garbage collector.

[0118] It can be understood that SmartTensor provides a constructor for tensors, that is, when creating a tensor, it can obtain it from the tensor pool. When there is no matching tensor, then call the torch.empty interface to create it, and then return a weak reference object of the tensor. If the reference does not increase the reference count of the original tensor, in this way, both the functions of the original tensor can be used and the life cycle of the tensor is rewritten. When the weak reference object is destructed, it will automatically put the tensor back into the tensor pool for subsequent reuse.

[0119] The method provided by the embodiments of the present invention isolates the life cycle management of tensors by the Allocator through the introduction of SmartTensor, and realizes the custom management of the life cycle of tensors. By introducing a tensor pool, the function of providing tensors with the same memory address for each communication operator is realized, improving the performance of model training. At the same time, automatic recycling of weak reference objects and tensor resources is realized, eliminating the need for manual recycling.

[0120] The tensor management device provided by the present invention will be described below. The tensor management device described below can be mutually corresponded and referred to the tensor management method described above.

[0121] Based on any of the above embodiments, Figure 3 is a schematic structural diagram of the tensor management device provided by the present invention. As Figure 3 shown, the device includes:

[0122] A tensor lookup unit 310, configured to look up whether there is a corresponding tensor queue in the tensor pool based on the attribute information of the tensor required by the current operator, where the attribute information of each tensor in the tensor queue is the same as the attribute information of the tensor required by the current operator, and instance objects of the tensors are stored in the tensor queue;

[0123] A tensor extraction unit 320, configured to, when it is found that there is a corresponding tensor queue in the tensor pool, extract a target tensor from the tensor queue based on the storage order of each tensor in the tensor queue;

[0124] A tensor usage unit 330, configured to use the target tensor to participate in model training based on the current operator, and automatically put the target tensor back into the tensor queue after use.

[0125] The device provided by the embodiment of the present invention, by virtue of the attribute information of the tensor required by the current operator, looks up whether there is a corresponding tensor queue in the tensor pool. When the corresponding tensor queue is found, the target tensor can be directly taken out from the tensor queue according to the storage order of each tensor in the tensor queue, without creating a new tensor, which can reduce the overhead of memory allocation and release. After taking out the target tensor, the current operator can use the target tensor to participate in model training, and automatically put the target tensor back into the tensor queue after use. Since the tensors in the tensor queue are arranged in a certain storage order, for example, arranged in the first-in-first-out order, and the instance objects of each tensor are stored in the tensor queue, the memory address of the tensor can be obtained through the instance object. Therefore, for the current operator, the position order of the target tensor taken out in each iteration of model training in the tensor queue is the same. Thus, the operator using the target tensor will obtain tensors with the same memory address in each iteration, which not only realizes the reuse of tensor memory resources, but also can reduce the complexity of data transmission and synchronization, and reduce the memory access latency, thereby playing a role in performance acceleration.

[0126] Based on any of the above embodiments, the device further includes a tensor creation unit, and the tensor creation unit is used for:

[0127] In the case where no corresponding tensor queue is found in the tensor pool, create a target tensor based on the attribute information of the tensor required by the current operator.

[0128] Based on any of the above embodiments, the tensor usage unit 330 is specifically used for:

[0129] Create a weak reference object of the target tensor, register the weak reference object into the tensor pool, and set the target tensor as the original tensor attribute of the weak reference object;

[0130] Based on the current operator, use the weak reference object to replace the target tensor to participate in model training.

[0131] Based on any of the above embodiments, the device further includes a tensor recycling unit, and the tensor recycling unit is used for:

[0132] In the case where it is detected that the life cycle of the weak reference object ends, automatically put the target tensor back into the tensor queue and recycle the weak reference object.

[0133] Based on any of the above embodiments, the tensor recycling unit is used for:

[0134] Check whether the weak reference object has been registered into the tensor pool;

[0135] If the weak reference object has been registered in the tensor pool, unregister the weak reference object from the tensor pool.

[0136] Based on any of the above embodiments, the tensor recycling unit is specifically configured to:

[0137] Obtain the target tensor based on the original tensor attributes of the weak reference object;

[0138] Based on the attribute information of the target tensor, check whether there is a corresponding tensor queue in the tensor pool;

[0139] When it is found that there is a corresponding tensor queue in the tensor pool, add the target tensor to the tail of the tensor queue;

[0140] When it is not found that there is a corresponding tensor queue in the tensor pool, create a new tensor queue in the tensor pool based on the attribute information of the target tensor, and add the target tensor to the new tensor queue.

[0141] Based on any of the above embodiments, the tensor pool is a data structure for storing and recycling tensors that need to be reused during model training. Each element in the tensor pool is stored in the form of a key-value pair. The key of any element is determined based on the attribute information of the tensor, and the value of any element is a first-in-first-out queue, and at least one instance object of the tensor is stored in the queue.

[0142] Figure 4 Illustrates a schematic physical structure diagram of an electronic device, as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a tensor management method, which includes: based on the attribute information of the tensor required by the current operator, check whether there is a corresponding tensor queue in the tensor pool, the attribute information of each tensor in the tensor queue is the same as the attribute information of the tensor required by the current operator, and instance objects of the tensors are stored in the tensor queue; when it is found that there is a corresponding tensor queue in the tensor pool, take out the target tensor from the tensor queue based on the storage order of each tensor in the tensor queue; based on the current operator, use the target tensor to participate in model training, and after use, automatically put the target tensor back into the tensor queue.

[0143] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0144] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the tensor management method provided by the above-mentioned various methods. The method includes: based on the attribute information of the tensors required by the current operator, searching in the tensor pool to find whether there is a corresponding tensor queue, where the attribute information of each tensor in the tensor queue is the same as the attribute information of the tensors required by the current operator, and instance objects of the tensors are stored in the tensor queue; in the case where a corresponding tensor queue is found in the tensor pool, based on the storage order of the tensors in the tensor queue, taking out the target tensor from the tensor queue; based on the current operator, using the target tensor to participate in model training, and after use, automatically putting the target tensor back into the tensor queue.

[0145] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the tensor management method provided by the above-mentioned various methods. The method includes: based on the attribute information of the tensors required by the current operator, searching in the tensor pool to find whether there is a corresponding tensor queue, where the attribute information of each tensor in the tensor queue is the same as the attribute information of the tensors required by the current operator, and instance objects of the tensors are stored in the tensor queue; in the case where a corresponding tensor queue is found in the tensor pool, based on the storage order of the tensors in the tensor queue, taking out the target tensor from the tensor queue; based on the current operator, using the target tensor to participate in model training, and after use, automatically putting the target tensor back into the tensor queue.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the relevant technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tensor management method, characterized in that: include: Based on the attribute information of the tensor required by the current operator, searching from the tensor pool whether there is a corresponding tensor queue, the attribute information of each tensor in the tensor queue is the same as the attribute information of the tensor required by the current operator, and the tensor queue stores instance objects of each tensor; When a corresponding tensor queue is found in the tensor pool, a target tensor is taken out from the tensor queue based on the storage order of the tensors in the tensor queue; Based on the current operator, the target tensor is used to participate in model training, and after use, the target tensor is automatically put back into the tensor queue. The target tensor is used to provide the current operator with computing or data transmission operations when participating in model training.

2. The tensor management method according to claim 1, characterized in that: The method of using the target tensor to participate in model training based on the current operator also includes: If no corresponding tensor queue is found in the tensor pool, a target tensor is created based on the attribute information of the tensor required by the current operator.

3. The tensor management method according to claim 1, characterized in that: The using the target tensor to participate in model training based on the current operator includes: Create a weak reference object of the target tensor, register the weak reference object in the tensor pool, and set the target tensor as the original tensor attribute of the weak reference object; Based on the current operator, the weak reference object is used to replace the target tensor to participate in model training.

4. The tensor management method according to claim 3, characterized in that: After the use is completed, automatically putting the target tensor back into the tensor queue includes: When it is detected that the life cycle of the weak reference object has ended, the target tensor is automatically put back into the tensor queue, and the weak reference object is recycled.

5. The tensor management method according to claim 4, characterized in that: The automatically placing the target tensor back into the tensor queue also includes: Check whether the weak reference object has been registered in the tensor pool; If the weak reference object has been registered in the tensor pool, the weak reference object is deregistered from the tensor pool.

6. The tensor management method according to claim 4, characterized in that: The automatically placing the target tensor back into the tensor queue includes: Based on the original tensor attribute of the weak reference object, obtain the target tensor; Based on the attribute information of the target tensor, searching from the tensor pool whether there is a corresponding tensor queue; When a corresponding tensor queue is found in the tensor pool, the target tensor is added to the tail of the tensor queue; If no corresponding tensor queue is found in the tensor pool, a new tensor queue is created in the tensor pool based on the attribute information of the target tensor, and the target tensor is added to the new tensor queue.

7. The tensor management method according to any one of claims 1 to 6, characterized in that: The tensor pool is a data structure used to store and recycle tensors that need to be reused during model training. Each element in the tensor pool is stored in the form of key-value pairs. The key of any element is determined based on the attribute information of the tensor. The value of any element is a first-in-first-out queue, and the queue stores at least one tensor instance object.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the tensor management method according to any one of claims 1 to 7 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the tensor management method according to any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the tensor management method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Operator processing method and related device

    CN119225918A

  • Method and device for multiplexing memory of neural network chip

    CN119512724A