Tensor splicing method and device, storage medium and program product

By storing data structure description information to continuous memory space during tensor splicing, the problem that the tensor data layout method does not match the hardware computing card is solved, and the computing performance is improved.

CN120256683APending Publication Date: 2025-07-04广州壁仞智能科技有限公司 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510289258.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When the prior art splicing multiple tensors into a continuous piece of memory and reorganizing the tensor using the shape + step length method, it may cause the data arrangement to not match the computing power of specific hardware, introduce performance overhead, and reduce computing power.

Method used

Store the description information of the data structure of the tensor in the data storage area of the tensor, and store it in the pre-allocated continuous memory space during splicing, ensuring that the spliced tensor has the data structure description information of the original tensor, and avoid rearrangement of data arrangement before each operation.

Benefits of technology

By maintaining the correct data arrangement of tensors, performance overhead is reduced and computing power is improved, especially in specific hardware computing cards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256683A_ABST
    Figure CN120256683A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a tensor splicing method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence, the method comprises the following steps: obtaining a plurality of tensors to be spliced, and storing description information of data structures of the tensors in data storage areas of the tensors; and storing the data of each tensor and the description information of the data structure recorded in the data storage area into a pre-allocated continuous memory space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a tensor splicing method, device, storage medium, and program product. Background Art

[0002] A tensor is a multi-dimensional array, which has a wide range of applications in the field of deep learning. In order to reduce video memory fragmentation, improve data throughput, or communication performance, multiple tensors are usually spliced into a continuous memory, and the data storage area of each tensor is a part of the continuous memory.

[0003] After splicing multiple tensors into a continuous memory, when reorganizing the tensors, the current common practice is to use the shape + stride method to represent the data layout of the tensors. Specifically, the first element of each tensor is determined from the continuous memory according to the address and offset, and then the tensors are reorganized according to the shape and stride of the tensors.

[0004] For a computing card that already uses the shape + stride to represent the tensor layout, the reorganized tensor is the correct data layout, and there will be no additional overhead when the tensor participates in the calculation. However, in actual applications, some computing cards will adopt some specific layouts to give full play to the hardware computing power and achieve higher computing capabilities.

[0005] For a computing card with a specific data layout for tensors, if the above splicing method is still used when splicing multiple tensors, when the tensor participates in the calculation, after reorganizing the tensor from the continuous memory according to the shape and stride, the reorganized tensor is organized in the shape + stride manner, and its data layout may not be the specific data layout. If you want to give full play to the hardware capabilities, when the tensor participates in each operation, you also need to readjust the data layout of the tensor to reorganize the tensor into a specific layout, which will introduce additional overhead in performance and greatly reduce the computing power. Summary of the Invention

[0006] Embodiments of the present application provide a tensor splicing method, device, storage medium, and program product, which are used to reduce the performance overhead introduced when tensors with a specific data layout participate in the calculation and improve the computing power.

[0007] On the one hand, embodiments of the present application provide a tensor splicing method, including:

[0008] Obtain multiple tensors to be concatenated, where the description information of the data structure of the tensors is stored in the data storage area of the tensors;

[0009] Store the data of each tensor and the description information of the data structure recorded in the data storage area into a pre-allocated continuous memory space.

[0010] On the one hand, an embodiment of the present application provides a tensor concatenation device, including:

[0011] An obtaining module, configured to obtain multiple tensors to be concatenated, where the description information of the data structure of the tensors is stored in the data storage area of the tensors;

[0012] A concatenation module, configured to store the data of each tensor and the description information of the data structure recorded in the data storage area into a pre-allocated continuous memory space.

[0013] Optionally, the concatenation module is specifically configured to:

[0014] Traverse each tensor in sequence. For a tensor, perform the following operations: determine the target storage space actually occupied by the tensor, create a target tensor in the continuous memory space whose storage space is greater than or equal to the target storage space, and store the data of the tensor and the description information of the data structure recorded in the data storage area into the target tensor.

[0015] Optionally, the concatenation module is specifically configured to:

[0016] Replace the description information stored in the data storage area of the target tensor with the description information stored in the data storage area of the tensor, and copy the data of the tensor to the target tensor.

[0017] Optionally, the concatenation module is specifically configured to:

[0018] At the code layer, replace the code of the description information stored in the data storage area of the target tensor with the code of the description information stored in the data storage area of the tensor.

[0019] Optionally, the concatenation module is further configured to:

[0020] Record the offset of the start address of the target tensor corresponding to each tensor to be concatenated relative to the start address of the continuous memory space.

[0021] Optionally, the pre-allocated continuous memory space is greater than or equal to the total storage space actually occupied by the multiple tensors, where the storage space actually occupied by each tensor is the storage space occupied by each tensor in the current data arrangement mode.

[0022] Optionally, the description information of the data structure of the tensor includes: the address information of the data arrangement mode of the tensor and the address information of the shape of the tensor.

[0023] On the one hand, an embodiment of the present application provides a computer device, including a memory, a processor chip, and a computer program stored on the memory and executable on the processor chip. When the processor chip executes the program, the steps of the above-mentioned tensor splicing method are implemented.

[0024] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is enabled to execute the steps of the above-mentioned tensor splicing method.

[0025] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is enabled to execute the steps of the above-mentioned tensor splicing method.

[0026] In the embodiment of the present application, the description information of the data structure of the tensor is stored in the data storage area of the tensor. When splicing multiple tensors, the data of each tensor and the description information of the data structure recorded in the data storage area can be stored in a pre-allocated continuous memory space. By storing the description information of the data structure of the tensor in the continuous memory space, it can be ensured that the spliced tensor (the tensor stored in the continuous memory space) carries the description information of the data structure of the original tensor. Even if the tensor adopts a specific data arrangement mode, when the spliced tensor participates in the calculation, the data arrangement mode of the tensor can be obtained according to the description information of the data structure, so that the tensor participates in the operation in the correct data arrangement mode, avoiding the need to rearrange the data arrangement of the tensor every time it participates in the operation, reducing the performance overhead introduced when a tensor with a specific data arrangement mode participates in the calculation, and improving the calculation ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a schematic diagram of the structure of a tensor in the related art;

[0029] Figure 2 Schematic diagram of the principle of reorganizing tensors from continuous memory in the related art;

[0030] Figure 3 Schematic diagram of the structure of the tensor provided by the embodiment of the present application;

[0031] Figure 4 Schematic flow chart of a tensor splicing method provided by the embodiment of the present application;

[0032] Figure 5 Schematic diagram of creating a target tensor in a continuous memory space provided by the embodiment of the present application;

[0033] Figure 6 Another schematic diagram of creating a target tensor in a continuous memory space provided by the embodiment of the present application;

[0034] Figure 7 Another schematic diagram of creating a target tensor in a continuous memory space provided by the embodiment of the present application;

[0035] Figure 8 Schematic flow chart of the specific implementation process of a tensor splicing method provided by the embodiment of the present application;

[0036] Figure 9 Schematic diagram of the structure of a tensor splicing device provided by the embodiment of the present application;

[0037] Figure 10 Schematic diagram of the structure of a computer device provided by the embodiment of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the technical solutions of the present application. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the technical solutions of the present application.

[0039] Some concepts involved in the embodiments of the present application are introduced below.

[0040] 1. Tensor, which is a multi-dimensional array and can push vectors and matrices to higher dimensions. Some attributes in the tensor are introduced below.

[0041] ① "shape" refers to the dimensionality of a tensor, which is a tuple of integers. These integers represent the size of the tensor along each dimension. For example, if a tensor has a shape of (3, 4, 5), it means the tensor is a three-dimensional array where the first dimension has 3 elements, the second dimension has 4 elements, and the third dimension has 5 elements.

[0042] ② "size" refers to the total number of elements in a tensor, sometimes also called the "total size" of the tensor. For example, if a tensor has a shape of (3, 4, 5), then the size of this tensor is 3 * 4 * 5 = 60, indicating that there are a total of 60 elements in the tensor.

[0043] 2. Data storage area (storage): A tensor is divided into a header information area and a data storage area. The header information area mainly stores the following information: shape, stride, data type dtype, etc. The actual data of the tensor is saved as a continuous array and stored in the data storage area. The specific structure of the tensor is as Figure 1 shown. Tensor A and Tensor B share the same data storage area. The data of both Tensor A and Tensor B are stored in the data storage area in the form of a continuous array. The data storage area also includes other contents, such as size and flag.

[0044] The following briefly introduces the design concept of the embodiments of this application:

[0045] A tensor is a multi-dimensional array that has extensive applications in the field of deep learning. To reduce video memory fragmentation, improve data throughput, or communication performance, multiple tensors are usually concatenated into a single continuous memory, and the storage of each tensor is a part of this continuous memory.

[0046] After concatenating multiple tensors into a single continuous memory, when reorganizing the tensors, the current common practice is to use the shape + stride method to represent the layout of the tensors. Specifically, the first element of each tensor is determined from the continuous memory based on the address and offset, and then the tensors are reorganized according to the shape and stride of the tensors. For example, as Figure 2As shown, tensor1 has two dimensions: dimension 1 and dimension 2. It has three elements in both dimension 1 and dimension 2. The stride for jumping from the previous element to the next element in dimension 1 is 1, and the stride for jumping from the previous element to the next element in dimension 2 is 3. Based on the address and offset in the continuous memory, the position of the first element 5 of tensor1 can be determined. Then, according to the shape (3, 3) and strides (1, 3) of tensor1, by looking up the elements at the corresponding positions with the strides, the data layout of tensor1 can be reorganized. Among them, the strides (1, 3) indicate that the stride of tensor1 in dimension 1 is 1 and the stride in dimension 2 is 3.

[0047] For a computing card that already uses shape + stride to represent the tensor layout, the reorganized tensor is the correct data layout, and there will be no additional overhead when the tensor participates in calculations. However, in practical applications, some computing cards use specific layouts to fully utilize the hardware computing power and achieve higher computing capabilities.

[0048] For a computing card with a specific data layout for tensors, if the above-mentioned splicing method is still used when splicing multiple tensors, when the tensor participates in calculations, after reorganizing the tensor from the continuous memory according to the shape and strides, the reorganized tensor is organized in the way of shape + stride, and its data layout may not be the specific data layout. If you want to utilize the hardware capabilities, when the tensor participates in each operation, you also need to readjust the data layout of the tensor to reorganize the tensor into a specific layout, which will introduce additional overhead in performance and greatly reduce the computing power.

[0049] In view of this, the embodiments of the present application provide a tensor splicing method, device, storage medium, and program product. The description information of the data structure of the tensor is stored in the data storage area of the tensor. When splicing multiple tensors, the data of each tensor and the description information of the data structure recorded in the data storage area can be stored in a pre-allocated continuous memory space. By storing the description information of the data structure of the tensor in the continuous memory space, it can be ensured that the spliced tensor (the tensor stored in the continuous memory space) carries the description information of the data structure of the original tensor. Even if the tensor uses a specific data layout, when the spliced tensor participates in calculations, the data layout of the tensor can be obtained according to the description information of the data structure, so that the tensor can participate in operations in the correct data layout, avoiding the need to rearrange the data layout of the tensor every time it participates in operations, reducing the performance overhead introduced when tensors with specific data layouts participate in calculations, and improving the computing power.

[0050] Due to the embodiments of the present application, in order to adapt to hardware devices, the data storage area of tensors is expanded. Therefore, before formally introducing the tensor splicing scheme provided by the embodiments of the present application, the structure of the tensors provided by the embodiments of the present application will be described first.

[0051] As Figure 3 shown, in the tensor provided by the embodiments of the present application, its storage stores the description information of the tensor data structure. The description information of this data structure at least includes the address pointers (date ptr) of the layout and shape. Through the corresponding interfaces, the layout and shape of the tensor can be queried and obtained according to these address pointers. When specifically storing, the storage of the tensor is stored in the device memory.

[0052] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0053] Refer to Figure 4 shown, which is the implementation flowchart of a tensor splicing method in the embodiments of the present application. The specific implementation process of this method is as follows in S401 - S402:

[0054] S401: Obtain multiple tensors to be spliced, and the description information of the data structure of the tensors is stored in the data storage area of the tensors.

[0055] When specifically implementing, the specific manner of obtaining multiple tensors to be spliced can adopt the same manner as in the related art, and the embodiments of the present application do not limit this.

[0056] Among them, as introduced above Figure 3 the description information of the data structure of the tensor at least includes: the address pointers of the layout and shape of the tensor, so that when reorganizing the tensor for calculation, the layout and shape of the tensor can be queried and obtained through the corresponding interfaces according to these address pointers. Of course, other information may also be included, such as memory architecture information, and the embodiments of the present application do not limit this.

[0057] S402: Store the data of each tensor and the description information of the data structure recorded in the data storage area into a pre - allocated continuous memory space.

[0058] In specific implementation, the pre-allocated continuous memory space can be allocated according to the total storage space actually occupied by multiple tensors to be concatenated. For example, if the allocated continuous memory space is greater than or equal to the total storage space actually occupied by the multiple tensors to be concatenated, of course, in order to save memory space, a continuous memory space with a size equal to the total storage space can be allocated. In other embodiments of the present application, the pre-allocated continuous memory space can also be directly allocated, as long as sufficient memory space is reserved. The specific allocation method of the continuous memory space in the embodiments of the present application is not limited.

[0059] Among them, the total storage space actually occupied by multiple tensors to be concatenated can be calculated by summing up the storage spaces actually occupied by the multiple tensors, and the storage space actually occupied by each tensor refers to the storage space occupied by each tensor in the current data arrangement method. For example, if the shape in the header information area of a tensor is (10, 10), and after aligning the tensor with the size in the data storage area, the actual shape of the tensor is determined to be (30, 30), then the insufficient data needs to be filled in the data storage area. After filling, the actual number of elements of the tensor is 30 * 30 = 900, that is, the storage space actually occupied by it is the storage space corresponding to 900 elements (assuming the data type of the elements is known).

[0060] It should be noted that the size of a tensor can be represented by the hardware storage space actually occupied by the tensor, or by the number of elements actually contained in the tensor + the data type of the elements, and the two can be converted to each other. In the embodiments of the present application, the hardware storage space is used for description. It should be understood that in the specific implementation of the present application, the size of the tensor can also be represented by the number of elements + the data type.

[0061] In specific implementation, after storing each tensor to be concatenated into the pre-allocated continuous memory space, the storage space occupied by each tensor to be concatenated can be released to reduce video memory fragmentation and reduce the additional storage overhead caused by the special data arrangement method of the tensors to be concatenated. When specifically storing each tensor into the continuous memory space, the data of each tensor and the description information of the data structure recorded in the data storage area need to be stored into the continuous memory space.

[0062] Specifically, when storing each tensor into the continuous memory space, each tensor is traversed in sequence. For a tensor, the following operations are performed: determine the target storage space actually occupied by a tensor, create a target tensor in the continuous memory space with a storage space greater than or equal to the target storage space, and store the data of a tensor and the description information of the data structure recorded in the data storage area into the target tensor.

[0063] For example, such as Figure 5As shown, assume that multiple tensors to be concatenated are respectively denoted as tensor A, tensor B, ……, tensor N. When storing each tensor in a continuous memory space, each tensor is traversed in sequence. First, for tensor A, determine the target storage space actually occupied by tensor A, which is assumed to be 100 megabytes (MB). Then, in the continuous memory space, create a target tensor a with a storage space greater than or equal to 100 MB, and store both the data of tensor A and the description information of the data structure recorded in the data storage area into the target tensor a. Then continue the traversal. For tensor B, determine the target storage space actually occupied by tensor B, which is assumed to be 40 megabytes (MB). Then, in the continuous memory space, create a target tensor b with a storage space greater than or equal to 40 MB, and store both the data of tensor B and the description information of the data structure recorded in the data storage area into the target tensor b. And so on, continue the traversal until the last tensor N is traversed. For tensor N, determine the target storage space actually occupied by tensor N, which is assumed to be 20 megabytes (MB). Then, in the continuous memory space, create a target tensor n with a storage space greater than or equal to 20 MB, and store both the data of tensor N and the description information of the data structure recorded in the data storage area into the target tensor n.

[0064] It should be noted that Figure 5 The method shown in Figure 6 is to create the target tensor corresponding to each tensor in sequence from front to back in the continuous memory space according to the traversal order. In other embodiments of the present application, according to the traversal order, the target tensor corresponding to each tensor can also be created at random positions in the continuous memory space. As Figure 7 shown, the target tensor corresponding to each tensor is created at random positions, and a certain space can also be left between the target tensors corresponding to different tensors. As

[0065] shown, a certain storage space is left between the target tensor a and the target tensor b.

[0065] In addition, taking tensor A as an example, after determining the target storage space actually occupied by tensor A, when creating the target tensor a with a storage space greater than or equal to the target storage space in the continuous memory space, if the size of the continuous memory space is equal to the total storage space actually occupied by all the tensors to be concatenated, then at this time, the storage space of the target tensor a should also be equal to the storage space actually occupied by tensor A. In this case, the storage space of each target tensor needs to be equal to the storage space actually occupied by the corresponding tensor to ensure that the continuous memory space can store all the tensors to be concatenated.

[0066] If the size of the continuous memory space is greater than the total storage space actually occupied by all the tensors to be concatenated, then at this time, the storage space of the target tensor a can be equal to the storage space actually occupied by tensor A, or can be greater than the storage space actually occupied by tensor A. In this case, the storage space of each target tensor can be equal to the storage space actually occupied by the corresponding tensor, or can be greater than the storage space actually occupied by the corresponding tensor. However, it is necessary to ensure that the total storage space occupied by all the target tensors is less than or equal to the continuous memory space to ensure that the continuous memory space can store all the tensors to be concatenated.

[0067] Of course, in practical applications, taking Figure 5 as an example, when traversing each tensor and creating the corresponding target tensor for each tensor in the continuous memory space, the starting address of the target tensor corresponding to each tensor to be concatenated and the offset relative to the starting address of the continuous memory space can be recorded.

[0068] Specifically, taking the offset of the starting position of the continuous memory space as 0, the starting offset and ending offset of the target tensor a can be determined according to the storage space of the allocated target tensor a. According to the creation method of the target tensors corresponding to each tensor to be concatenated, the starting offset and ending offset of the target tensor b, ……, the target tensor n can also be determined.

[0069] In the embodiments of the present application, by recording the starting address of the target tensor corresponding to each tensor to be concatenated and the offset relative to the starting address of the continuous memory space, when the concatenated tensor participates in the operation, it is convenient to split out each tensor according to this offset.

[0070] Specifically, when storing both the data of each tensor to be concatenated and the description information of the data structure recorded in the data storage area into the target tensor, the description information stored in the data storage area of each tensor is used to replace the description information stored in the data storage area of the corresponding target tensor, and the data of each tensor is copied into the corresponding target tensor.

[0071] It should be noted that in the embodiments of the present application, when using the description information stored in the data storage area of each tensor to replace the description information stored in the data storage area of the corresponding target tensor, at the code layer, the code of the description information stored in the data storage area of one tensor can be used to replace the code of the description information stored in the data storage area of the target tensor.

[0072] By replacing the code of the description information of the data structure of the tensor at the code layer, it is thus unnecessary to modify the data pointer of the description information of the data structure of the tensor at the framework layer, and the data structure information of the tensor can still be queried and obtained through the corresponding interface according to the description information of the data structure.

[0073] Specifically, the data of each tensor can be copied to the corresponding target tensor. The tensor can be modified to be one-dimensional and then copied to the storage space of the corresponding target tensor one by one. The specific data copying method in the embodiments of the present application is not specifically limited.

[0074] Generally speaking, in the embodiments of the present application, by determining the storage space actually occupied by each tensor to be concatenated in the current data arrangement, and allocating corresponding target tensors based on the storage space actually occupied by each tensor, it can ensure that the length of the data segment after concatenation is correct. By exchanging the description information of the data structure stored in the data storage area of the tensor and copying the tensor data, it can ensure that the concatenated tensor has the expected data arrangement and correct data. When the subsequent concatenated tensor participates in calculations, each tensor will have the correct data arrangement expected under its specific shape (shape), thus avoiding the performance overhead introduced by rearranging the data arrangement of the tensor every time and greatly improving the computing power.

[0075] Taking an actual application scenario as an example, in the second-generation generative pre-training model (GPT-2), the optimizer adopted the ZeRO-1 (large model video memory optimization technology) optimization scheme. This optimization scheme concatenates all the model parameters. In the original scheme, there is a lot of data rearrangement in the model, which affects the end-to-end training speed of the model. After adopting the tensor concatenation scheme provided by the embodiments of the present application, the performance of GPT-2 in one step of end-to-end training is improved by nearly 20%, which can significantly improve the training speed.

[0076] Next, in combination with Figure 8 , the specific implementation process of the tensor concatenation method provided by the embodiments of the present application will be described in detail. As Figure 8 shown, the specific implementation process of the tensor concatenation method provided by the embodiments of the present application includes:

[0077] Step 801: Obtain multiple tensors to be concatenated. Among them, the description information of the data structure of each tensor is stored in the data storage area of the tensor.

[0078] Step 802: Determine the storage space actually occupied by each tensor in the current data arrangement.

[0079] Step 803: Sum the total storage space actually occupied by the multiple tensors to obtain the total storage space actually occupied by all the tensors to be concatenated.

[0080] Step 804: Create a continuous memory space in the device memory with the same size as the total storage space.

[0081] Step 805: Traverse each tensor to be concatenated.

[0082] Step 806: Determine the storage space actually occupied by the current tensor.

[0083] Step 807: Create a target tensor in the continuous memory space with the same size as the storage space actually occupied by the current tensor.

[0084] Step 808: Record the starting address of the target tensor and the offset relative to the starting address of the continuous memory space.

[0085] Step 809: Replace the description information stored in the data storage area of the target tensor with the description information stored in the data storage area of the current tensor.

[0086] Step 810: Copy the data of the current tensor to the target tensor.

[0087] Step 811: Determine whether all the tensors to be concatenated have been traversed. If so, execute Step 812; if not, continue to execute Step 805 to traverse the next tensor to be concatenated.

[0088] Step 812: The concatenation is completed.

[0089] Based on the same technical concept, the embodiment of the present application provides a schematic structural diagram of a tensor concatenation device, as Figure 9 shown. The tensor concatenation device 900 includes:

[0090] An acquisition module 901, configured to acquire multiple tensors to be concatenated, and the description information of the data structure of the tensors is stored in the data storage area of the tensors;

[0091] A concatenation module 902, configured to store both the data of each tensor and the description information of the data structure recorded in the data storage area into a pre-allocated continuous memory space.

[0092] Optionally, the concatenation module 902 is specifically configured to:

[0093] Traverse each tensor in sequence. For a tensor, perform the following operations: determine the target storage space actually occupied by a tensor, create a target tensor in the continuous memory space with a storage space greater than or equal to the target storage space, and store both the data of a tensor and the description information of the data structure recorded in the data storage area into the target tensor.

[0094] Optionally, the concatenation module 902 is specifically configured to:

[0095] Replace the description information stored in the data storage area of the target tensor with the description information stored in the data storage area of a tensor, and copy the data of a tensor to the target tensor.

[0096] Optionally, the concatenation module 902 is specifically configured to:

[0097] At the code layer, the code storing the description information in the data storage area of a tensor is used to replace the code storing the description information in the data storage area of the target tensor.

[0098] Optionally, the splicing module 902 is further configured to:

[0099] Record the starting address of the target tensor corresponding to each tensor to be spliced, and the offset relative to the starting address of the continuous memory space.

[0100] Optionally, the pre-allocated continuous memory space is greater than or equal to the total storage space actually occupied by multiple tensors, where the storage space actually occupied by each tensor is the storage space occupied by each tensor in the current data layout.

[0101] Optionally, the description information of the data structure of the tensor includes: the address information of the data layout mode of the tensor, and the address information of the shape of the tensor.

[0102] Based on the same technical concept, an embodiment of the present application provides a computer device, as Figure 10 shown, including at least one processor chip 1001 and a memory 1002 connected to the at least one processor chip. In the embodiment of the present application, the specific connection medium between the processor chip 1001 and the memory 1002 is not limited. Figure 10 Taking the example that the processor chip 1001 and the memory 1002 are connected by a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0103] In the embodiment of the present application, the memory 1002 stores instructions executable by the at least one processor chip 1001. By executing the instructions stored in the memory 1002, the at least one processor chip 1001 can execute the steps of the above tensor splicing method.

[0104] Among them, the processor chip 1001 is the control center of the computer device, and can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 1002 and calling the data stored in the memory 1002, tensor splicing can be achieved. Optionally, the processor chip 1001 may include one or more processing units. The processor chip 1001 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor chip 1001. In some embodiments, the processor chip 1001 and the memory 1002 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.

[0105] The processor chip 1001 can be a general-purpose processor, such as a Graphics Processing Unit (GPU), a General-Purpose computing on Graphics Processing Units (GPGPU), a Central Processing Unit (CPU), a Digital Signal Processor, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0106] The memory 1002, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1002 can include at least one type of storage medium, for example, it can include flash memory, hard disk, multimedia card, card-type memory, Random Access Memory (RAM), Static Random Access Memory (SRAM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), magnetic memory, magnetic disk, optical disc, and so on. The memory 1002 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 1002 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0107] Based on the same inventive concept, the embodiments of the present application provide a computer-readable storage medium that stores a computer program executable by a computer device. When the program runs on the computer device, it causes the computer device to execute the steps of the above-mentioned tensor splicing method.

[0108] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to perform the steps of the above-mentioned tensor splicing method.

[0109] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a computer program product, or a combination thereof. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable tensor splicing devices to generate a machine, such that the instructions executed by the processor of the computer device or other programmable tensor splicing devices generate means for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0111] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable tensor splicing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0112] These computer program instructions can also be loaded onto a computer device or other programmable tensor splicing devices, such that a series of operation steps are performed on the computer device or other programmable devices to generate a process implemented by the computer device. Thus, the instructions executed on the computer device or other programmable devices provide steps for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0113] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0114] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A tensor splicing method, characterized in that, Including: Obtain a plurality of tensors to be spliced, and description information of the data structure of the tensors is stored in the data storage area of the tensors; Store the data of each tensor and the description information of the data structure recorded in the data storage area into a pre-allocated continuous memory space.

2. The method according to claim 1, wherein The storing the data of each tensor and the description information of the data structure recorded in the data storage area into a pre-allocated continuous memory space includes: Traverse each tensor in sequence. For a tensor, perform the following operations: determine the target storage space actually occupied by the tensor, create a target tensor in the continuous memory space with a storage space greater than or equal to the target storage space, and store the data of the tensor and the description information of the data structure recorded in the data storage area into the target tensor.

3. The method according to claim 2, wherein The storing the data of the tensor and the description information of the data structure recorded in the data storage area into the target tensor includes: Replace the description information stored in the data storage area of the target tensor with the description information stored in the data storage area of the tensor, and copy the data of the tensor to the target tensor.

4. The method according to claim 3, wherein The replacing the description information stored in the data storage area of the target tensor with the description information stored in the data storage area of the tensor includes: At the code layer, replace the code of the description information stored in the data storage area of the target tensor with the code of the description information stored in the data storage area of the tensor.

5. The method according to any one of claims 2-4, characterized in that, The method further includes: Record the starting address of the target tensor corresponding to each tensor to be spliced and the offset relative to the starting address of the continuous memory space.

6. The method according to any one of claims 1-4, characterized in that, The pre-allocated continuous memory space is greater than or equal to the total storage space actually occupied by the plurality of tensors, where the storage space actually occupied by each tensor is the storage space occupied by each tensor in the current data arrangement mode.

7. The method according to any one of claims 1 to 4, characterized in that, The description information of the data structure of the tensor includes: the address information of the data arrangement mode of the tensor and the address information of the shape of the tensor.

8. A computer device, comprising a memory, a processor chip, and a computer program stored on the memory and executable on the processor chip, characterized in that, When the processor chip executes the program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Tensor splicing method and device, equipment, storage medium and program product

    CN121257615A