Data operation method and device, electronic equipment and storage medium
By using TenosrInfor variables and CastHelper instances in the intelligent question and answer model for memory operation optimization, the problem of low operator operation performance is solved and more efficient computing performance is achieved.
Patent Information
- Application Number
- CN202510231448.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
During the training and calculation process of intelligent question-answer models, the operator's operation performance is poor due to frequent memory operations.
By obtaining the target operator and its tensor parameters, creating TenosrInfor variables and CastHelper instances and transferring them to the unified computing device architecture, implicit content copying and type conversion are automatically performed to avoid explicit memory creation, copying and freeing, as well as memory development and copying.
The computing performance of the operator is improved, and the computing efficiency is improved by reducing the number and complexity of memory operations.
Smart Images

Figure CN120066732A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer science and technology, and particularly to data operation methods, devices, electronic devices, and storage media. Background Art
[0002] Intelligent question-and-answer models are usually built based on deep learning frameworks. In deep learning frameworks, a Graphics Processing Unit (GPU) chip is usually connected, and various Compute Unified Device Architecture (CUDA) kernels are used to train and calculate the models.
[0003] During the training and calculation of intelligent question-and-answer models, due to the frequent invocation of kernels for operations, a large number of memory operations are inevitably generated during the kernel operations, resulting in poor operation performance of the kernels. Summary of the Invention
[0004] This application provides data operation methods, devices, electronic devices, and storage media to improve the operation performance of kernels.
[0005] This application provides a data operation method applied to a computing device, including:
[0006] Obtain a target kernel;
[0007] Determine the tensor parameters corresponding to the target kernel;
[0008] Create a TenosrInfor variable corresponding to the tensor parameters;
[0009] Obtain the target data type corresponding to the tensor parameters;
[0010] Based on the target data type, create a CastHelper instance corresponding to the tensor parameters;
[0011] Transfer the TenosrInfor variable and the CastHelper instance to the Compute Unified Device Architecture to call the code of the Compute Unified Device Architecture to obtain the data to be operated, perform type conversion and operation on the data to be operated, and obtain the result data;
[0012] Output the result data.
[0013] This application also provides a data operation device applied to a computing device, including:
[0014] An obtaining module, configured to obtain a target kernel;
[0015] A determining module, configured to determine the tensor parameters corresponding to the target kernel;
[0016] A processing module, configured to create a TenosrInfor variable corresponding to the tensor parameter;
[0017] An acquisition module, further configured to acquire a target data type corresponding to the tensor parameter;
[0018] The processing module is further configured to create a CastHelper instance corresponding to the tensor parameter based on the target data type;
[0019] The processing module is further configured to transfer the TenosrInfor variable and the CastHelper instance to the unified computing device architecture, so as to call the code of the unified computing device architecture to obtain the data to be operated, perform type conversion and operation on the data to be operated, so as to obtain the result data;
[0020] An output module, configured to output the result data.
[0021] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above data operation methods when executing the computer program.
[0022] This application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above data operation methods are implemented.
[0023] This application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above data operation methods are implemented.
[0024] The data operation method, device, electronic device and storage medium provided by this application, by transferring the TenosrInfor variable to the unified computing device architecture, the unified computing device architecture will automatically perform implicit content copying, and can copy the TenosrInfor value on the central processing unit (Central Processing Unit, CPU for short) to the device device, avoiding explicit memory creation, copying and release. Transferring the CastHelper instance to the unified computing device architecture, the unified computing device architecture can perform type conversion in the unified computing device architecture according to the type of the tensor parameter in the CastHelper and the target data type, avoiding memory allocation and a large amount of memory copying. Therefore, the method of avoiding explicit memory creation, copying, release and avoiding memory allocation and a large amount of memory copying can improve the operation performance of the operator. Description of the Drawings
[0025] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0026] Figure 1 It is a schematic diagram of the scenario provided by the embodiment of the present application;
[0027] Figure 2 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 1 ;
[0028] Figure 3 It is an optimization schematic diagram of the target operator for example;
[0029] Figure 4 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 2 ;
[0030] Figure 5 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 3 ;
[0031] Figure 6 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 4 ;
[0032] Figure 7 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 5 ;
[0033] Figure 8 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 6 ;
[0034] Figure 9 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 7 ;
[0035] Figure 10 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 8 ;
[0036] Figure 11 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 9 ;
[0037] Figure 12 It is a schematic flowchart of the data operation method provided by the embodiment of the present application Figure 10 ;
[0038] Figure 13Schematic diagram of the data operation device provided by the embodiment of the present application;
[0039] Figure 14 Schematic diagram of the electronic device provided by the present application. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0041] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0042] The data operation method provided by the present application first obtains a target operator, then determines the tensor parameters and the target data type corresponding to the target operator, creates a TenosrInfor variable and a CastHelper instance corresponding to the tensor parameters, and then transfers the TenosrInfor variable and the CastHelper instance to the unified computing device architecture. The data to be operated is obtained by calling the code of the unified computing device architecture, and the type conversion and operation are performed on the data to be operated, and finally the result data obtained by the operation is output. Based on the data operation method provided by the present application, the operation performance of the operator can be improved by avoiding explicit memory creation, copying, release, and avoiding memory allocation and a large number of memory copies.
[0043] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0044] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the data operation method depends, the specific application environment architecture or specific hardware architecture is described herein. Refer to Figure 1 , Figure 1 Schematic diagram of the scenario provided by the embodiment of the present application, as Figure 1As shown, the computing device ① is the execution subject of this application, which can be selected as a server, a chip, or a board. The chip can be an Artificial Intelligence (AI) chip, and the AI chip includes at least one hardware processor.
[0045] Figure 2 It is a schematic flow of the data operation method provided by the embodiments of this application Figure 1 , such as Figure 2 shown, including:
[0046] S201. Obtain the target operator.
[0047] Combined with the scenario example, an operator usually refers to a mathematical or logical operator used for data operation. Operator operation includes two parts. One part is cpp code for preparatory work such as parameter parsing, and this part of the code runs on the cpu. The other part is CUDA code for operator calculation, and the data for calculation is all in the device device. The target operator is a preset operator, such as the silu operator.
[0048] S202. Determine the tensor parameters corresponding to the target operator.
[0049] Combined with the scenario example, a tensor is a mathematical concept, which is a generalization of the concepts of vectors and matrices. The target operator includes multiple tensor parameters, and the number of tensor parameters is determined by the operation type of the target operator.
[0050] S203. Create a TenosrInfor variable corresponding to the tensor parameter.
[0051] Combined with the scenario example, Figure 3 is an optimization schematic diagram of the target operator for the example. As Figure 3 shown, the TenosrInfor variable corresponding to the tensor parameter of the target operator can be determined through cpp code, and then the TenosrInfor variable is transmitted to the CUDA kernel by value passing. The CUDA kernel will automatically perform an implicit memory copy to copy the TenosrInfo value on the CPU to the device device, avoiding explicit memory creation, copying, and release.
[0052] S204. Obtain the target data type corresponding to the tensor parameter.
[0053] Combined with the scenario example, internal data types can be defined, denoted as TGDataType-t, to abstract the types in the front end. The target data type is the optimal data type corresponding to the tensor parameter. Additionally, the basic type conversion rules between data types can be defined, denoted as CommonTypeMap, which can be encapsulated using std::map. std::map is an associative container in the C++ standard library, implemented based on a red-black tree, used to store key-value pairs. The functions of std::map include quickly retrieving, inserting, and deleting elements, and its elements are automatically sorted according to the order of the keys, usually in ascending order of the keys.
[0054] S205. Create an instance of CastHelper corresponding to the tensor parameter based on the target data type.
[0055] Combined with the scenario example, CastHelper usually refers to an auxiliary tool or library in a computer, which provides the function of type conversion during programming. The CastHelper instance is used to encapsulate the basic type conversion rules, the data types of each tensor parameter, and the target data type.
[0056] S206. Transfer the TenosrInfor variable and the CastHelper instance to the unified computing device architecture to call the code of the unified computing device architecture to obtain the data to be operated on, perform type conversion and operation on the data to be operated on, so as to obtain the result data.
[0057] Combined with the scenario example, combined Figure 3 , the obtained CastHelper instance can be transferred to the CUDA kernel together. The transfer method can be by value. The CUDA kernel performs type conversion and operation on the data type of the data to be operated on according to the basic type conversion rules, the data types of each tensor parameter, and the target data type encapsulated in the CastHelper instance, so as to avoid memory allocation and memory copy.
[0058] It is worth mentioning that the target data type can also not be encapsulated into the CastHelper instance and can be transferred to CUDA separately. At this time, CUDA can select the corresponding CUDA kernel for data operation according to the target data type to achieve data operation with the target data type.
[0059] S207. Output the result data.
[0060] Combined with the scenario example, finally output the obtained result data.
[0061] The way this example transfers the TenosrInfor variable and the CastHelper instance to the CUDA kernel can improve the computing performance of the operator by avoiding explicit memory creation, copying, and release, as well as avoiding memory allocation and a large number of memory copies.
[0062] Optionally, Figure 4 is the process schematic of the data operation method provided by the embodiment of this application Figure 2 , as Figure 4 shown, the tensor parameters include input tensor parameters and output tensor parameters;
[0063] Correspondingly, S202 includes:
[0064] S401. Determine the parameters and result parameters involved in the operation according to the operation type of the target operator.
[0065] Combined with the scenario example, if the operation type of the target operator is the add operator, for example, c = a + b, then "a" and "b" are the parameters involved in the operation, and "c" is the result parameter.
[0066] S402. Determine the parameters involved in the operation as input tensor parameters and the result parameter as the output tensor parameter.
[0067] Combined with the scenario example, the parameters "a" and "b" involved in the operation can be determined as the input tensor parameters corresponding to the add operator, and the result parameter "c" can be determined as the output tensor parameter corresponding to the add operator.
[0068] Based on the method provided in this example, all the input tensor parameters and output tensor parameters corresponding to the target operator can be determined.
[0069] Optionally, Figure 5 is the process schematic of the data operation method provided by the embodiment of this application Figure 3 , as Figure 5 shown, S203 includes:
[0070] S501. Define the TenosrInfor structure.
[0071] Combined with the scenario example, TenosrInfor is tensor information, and the TenosrInfor structure is used to encapsulate the basic information of the tensor to obtain the TenosrInfor variable. For each tensor parameter, its corresponding TenosrInfor variable can be created.
[0072] S502. Based on the TenosrInfor structure, encapsulate the basic information corresponding to the tensor parameter to obtain the TenosrInfor variable corresponding to the tensor parameter, where the basic information includes shape, stride, and ndim.
[0073] Combined with the scenario example, the shape of a tensor is a tuple of integers, representing the size of the tensor in each dimension. The stride of a tensor is a tuple of integers with the same length as the shape, representing the number of elements to skip when striding across a single dimension in memory, that is, it is used to calculate the position of the next element through the memory address when accessing a specific dimension of the tensor. ndim represents the number of dimensions of the tensor.
[0074] For example, taking the silu operator as an example, the silu operator includes an input tensor parameter and an output tensor parameter. Therefore, by creating TenosrInfor variables corresponding to the input tensor parameter and the output tensor parameter respectively, they can be denoted as info-x and info-y respectively. info-x is used to encapsulate the shape, stride, and ndim corresponding to the input tensor parameter, and info-y is used to encapsulate the shape, stride, and ndim corresponding to the output tensor parameter.
[0075] Based on the method provided in this example, information such as the hape, stride, and ndim of each tensor can be encapsulated through the TenosrInfor structure to obtain the TenosrInfor variable corresponding to each tensor.
[0076] Optionally, Figure 6 is the flowchart of the data operation method provided in the embodiment of the present application Figure 4 , as Figure 6 shown, S204 includes:
[0077] S601. Define basic type conversion rules, where the basic type conversion rules include multiple preset conversion rules.
[0078] Combined with the scenario example, different rules can be configured according to different operators by maintaining a configuration file. Specifically, a corresponding fixed value can be defined for different data types through enum, as follows:
[0079] TG-Dtype-inwalid = 0, TG-Dtype-half = 1, TG-Dtype-FLOAT = 2, TG-Dtype-DOUBLE = 3, TG-Dtype-INT8 = 4, TG-Dtype-INT16 = 5, TG-Dtype-INT31 = 6, TG-Dtype-INT32 = 7, TG-Dtype-INT64 = 8, TG-Dtype-UINT8 = 9, TG-Dtype-UINT16 = 10, TG-Dtype-UINT32 = 11, TG-Dtype-UINT64 = 12, TG-Dtype-BOOL = 13, TG-Dtype-BHALF = 14。
[0080] The multiple preset conversion rules included in the basic type conversion rule CommonTypeMap are as follows:
[0081] CommonTypeMap =
[0082] {{{TG-Dtype-INT32, TG-Dtype-INT64}, TG-Dtype-INT32},
[0083] {{TG-Dtype-INT32, TG-Dtype-DOUBLE}, TG-Dtype-DOUBLE},
[0084] {{TG-Dtype-INT64, TG-Dtype-DOUBLE}, TG-Dtype-DOUBLE},
[0085] {{TG-Dtype-FLOAT, TG-Dtype-INT32}, TG-Dtype-FLOAT},
[0086] {{TG-Dtype-FLOAT, TG-Dtype-INT64}, TG-Dtype-FLOAT},
[0087] {{TG-Dtype-FLOAT, TG-Dtype-DOUBLE}, TG-Dtype-FLOAT},
[0088] {{TG-Dtype-UINT8, TG-Dtype-FLOAT}, TG-Dtype-FLOAT},
[0089] {{TG-Dtype-BOOL, TG-Dtype-INT8}, TG-Dtype-INT8},
[0090] {{TG-Dtype-BOOL, TG-Dtype-UINT8}, TG-Dtype-UINT8},
[0091] {{TG-Dtype-BOOL, TG-Dtype-FLOAT}, TG-Dtype-FLOAT},
[0092] {{TG-Dtype-BOOL, TG-Dtype-HALF}, TG-Dtype-HALF},
[0093] {{TG-Dtype-BOOL, TG-Dtype-BHALF}, TG-Dtype-BHALF},
[0094] {{TG-Dtype-BOOL, TG-Dtype-INT64}, TG-Dtype-INT64},
[0095] {{TG-Dtype-BOOL, TG-Dtype-INT32}, TG-Dtype-INT32},
[0096] {{TG-Dtype-BOOL, TG-Dtype-DOUBLE}, TG-Dtype-DOUBLE},
[0097] {{TG-Dtype-HALF, TG-Dtype-DOUBLE}, TG-Dtype-HALF},
[0098] {{TG-Dtype-HALF, TG-Dtype-FLOAT}, TG-Dtype-HALF},
[0099] {{TG-Dtype-HALF, TG-Dtype-INT32}, TG-Dtype-HALF},
[0100] {{TG-Dtype-HALF, TG-Dtype-INT64}, TG-Dtype-HALF},
[0101] {{TG-Dtype-HALF, TG-Dtype-UINT8}, TG-Dtype-HALF},
[0102] {{TG-Dtype-HALF, TG-Dtype-BOOL}, TG-Dtype-HALF},
[0103] {{TG-Dtype-BHALF, TG-Dtype-DOUBLE}, TG-Dtype-BHALF},
[0104] {{TG-Dtype-BHALF, TG-Dtype-FLOAT}, TG-Dtype-BHALF},
[0105] {{TG-Dtype-BHALF, TG-Dtype-INT32}, TG-Dtype-BHALF},
[0106] {{TG-Dtype-BHALF, TG-Dtype-INT64}, TG-Dtype-BHALF},
[0107] {{TG-Dtype-BHALF, TG-Dtype-UINT8}, TG-Dtype-BHALF},
[0108] {{TG-Dtype-INT8, TG-Dtype-FLOAT}, TG-Dtype-FLOAT},
[0109] {{TG-Dtype-FLOAT, TG-Dtype-INT8}, TG-Dtype-FLOAT}}
[0110] Among them, each line represents a conversion rule. For example, taking "{TG-Dtype-INT32, TG-Dtype-INT64}, TG-Dtype-INT32}" as an example, when the data type of the first tensor parameter for comparison is INT32 and the data type of the second tensor parameter is INT64, using the upward compatibility approach, the data types of the first tensor parameter and the second tensor parameter can be determined as INT32.
[0111] All the conversion rules represented by the lines together constitute a rule set, and the obtained rule set can be determined as the basic type conversion rule CommonTypeMap.
[0112] S602, determine the data type corresponding to the input tensor parameter and the data type corresponding to the output tensor parameter.
[0113] Combined with the scenario example, taking the above add operator as an example, "a" and "b" are the input tensor parameters corresponding to the add operator, and "c" is the output tensor parameter corresponding to the add operator. Determine the data types corresponding to the two input tensor parameters "a" and "b", and determine the data type corresponding to the output tensor parameter "c".
[0114] S603. Adopt the principle of pairwise comparison, traverse the tensor parameters based on multiple preset conversion rules to obtain the target data type.
[0115] Combined with the scenario example, for the three tensor parameters of the add operator, based on the above basic type conversion rules, use the polling pairwise comparison method to determine the optimal target data type corresponding to the target operator, and the optimal target data type corresponding to the target operator can be denoted as common-type. Specifically, compare the data types corresponding to "a" and "b" in turn, and compare the data types corresponding to "b" and "c" to obtain the common-type corresponding to the three tensor parameters of the add operator.
[0116] Based on the method provided in this example, the optimal target data type corresponding to the target operator can be determined.
[0117] Optionally, Figure 7 is the flowchart of the data operation method provided by the embodiment of the present application Figure 5 , as Figure 7 shown, S205 includes:
[0118] S701. Obtain the data type conversion method.
[0119] Combined with the scenario example, the data type conversion method is used to convert the data type of the tensor parameter into the optimal target data type common-type.
[0120] S702. Define the CastHelper module.
[0121] Combined with the scenario example, CastHelper usually refers to an auxiliary tool or library used to provide type conversion functions during programming.
[0122] S703. Package the data type conversion method, the data type corresponding to the tensor parameter, the size of the data type, and the target data type into the CastHelper module to obtain a CastHelper instance.
[0123] Combined with the scenario example, loop through all tensor parameters in the target operator to obtain the data types corresponding to all tensor parameters respectively, call the method PushTypeAndSize of CastHelper, and package the data types of all tensor parameters, the size of each data type, the data type conversion method, and the optimal target data type common-type into the CastHelper module together.
[0124] Based on the method provided in this example, a CastHelper instance corresponding to the target operator can be obtained.
[0125] Optionally, Figure 8 is a flowchart of the data operation method provided by the embodiment of the present application Figure 6 , as Figure 8 shown, S206 includes:
[0126] S801. Invoke the code of the unified computing device architecture to obtain the data to be operated based on the TenosrInfor variable and the CastHelper instance.
[0127] Combined with the scenario example, invoke the CUDA code to start the CUDA core. The CUDA core determines the corresponding data to be operated according to the received TenosrInfor variable and CastHelper instance.
[0128] S802. Convert the data to be operated into the target data type.
[0129] Combined with the scenario example, the CUDA core can determine the optimal target data type common-type corresponding to the target operator according to the CastHelper instance, and at the same time determine the data type of the data to be operated. The CUDA core can convert the data format of the data to be operated into common-type according to the data type conversion method encapsulated in the CastHelper instance.
[0130] S803. Operate on the data to be operated in the target data type to obtain the result data.
[0131] Combined with the scenario example, after converting the data format of the data to be operated into common-type, perform corresponding operations on the data to be operated in the common-type format based on the operation type of the target operator, and then obtain the corresponding result data. At this time, the data type of the obtained result data is also the optimal target data type common-type.
[0132] Figure 9 is a flowchart of the data operation method provided by the embodiment of the present application Figure 7 , as Figure 9 shown, S207 includes:
[0133] S901. Obtain the data type corresponding to the output tensor parameter.
[0134] Combined with the scenario example, determine the data type corresponding to the output tensor parameter in the target operator.
[0135] S902. Convert the result data into the data type corresponding to the output tensor parameter and then output it.
[0136] When outputting the result data in combination with the scenario example, it is necessary to convert the result data into the data type corresponding to the output tensor parameter in the target operator. For example, if the data type corresponding to the output tensor parameter in the target operator is int8, the data type of the result data needs to be converted to int8 before outputting.
[0137] Based on the method provided in this example, performing operations on the data to be operated based on the optimal target data type common-type can improve the operation efficiency of the operator. Converting the data type of the result data into the data type corresponding to the output tensor parameter can ensure the consistency of the data type.
[0138] Optionally, Figure 10 is the flowchart of the data operation method provided by the embodiment of the present application Figure 8 , the unified computing device architecture includes multiple kernel functions, and the CastHelper instance includes a data loading instruction;
[0139] As Figure 10 shown, in S801, obtaining the data to be operated based on the TenosrInfor variable and the CastHelper instance includes:
[0140] S1001. Determine the target data type based on the CastHelper instance.
[0141] Combined with the scenario example, since the optimal target data type common-type is encapsulated in the CastHelper instance, the corresponding optimal target data type common-type can be determined according to the CastHelper instance.
[0142] S1002. Determine the target kernel function corresponding to the target data type.
[0143] Combined with the scenario example, since the CUDA kernel is multiplexed with multiple types, the kernel function corresponding to this type can be determined according to the optimal target data type common-type, and this kernel function can be determined as the target kernel function.
[0144] S1003. Execute the target kernel function to determine the shape, stride, and ndim information in the corresponding tensor parameter based on the TenosrInfor variable.
[0145] Combined with the scenario example, since the shape, stride, and ndim information corresponding to the tensor parameter are encapsulated in the TenosrInfor variable corresponding to the tensor parameter, the shape, stride, and ndim information corresponding to the corresponding tensor parameter can be determined according to the TenosrInfor variable.
[0146] S1004. Obtain the corresponding offset based on the shape, stride, and ndim information in the tensor parameter.
[0147] Combined with the scenario example, according to the shape, stride, and ndim information corresponding to the tensor parameter, the corresponding offset can be calculated.
[0148] S1005. Obtain the data type corresponding to the input tensor parameter.
[0149] Combined with the scenario example, since the data type is encapsulated in the CastHelper instance, the data type corresponding to the input tensor parameter can be determined according to the CastHelper instance. The CastHelper instance contains the tg-types variable, and tg-types represents the data type of the tensor parameter, so CastHelper can determine the data type corresponding to the input tensor parameter from tg-types.
[0150] S1006. Determine the first memory size occupied by the data type corresponding to the input tensor parameter.
[0151] Combined with the scenario example, the CastHelper instance also contains the element-sizes variable, and element-sizes represents the memory size occupied by the data type. For example, the int type is 4 bytes, and int64 is 8 bytes. Pass the calculated offset above to the LoadData instruction of CastHelper. The function of LoadData is to obtain the data in src-ptr. The element-sizes corresponding to the data type of the input tensor parameter can be obtained through arg-index to get the first memory size corresponding to the input tensor parameter.
[0152] S1007. Determine the first memory location based on the offset and the first memory size.
[0153] Combined with the scenario example, LoadData calculates the first memory location corresponding to the input tensor parameter according to the first memory size and the offset obtained above.
[0154] S1008. Obtain the data to be operated on based on the first memory location.
[0155] Combined with the scenario example, based on the first memory location obtained above, extract the corresponding data to be operated on from the first memory location.
[0156] Figure 11 Flow schematic of the data operation method provided by the embodiment of the present application Figure 9 , such as Figure 11 shown, S902 includes:
[0157] S1101. Determine the second memory size occupied by the data type corresponding to the output tensor parameter.
[0158] In combination with the scenario example, call the LoadData instruction, and obtain the element-sizes corresponding to the data type of the output tensor parameter through the arg-index to obtain the second memory size corresponding to the input tensor parameter.
[0159] S1102. Determine the second memory location based on the offset and the second memory size.
[0160] In combination with the scenario example, LoadData calculates the second memory location corresponding to the output tensor parameter according to the obtained second memory size and the offset above.
[0161] S1103. After converting the result data into the data type corresponding to the output tensor parameter, write it to the second memory location to output the result data.
[0162] In combination with the scenario example, CastHelper can determine the data type corresponding to the output tensor parameter from tg-types. When outputting the result data, it can first convert the result data from the optimal target data type common-type into the data type corresponding to the output tensor parameter, and then write the result data converted into the data type corresponding to the output tensor parameter to the determined second memory location above. It is worth mentioning that if the optimal target data type common-type is the same as the data type corresponding to the output tensor parameter, no conversion is required.
[0163] Based on the method provided in this example, the purpose of accurately converting, running, and outputting the data type of the data to be operated on can be achieved.
[0164] Optionally, Figure 12 The flowchart of the data operation method provided in the embodiment of the present application Figure 10 , as Figure 12 shown, further includes:
[0165] S1201. Determine whether the target operator is a special operator.
[0166] In combination with the scenario example, the special operator is a preset operator, such as the Add and Silu operators.
[0167] S1202. If the target operator is a special operator, determine the target data type corresponding to the tensor parameter in the target operator based on the special type conversion rule.
[0168] Combined with the scenario example, for special operators, there are corresponding special type conversion rules. For example, for the Add operator, the special type conversion rules include the following two special type conversion rules:
[0169] {{TG-Dtype-BHALF, TG-Dtype-INT64}, TG-Dtype-BHALF};
[0170] {{TG-Dtype-HALF, TG-Dtype-FLOAT}, TG-Dtype-HALF}.
[0171] For the Silu operator, the special type conversion rules include the following one special type conversion rule:
[0172] {{TG-Dtype-HALF, TG-Dtype-FLOAT}, TG-Dtype-HALF}.
[0173] Therefore, for the Add and Silu operators, the optimal target data type corresponding to the tensor parameters in the operator can be determined based on the corresponding special type conversion rules.
[0174] S1203. If the target operator is not a special operator, then determine the target data type corresponding to the tensor parameters in the target operator based on the basic type conversion rules.
[0175] Combined with the scenario example, if the target operator is not a special operator such as Add and Silu, then determine the optimal target data type corresponding to the tensor parameters in the target operator based on the basic type conversion rules defined above.
[0176] Based on the method provided in this example, ordinary operators and special operators can be distinguished, and the optimal target data type corresponding to the tensor parameters can be determined more accurately.
[0177] This embodiment can improve the operation performance of the operator by avoiding explicit memory creation, copying, release, and avoiding memory allocation and a large number of memory copies.
[0178] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0179] The embodiment of the present application also provides a data operation device. Figure 13 For the structural schematic diagram of the data operation device provided by the embodiment of the present application, as Figure 13 shown, the data operation device includes:
[0180] An acquisition module 131, configured to acquire a target operator;
[0181] A determination module 132, configured to determine tensor parameters corresponding to the target operator;
[0182] A processing module 133, configured to create a TenosrInfor variable corresponding to the tensor parameters;
[0183] The acquisition module 131 is further configured to acquire a target data type corresponding to the tensor parameters;
[0184] The processing module 133 is further configured to create a CastHelper instance corresponding to the tensor parameters based on the target data type;
[0185] The processing module 133 is further configured to transfer the TenosrInfor variable and the CastHelper instance to a unified computing device architecture, so as to call the code of the unified computing device architecture to obtain data to be operated on, perform type conversion and operation on the data to be operated on, so as to obtain result data;
[0186] An output module 134, configured to output the result data.
[0187] Optionally, the tensor parameters include input tensor parameters and output tensor parameters;
[0188] The determination module 132 is specifically configured to determine the parameters participating in the operation and the result parameters according to the operation type of the target operator;
[0189] The determination module 132 is further specifically configured to determine the parameters participating in the operation as input tensor parameters and the result parameters as output tensor parameters.
[0190] Optionally, the processing module 133 is specifically configured to define a TenosrInfor structure;
[0191] The processing module 133 is further specifically configured to encapsulate the basic information corresponding to the tensor parameters based on the TenosrInfor structure, so as to obtain a TenosrInfor variable corresponding to the tensor parameters, where the basic information includes shape, stride, and ndim.
[0192] Optionally, the acquisition module 131 is specifically configured to define basic type conversion rules, where the basic type conversion rules include multiple preset conversion rules;
[0193] The acquisition module 131 is further specifically configured to determine the data type corresponding to the input tensor parameters and the data type corresponding to the output tensor parameters;
[0194] The acquisition module 131 is further specifically configured to traverse the tensor parameters based on multiple preset conversion rules by adopting the principle of pairwise comparison, so as to obtain the target data type.
[0195] Optionally, the processing module 133 is further specifically configured to obtain a data type conversion method;
[0196] The processing module 133 is further specifically configured to define a CastHelper module;
[0197] The processing module 133 is further specifically configured to encapsulate the data type conversion method, the data type corresponding to the tensor parameter, the size of the data type, and the target data type into the CastHelper module to obtain a CastHelper instance.
[0198] Optionally, the processing module 133 is further specifically configured to call the code of the unified computing device architecture to obtain the data to be computed based on the TenosrInfor variable and the CastHelper instance;
[0199] The processing module 133 is further specifically configured to convert the data to be computed into the target data type;
[0200] The processing module 133 is further specifically configured to compute the data to be computed of the target data type to obtain result data;
[0201] The processing module 133 is further specifically configured to obtain the data type corresponding to the output tensor parameter;
[0202] The processing module 133 is further specifically configured to convert the result data into the data type corresponding to the output tensor parameter and then output it.
[0203] Optionally, the unified computing device architecture includes multiple kernel functions, and the CastHelper instance includes a data loading instruction;
[0204] Correspondingly, the processing module 133 is further specifically configured to determine the target data type based on the CastHelper instance;
[0205] The processing module 133 is further specifically configured to determine the target kernel function corresponding to the target data type;
[0206] The processing module 133 is further specifically configured to execute the target kernel function to determine the shape, stride, and ndim information in the corresponding tensor parameter based on the TenosrInfor variable;
[0207] The processing module 133 is further specifically configured to obtain the corresponding offset based on the shape, stride, and ndim information in the tensor parameter;
[0208] The processing module 133 is further specifically configured to obtain the data type corresponding to the input tensor parameter;
[0209] The processing module 133 is further specifically configured to determine the first memory size occupied by the data type corresponding to the input tensor parameter;
[0210] The processing module 133 is further specifically configured to determine the first memory location based on the offset and the first memory size;
[0211] The processing module 133 is further specifically configured to obtain the data to be operated on based on the first memory location;
[0212] The processing module 133 is further specifically configured to determine the second memory size occupied by the data type corresponding to the output tensor parameter;
[0213] The processing module 133 is further specifically configured to determine the second memory location based on the offset and the second memory size;
[0214] The processing module 133 is further specifically configured to convert the result data into the data type corresponding to the output tensor parameter and then write it to the second memory location to output the result data.
[0215] For the description of the features in the corresponding embodiments of the data processing device, reference can be made to the relevant descriptions in the corresponding embodiments of the data processing method, which will not be elaborated here one by one.
[0216] Figure 14 This is a schematic structural diagram of the electronic device provided by this application. As Figure 14 shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus.
[0217] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned data processing method embodiment.
[0218] For the specific implementation process of the processor 501, reference can be made to the above-mentioned method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0219] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or implemented by the combination of the hardware and software modules in the processor.
[0220] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0221] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.
[0222] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above data processing method embodiments when running.
[0223] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other media that can store computer programs.
[0224] The embodiments of the present application also provide a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above data processing method embodiments are implemented.
[0225] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described data processing method embodiments.
[0226] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0227] The above has introduced in detail a data processing method provided by the present application. Specific examples are used herein to illustrate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data computing method, characterized in that: Applied to computing devices, including: Get the target operator; Determining tensor parameters corresponding to the target operator; Create a TenosrInfor variable corresponding to the tensor parameter; Get the target data type corresponding to the tensor parameter; Based on the target data type, create a CastHelper instance corresponding to the tensor parameter; The TenosrInfo variable and the CastHelper instance are transferred to the unified computing device architecture to call the code of the unified computing device architecture to obtain the data to be calculated, and the data to be calculated is converted and calculated to obtain the result data; The result data is output.
2. The method according to claim 1, characterized in that The tensor parameters include input tensor parameters and output tensor parameters; Accordingly, determining the tensor parameters corresponding to the target operator includes: Determine the parameters involved in the operation and the result parameters according to the operation type of the target operator; The parameters involved in the operation are determined as the input tensor parameters, and the result parameters are determined as the output tensor parameters.
3. The method according to claim 1, characterized in that The creating of the TenosrInfor variable corresponding to the tensor parameter comprises: Define TenosrInfor structure; Based on the TenosrInfor structure, basic information corresponding to the tensor parameter is encapsulated to obtain a TenosrInfor variable corresponding to the tensor parameter, wherein the basic information includes shape, stride and ndim.
4. The method according to claim 2, characterized in that: The obtaining the target data type corresponding to the tensor parameter includes: Defining basic type conversion rules, wherein the basic type conversion rules include multiple preset conversion rules; Determine the data type corresponding to the input tensor parameter and the data type corresponding to the output tensor parameter; The principle of pairwise comparison is adopted, and the tensor parameters are traversed based on the plurality of preset conversion rules to obtain the target data type.
5. The method according to claim 4, characterized in that The creating, based on the target data type, a CastHelper instance corresponding to the tensor parameter includes: Get the data type conversion method; Define the CastHelper module; The data type conversion method, the data type corresponding to the tensor parameter, the size of the data type, and the target data type are encapsulated into the CastHelper module to obtain the CastHelper instance.
6. The method according to claim 2, characterized in that The transmitting the TenosrInfo variable and the CastHelper instance to the unified computing device architecture to call the code of the unified computing device architecture to obtain the data to be calculated, and performing type conversion and calculation on the data to be calculated to obtain result data, including: Calling the code of the unified computing device architecture to obtain the data to be calculated based on the TenosrInfo variable and the CastHelper instance; Converting the data to be calculated into the target data type; Operate the data to be operated of the target data type to obtain result data; The outputting the result data comprises: Get the data type corresponding to the output tensor parameter; The result data is converted into the data type corresponding to the output tensor parameter and then output.
7. The method according to claim 6, characterized in that The unified computing device architecture includes a plurality of kernel functions, and the CastHelper instance includes a load data instruction; Accordingly, the obtaining of data to be calculated based on the TenosrInfor variable and the CastHelper instance includes: Determine the target data type based on the CastHelper instance; Determine a target kernel function corresponding to the target data type; Execute the target kernel function to determine shape, stride and ndim information in the corresponding tensor parameters based on the TenosrInfor variable; Based on the shape, stride and ndim information in the tensor parameter, the corresponding offset is obtained; Get the data type corresponding to the input tensor parameter; Determine a first memory size occupied by a data type corresponding to the input tensor parameter; Determining a first memory location based on the offset and the first memory size; Based on the first memory location, obtaining the data to be calculated; Accordingly, the converting the result data into a data type corresponding to the output tensor parameter and then outputting the data includes: Determine a second memory size occupied by the data type corresponding to the output tensor parameter; determining a second memory location based on the offset and the second memory size; After converting the result data into the data type corresponding to the output tensor parameter, the result data is written to the second memory location to output the result data.
8. A data computing device, characterized in that: Applied to computing devices, including: An acquisition module is used to acquire the target operator; A determination module, used to determine the tensor parameters corresponding to the target operator; A processing module, used for creating a TenosrInfor variable corresponding to the tensor parameter; The acquisition module is further used to obtain the target data type corresponding to the tensor parameter; The processing module is further used to create a CastHelper instance corresponding to the tensor parameter based on the target data type; The processing module is further used to transfer the TenosrInfo variable and the CastHelper instance to the unified computing device architecture, so as to call the code of the unified computing device architecture to obtain the data to be calculated, and perform type conversion and calculation on the data to be calculated to obtain result data; An output module is used to output the result data.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data operation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data operation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Operator processing method and device, chip, computing equipment and storage medium
CN118246497A
Tensor data processing method and device and storage medium
CN119105801A
Method for optimizing reasoning performance of neural network model and computing equipment
CN119106705A
Data processing method and system, and related device
US20240143496A1
Performing dynamic sparse computation on dense computation-efficient computing devices
US20240403618A1
Cited By
Data access method and device and readable storage medium
CN120508412A