Llogarithm summation index operator optimization method and system for mercuric chloride heterogeneous processor
By optimizing the LogSumExp operator on Astend heterogeneous processor, using the SIMD instruction set for vectorized computing and data transfer strategies, the problems of low memory access efficiency and insufficient computing resource utilization are solved, and efficient computing and accuracy guarantees are achieved.
Patent Information
- Application Number
- CN202510539373.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
The existing LogSumExp operator has low memory access efficiency on Ascend processors, low data transfer efficiency, and insufficient computing resource utilization, so it is impossible to fully utilize the parallel computing capabilities of heterogeneous processors.
By receiving the input data of the LogSumExp operator, optimize it based on the SIMD calculation characteristics of Ascend chip, it determines the parallel computing strategy and data transfer strategy suitable for non-continuous data axis in memory, and uses the SIMD instruction set to perform vectorization operations to optimize the calculation process.
Improves calculation efficiency, reduces the risk of numerical overflow, and ensures the accuracy of calculation results.
Smart Images

Figure CN120447868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically to a logarithmic sum exponential operator optimization method and system for Ascend heterogeneous processors. Background Art
[0002] With the rapid development of artificial intelligence technology, deep learning models have been widely used in various fields. Among them, the LogSumExp operator, as an important mathematical operation, is often used in neural network models, especially in the calculation of the Softmax function.
[0003] Existing implementation methods of the LogSumExp operator face the following challenges: (1) Low memory access efficiency: When processing the reduction sum operation on non-contiguous data axes, the data is discontinuously distributed in memory, resulting in inefficient data movement; (2) Insufficient computing resource utilization: Traditional point-by-point calculation methods cannot fully utilize the parallel computing capabilities of Ascend processors, especially on heterogeneous processors. Summary of the Invention
[0004] To address the deficiencies mentioned in the above background technology, the present invention aims to provide a logarithmic sum exponential operator optimization method and system for Ascend heterogeneous processors.
[0005] In a first aspect, the objectives of the present invention can be achieved through the following technical solution: a logarithmic sum exponential operator optimization method for Ascend heterogeneous processors, the method comprising the following steps:
[0006] Receive the input data of the LogSumExp operator and optimize it based on the SIMD computing characteristics of the Ascend chip to obtain the optimized input data of the LogSumExp operator.
[0007] Based on the shape of the input data of the optimized LogSumExp operator, a parallel computing strategy and a corresponding data movement strategy suitable for the reduction and summation of non-continuous data axes in memory are determined. The input data based on the optimized LogSumExp operator is calculated to obtain the output data of the LogSumExp operator.
[0008] In combination with the first aspect, in some implementations of the first aspect, the method further includes: a calculation formula of the LogSumExp operator is as follows:
[0009]
[0010] Among them, x ij It is an element in the input tensor, and the dim parameter specifies the axis to be reduced and summed.
[0011] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: when the shape of the input data of the optimized LogSumExp operator is [c, h, w], and dim = [0], it means that the sum is reduced on the c-axis, and the shape of the output data of the LogSumExp operator is [1, h, w].
[0012] In combination with the first aspect, in some implementations of the first aspect, the method further includes: within the shape of the input data of the optimized LogSumExp operator, when h×w×size≤ub_size, moving data of h×w×size into the UnifiedBuffer at a single time, calculating the exp of the data, and then traversing each layer along the c-axis direction to calculate the exp and accumulate, and after the traversal is completed, calculating the log of the accumulated data to obtain the output result.
[0013] In combination with the first aspect, in some implementations of the first aspect, the method further includes: within the shape of the input data of the optimized LogSumExp operator, when h×w×size>ub_size, the plane is divided into botileNum groups, and each group of data is moved into UB at a time with data of ub_size. After calculating the exp of the data, each layer in the group is traversed along the c-axis direction to calculate the exp and accumulate it. After the traversal is completed, the accumulated data is logarithmized, and finally the calculation of each group of data is completed to obtain a complete result.
[0014] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: when the shape of the input data of the optimized LogSumExp operator is [n, c, h, w], and dim = [0, 2], it means that the sum is reduced on the n-axis and the h-axis, and the shape of the output data is [1, c, 1, w].
[0015] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: within the shape of the input data of the optimized LogSumExp operator, when h×w×size≤ub_size, moving data of h×w×size into the UnifiedBuffer at a single time, calculating the exp of the data, and accumulating the results along the h-axis direction through the Add iteration instruction. After completion, accumulating the c×w plane in the n-axis direction, performing a log calculation on the accumulated results, and obtaining the final calculation result.
[0016] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the process of optimizing the input data of the LogSumExp operator based on the SIMD computing characteristics of the Ascend chip includes:
[0017] Optimizing computing using the SIMD instruction set: Ascend chips provide the SIMD instruction set for performing vectorized exponential and logarithmic operations, converting mathematical operations into SIMD instructions and performing calculations on multiple data points simultaneously.
[0018] In a second aspect, to achieve the above-mentioned objectives, the present invention discloses a logarithmic sum exponential operator optimization system for Ascend heterogeneous processors, comprising:
[0019] The data optimization module receives the input data of the LogSumExp operator and optimizes the input data based on the SIMD computing characteristics of the Ascend chip to obtain the optimized input data of the LogSumExp operator.
[0020] The data processing module is used to determine the parallel computing strategy and corresponding data movement strategy suitable for the reduction and summation of non-continuous data axes in memory based on the shape of the input data of the optimized LogSumExp operator, calculate the input data based on the optimized LogSumExp operator, and obtain the output data of the LogSumExp operator.
[0021] In another aspect of the present invention, in order to achieve the above-mentioned objectives, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, it adopts the logarithmic sum exponential operator optimization method for Ascend heterogeneous processors as described above.
[0022] Beneficial effects of the present invention:
[0023] The present invention can effectively improve the computational efficiency of the logarithmic sum exponential operator, reduce the risk of numerical overflow, and ensure the accuracy of the calculation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0025] Figure 1 It is a schematic flow chart of the method of the present invention;
[0026] Figure 2 It is a schematic diagram of the first calculation process of the present invention;
[0027] Figure 3 2 is a schematic diagram of the second calculation process of the present invention;
[0028] Figure 4 Schematic diagram of the third calculation process of the present invention;
[0029] Figure 5 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0031] Example 1:
[0032] like Figure 1 As shown in FIG, a logarithmic sum exponential operator optimization method for Ascend heterogeneous processors includes the following steps:
[0033] S101: Receive input data of the LogSumExp operator, optimize the input data of the LogSumExp operator based on the SIMD computing characteristics of the Ascend chip, and obtain optimized input data of the LogSumExp operator;
[0034] LogSumExp operator basic calculation formula
[0035]
[0036] Among them, x ij It is an element in the input tensor, and the dim parameter specifies the axis to be reduced and summed.
[0037] The process of optimizing the input data of the LogSumExp operator based on the SIMD computing characteristics of the Ascend chip includes:
[0038] Optimizing Computation with the SIMD Instruction Set: Ascend chips provide an efficient SIMD instruction set for performing vectorized exponential and logarithmic operations. Converting these mathematical operations into SIMD instructions allows simultaneous calculations on multiple data points, significantly improving computational efficiency. During computations, using SIMD instructions to calculate the exp and log values of data allows processing multiple data items at once, fully leveraging the parallel computing capabilities of the Ascend chip. For example, when processing input data of shape [c, h, w], SIMD instructions calculate the exp value for all elements in each h×w plane in parallel, then accumulate them along the c-axis. Finally, the log value of the accumulated results is calculated in parallel, achieving efficient data processing.
[0039] S102: Based on the shape of the input data of the optimized LogSumExp operator, a parallel computing strategy and a corresponding data movement strategy suitable for summing the non-continuous data axes in the memory are determined, and the input data based on the optimized LogSumExp operator is calculated to obtain the output data of the LogSumExp operator.
[0040] Optimization strategy for summing the shape [c,h,w] on the c-axis
[0041] The shape of input x is [c, h, w], and dim = [0], indicating a reduced summation along the c-axis. The shape of output y is [1, h, w]. The two data along the c-axis (dim = 0) are discontinuous in memory, making it difficult to load them simultaneously for summation. Therefore, a data movement strategy is required to achieve parallel computing.
[0042] Case 1: h × w × size (dataType) ≤ ub_size
[0043] Load h×w×size(dataType) data into UnifiedBuffer(UB) at a time (one plane at a time), calculate the exp of the data (calculate h×w data at a time), then traverse each layer along the c-axis to calculate the exp and accumulate. After the traversal is completed, calculate the log of the accumulated data to get the output y. Compared with point-by-point calculation, this method speeds up by h×w times. The calculation process is as follows Figure 2 .
[0044] Case 2: h × w × size (dataType) > ub_size
[0045] At this time, it is not possible to move the plane into UB at once. The plane needs to be divided into botileNum groups. The formula is as follows:
[0046]
[0047] For each group of data after segmentation, move data of ub_size into UB at a time, calculate exp for this data, and then traverse each layer in the group along the c-axis direction, calculate exp and accumulate the results. After the traversal is completed, calculate the log of the accumulated data. Finally, the calculation of each group of data is completed to obtain the complete result y. The calculation process is as follows Figure 3 .
[0048] Optimization strategy for summing the shape [n,c,h,w] on the n-axis and h-axis
[0049] The shape of input x is [n, c, h, w], and dim = [0, 2], indicating a reduced summation along the n-axis and h-axis. The output y has a shape of [1, c, 1, w]. The data on the n-axis (dim = 0) and the h-axis (dim = 2) are discontinuous in memory, making it difficult to load the data simultaneously for summation. A data movement strategy is required to achieve parallel computation. This discussion focuses on the case where h × w × size (T) < = ub_size. The logic for processing large planes is consistent with the grouping approach described above.
[0050] For h×w×size(T)<=ub_size, move data of size h×w×size(dataType) into UB at a time (move one h×w plane at a time), calculate exp for the data (calculate h×w data at a time), then accumulate the results along the h-axis using the Add iterative instruction. After completion, accumulate the c×w planes along the n-axis, perform a logarithm calculation on the accumulated results, and obtain the final calculation result y. The calculation process is as follows Figure 4 .
[0051] Example 2: Figure 5 As shown, in order to achieve the above objectives, the present invention discloses a logarithmic sum exponential operator optimization system for Ascend heterogeneous processors, including:
[0052] The data optimization module 11 is configured to receive input data of the LogSumExp operator and optimize the input data of the LogSumExp operator based on the SIMD computing characteristics of the Ascend chip to obtain optimized input data of the LogSumExp operator.
[0053] The data processing module 12 is used to determine the parallel computing strategy and the corresponding data movement strategy suitable for the reduction and summation of non-continuous data axes in the memory based on the shape of the input data of the optimized LogSumExp operator, calculate the input data based on the optimized LogSumExp operator, and obtain the output data of the LogSumExp operator.
[0054] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0055] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which executes the above method when executed by a processor. The storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0056] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0057] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present disclosure. Various changes and improvements may be made to the present disclosure without departing from the spirit and scope of the present disclosure, and such changes and improvements shall fall within the scope of the present disclosure.
Claims
1. A logarithmic sum exponential operator optimization method for Ascend heterogeneous processors, characterized by: The method comprises the following steps: Receive the input data of the LogSumExp operator and optimize it based on the SIMD computing characteristics of the Ascend chip to obtain the optimized input data of the LogSumExp operator. Based on the shape of the input data of the optimized LogSumExp operator, a parallel computing strategy and a corresponding data movement strategy suitable for the reduction and summation of non-continuous data axes in memory are determined. The input data based on the optimized LogSumExp operator is calculated to obtain the output data of the LogSumExp operator.
2. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 1, characterized in that: The calculation formula of the LogSumExp operator is as follows: Among them, x ij It is an element in the input tensor, and the dim parameter specifies the axis to be reduced and summed.
3. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 1, characterized in that: When the shape of the input data of the optimized LogSumExp operator is [c, h, w], and dim=[0], it means that the sum is reduced on the c-axis, and the shape of the output data of the LogSumExp operator is [1, h, w].
4. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 3, characterized in that: In the shape of the input data of the optimized LogSumExp operator, when h×w×size≤ub_size, data of size h×w×size is moved into the UnifiedBuffer at a time. After the exp is calculated for the data, the exp is traversed along the c-axis for each layer and accumulated. After the traversal is completed, the accumulated data is logarithmized to obtain the output result.
5. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 4, characterized in that: In the shape of the input data of the optimized LogSumExp operator, when h×w×size>ub_size, the plane is divided into botileNum groups. Each group of data is loaded with data of ub_size into UB at a time. After the exp is calculated for the data, each layer in the group is traversed along the c-axis direction to calculate the exp and accumulate it. After the traversal is completed, the accumulated data is logarithmized. Finally, the calculation of each group of data is completed to obtain a complete result.
6. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 5, characterized in that: When the shape of the input data of the optimized LogSumExp operator is [n, c, h, w], and dim=[0, 2], it means that the sum is reduced on the n-axis and the h-axis, and the shape of the output data is [1, c, 1, w].
7. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 6, characterized in that: In the shape of the input data of the optimized LogSumExp operator, when h×w×size≤ub_size, data of size h×w×size is moved into the UnifiedBuffer at a time. After calculating the exp of the data, the results are accumulated along the h-axis direction through the Add iterative instruction. After completion, the c×w plane in the n-axis direction is accumulated, and a log calculation is performed on the accumulated results to obtain the final calculation result.
8. The logarithmic sum exponential operator optimization method for Ascend heterogeneous processors according to claim 1, characterized in that: The process of optimizing the input data of the LogSumExp operator based on the SIMD computing characteristics of the Ascend chip includes: Optimizing computing using the SIMD instruction set: Ascend chips provide the SIMD instruction set for performing vectorized exponential and logarithmic operations, converting mathematical operations into SIMD instructions and performing calculations on multiple data points simultaneously.
9. The logarithmic sum exponential operator optimization system for Ascend heterogeneous processors is characterized by: include: The data optimization module receives the input data of the LogSumExp operator and optimizes the input data based on the SIMD computing characteristics of the Ascend chip to obtain the optimized input data of the LogSumExp operator. The data processing module is used to determine the parallel computing strategy and corresponding data movement strategy suitable for the reduction and summation of non-continuous data axes in memory based on the shape of the input data of the optimized LogSumExp operator, calculate the input data based on the optimized LogSumExp operator, and obtain the output data of the LogSumExp operator.
10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the logarithmic sum exponential operator optimization method for Ascend heterogeneous processors described in any one of claims 1 to 8 is adopted.