A method, device, equipment and medium for implementing a fusion operator

By performing semantic transformation and algebraic optimization of the advanced intermediate expression fusion operators in deep learning networks, combining cache size, data flow cycling polyhedral structure and computational flow, the underlying code is generated, and efficient automatic implementation of fusion operators is solved, and the problems of unsatisfactory fusion effect and large manpower investment in the existing technology are solved.

CN119806540BActive Publication Date: 2025-06-10SHANGHAI ENFLAME INTELLIGENCE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510300598.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-10
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

When the prior art realizes operator fusion in deep learning networks, it encounters complex cyclic polyhedral structures, resulting in unsatisfactory fusion effect. The implementation of handwritten fusion operators requires huge manpower investment, making it difficult to meet the complex and rapidly evolving deep learning network needs.

Method used

By semantic conversion of the advanced intermediate expression fusion operator on the multi-level memory chip, the own semantic fusion operator is obtained, and algebraic optimization is performed to obtain the optimized fusion operator. Based on the information of the optimized fusion operator, the cache size, data flow cycles the polyhedral structure and the calculation flow are obtained, and the underlying code is generated in combination to realize the optimized fusion operator.

Benefits of technology

The efficiency and accuracy of the implementation of fusion operators is improved, and the dependence on user manual writing is reduced. The fusion operator can be automatically implemented to adapt to the complex and rapidly evolving deep learning network needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806540B_ABST
    Figure CN119806540B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and medium for implementing a fusion operator. The method includes: performing semantic conversion on a high-level intermediate expression fusion operator on a multi-level storage chip to obtain a self-owned semantic fusion operator; performing algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator; obtaining a cache size, a data flow loop polyhedron structure and a computation flow corresponding to the optimized fusion operator; combining the cache size, the data flow loop polyhedron structure and the computation flow to obtain underlying code, and implementing the optimized fusion operator by running the underlying code. By combining the obtained cache size, data flow loop polyhedron structure and computation flow through the information in the optimized fusion operator, the underlying code corresponding to the fusion operator is automatically determined, and the fusion operator is implemented by running the underlying code without manual writing by the user, thereby improving the efficiency and accuracy of implementing the fusion operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to computer technology, and in particular, to a method, device, equipment and medium for implementing a fused operator. Background Art

[0002] A deep learning network usually contains a large number of operators. An artificial intelligence compiler usually adopts the method of operator fusion to improve the operator execution efficiency. Operator fusion refers to the process of combining multiple related operators into a larger operator for execution.

[0003] Currently, there are two ways to implement operator fusion: The first way is to perform operator fusion during the process of automatically generating operator implementations. This type of operator fusion generally analyzes the loop polyhedron structure of operators with memory dependencies and performs fusion at the loop polyhedron level. When the loop polyhedron structure of the operator is relatively complex, it is difficult to perform fusion at the loop polyhedron level, and thus it is impossible to achieve an ideal operator fusion result; The second way is to separate operator fusion from operator implementation. The operator fusion process occurs at the high-level intermediate representation layer. That is, the high-level intermediate representation layer decides which operators to fuse together to generate a fused operator according to certain rules. Since the fused operator generated by the high-level intermediate representation layer only contains the semantic expression of the operator and does not contain the specific implementation of the operator, the underlying operator implementation is required to support the semantics of the upper-layer fused operator to implement the fused operator.

[0004] Although the second method described above is more flexible in implementing fused operators, it is also more dependent on the number of fused operator implementations included in the underlying layer. Only when the number of underlying fused operator implementations is rich enough can the upper layer have a richer and more powerful fusion ability, and thus the network may have a higher execution efficiency. In a deep learning network, there are various types and a large number of operators. Correspondingly, the high-level intermediate representation layer HLIR can generate a wide variety of fused operators. If the implementation of the underlying fused operators depends on manual writing, it requires a huge amount of human input. Therefore, manually writing fused operator implementations usually can only meet the needs of certain specific networks and lack generalization. Especially in the case where current deep learning networks are becoming more and more complex and evolving faster and faster, manually written operators are difficult to meet the current needs. Summary of the Invention

[0005] The embodiments of the present invention provide a method for implementing a fused operator to automatically implement a fused operator.

[0006] In a first aspect, the embodiments of the present invention provide a method for implementing a fused operator, including: performing semantic conversion on a high-level intermediate representation fused operator on a multi-level storage chip to obtain a self-semantic fused operator;

[0007] performing algebraic optimization on the self-semantic fused operator to obtain an optimized fused operator;

[0008] Obtain the cache size, data flow loop polyhedron structure, and computation flow corresponding to the optimized fusion operator;

[0009] Combine the cache size, the data flow loop polyhedron structure, and the computation flow to obtain low-level code, and implement the optimized fusion operator by running the low-level code.

[0010] In a second aspect, an embodiment of the present invention provides an apparatus for implementing a fusion operator, including: a semantic conversion module, configured to perform semantic conversion on a high-level intermediate representation fusion operator on a multi-level storage chip to obtain a self-owned semantic fusion operator;

[0011] An algebraic optimization module, configured to perform algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator;

[0012] A data flow loop polyhedron structure and computation flow acquisition module, configured to obtain the cache size, data flow loop polyhedron structure, and computation flow corresponding to the optimized fusion operator;

[0013] A fusion operator implementation module, configured to combine the cache size, the data flow loop polyhedron structure, and the computation flow to obtain low-level code, and implement the optimized fusion operator by running the low-level code.

[0014] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and running on the processor, wherein the processor implements the method as described above when executing the program.

[0015] In a fourth aspect, an embodiment of the present invention provides a storage medium storing computer-executable instructions, on which a computer program is stored, wherein the program implements the method as described above when executed by a processor.

[0016] The present invention semantically converts a high-level intermediate representation fusion operator into a self-owned semantic fusion operator, performs algebraic optimization to obtain an optimized fusion operator, combines the obtained cache size, data flow loop polyhedron structure, and computation flow based on the information in the optimized fusion operator to automatically determine the low-level code corresponding to the fusion operator, and implements the fusion operator by running the low-level code without the need for manual writing by the user, thereby improving the efficiency and accuracy of implementing the fusion operator. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of a method for implementing a fusion operator provided in Embodiment 1 of the present invention;

[0018] Figure 2It is a flowchart of a method for implementing a fusion operator provided in the first embodiment of the present invention;

[0019] Figure 3 It is a flowchart of a method for implementing a fusion operator provided in the second embodiment of the present invention;

[0020] Figure 4 It is a schematic structural diagram of a device for implementing a fusion operator provided in the third embodiment of the present invention;

[0021] Figure 5 It is a schematic structural diagram of a computer device provided in the fourth embodiment of the present invention. Specific embodiments

[0022] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.

[0023] Embodiment 1

[0024] Figure 1 It is a flowchart of a method for implementing a fusion operator provided in the first embodiment of the present invention. This embodiment is applicable to the situation of implementing a fusion operator. This method can be executed by a fusion operator implementation device, and this device can be implemented in the form of hardware and / or software. As Figure 1 shown, this method includes:

[0025] Step S101, perform semantic conversion on the high-level intermediate expression fusion operator on the multi-level storage chip to obtain a self-owned semantic fusion operator.

[0026] Optionally, performing semantic conversion on the high-level intermediate expression fusion operator on the multi-level storage chip to obtain a self-owned semantic fusion operator includes: extracting the input and output information, sub-operator information, and memory access mapping relationship of the high-level intermediate expression fusion operator on the multi-level storage chip, where the sub-operator information includes sub-operator types and computational dependency relationships between sub-operators; obtaining a self-owned semantic fusion operator based on the input and output information, sub-operator information, and memory access mapping relationship.

[0027] Optionally, extracting the input and output information, sub-operator information, and memory access mapping relationship of the high-level intermediate expression fusion operator on the multi-level storage chip includes: directly extracting the input and output information and sub-operator information from the high-level intermediate expression fusion operator; determining whether the high-level intermediate expression fusion operator contains non-one-to-one type sub-operators. If so, using the non-one-to-one type sub-operator as a typical operator, otherwise, using any one-to-one type sub-operator in the fusion operator as a typical operator; using the memory access mapping relationship of the typical operator as the memory access mapping relationship of the high-level intermediate expression fusion operator.

[0028] Specifically, in this embodiment, the well - fused high - level intermediate representation fusion operator can be obtained from the framework layer of the multi - level storage chip, such as the high - level intermediate representation layer. In this embodiment, each sub - operator in the high - level intermediate representation fusion operator will be converted into a self - defined semantic operator and then wrapped with a func operator to jointly form a self - defined semantic fusion operator. Among them, when performing semantic conversion, the specific operation is to extract the input and output information, sub - operator information including sub - operator types and computational dependency relationships between sub - operators, and memory access mapping relationships in the high - level intermediate representation fusion operator, and determine the self - defined semantic fusion operator based on the above - obtained three contents.

[0029] It should be noted that the input and output information and sub - operator information can be directly extracted from the high - level intermediate representation fusion operator, but the memory access mapping relationship needs to be obtained after processing the high - level intermediate representation fusion operator. Different from other types of operators with fixed memory access mapping relationships, the types of sub - operators in the high - level intermediate representation fusion operator are diverse and the quantity is uncertain. In this embodiment, the memory access mapping relationship of the fusion operator is determined by traversing the sub - operators and finding the typical operators in the sub - operators. Among them, the typical operators are divided into two categories: the first - type operators are one - to - one type operators, such as simple Elementwise operators; the second - type operators are non - one - to - one type operators, such as complex operators like softmax, layernorm, and dot. In this embodiment, the fusion of the first - type operator with the first - type operator, the fusion of the first - type operator with the second - type operator, and the fusion of the second - type operator with the first - type operator are supported, but the fusion of the second - type operator with the second - type operator is not supported. Therefore, in the high - level intermediate representation fusion operator, there are cases where only one - to - one type sub - operators are included, or cases where both one - to - one type sub - operators and non - one - to - one type sub - operators are included, and the number of non - one - to - one type sub - operators is only one. Among them, when only one - to - one type sub - operators are included in the high - level intermediate representation fusion operator, any one - to - one type sub - operator is taken as the typical operator; when non - one - to - one type sub - operators are included in the high - level intermediate representation fusion operator, the non - one - to - one type sub - operator is taken as the typical operator. After determining the typical operator, the memory access mapping relationship of the typical operator is extracted, and the memory access mapping relationship of the typical operator is used as the memory access mapping relationship of the high - level intermediate representation fusion operator, so as to subsequently construct the self - defined semantic fusion operator based on the memory access mapping relationship of the high - level intermediate representation fusion operator. Therefore, although the types of sub - operators in the high - level intermediate representation fusion operator are diverse, as long as the memory access mapping relationship of the typical operator is defined, the memory access mapping relationship of the fusion operator can be generated.

[0030] Step S102, perform algebraic optimization on the self - defined semantic fusion operator to obtain the optimized fusion operator.

[0031] Optionally, perform algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator, including: obtaining the typical operators in the self-owned semantic fusion operator, and determining the dimension adjustment method of the input and output information according to the typical operators, where the dimension adjustment method includes the adjustment of the number of dimensions and the adjustment of the dimension size; adjusting the dimensions of the input and output information of the self-owned semantic fusion operator according to the dimension adjustment method.

[0032] Specifically, in this embodiment, after obtaining the self-owned semantic fusion operator through semantic conversion, algebraic optimization can be performed on the self-owned semantic fusion operator to obtain an optimized fusion operator. Among them, the algebraic optimization mainly adjusts the dimensions of the input and output information of the self-owned semantic fusion operator, and the dimension adjustment method for the input and output information is mainly determined according to the type of the typical operator. For example, when the typical operator in the self-owned semantic fusion operator is a one-to-one type sub-operator, during algebraic optimization, the input and output information is merged into one dimension. For example, when the input information is [d0, d1, …, dn], it is merged into one dimension, that is, [d0 * d1 * … * dn]. Of course, the dimension adjustment method for the output information is similar, and will not be elaborated in this embodiment. After the dimension adjustment, the subsequent generated data flow loop polyhedron structure can be made simpler and it is easier to perform data flow movement optimization. When the typical operator is a non-one-to-one type sub-operator, since the adjustment of the number of dimensions and the dimension size corresponding to different non-one-to-one type sub-operators are different, when the typical operator is a non-one-to-one type operator, it can include various forms of dimension adjustment methods, which are specifically determined according to the specific form of the non-one-to-one type operator.

[0033] Step S103, obtain the cache size, data flow loop polyhedron structure, and computational flow corresponding to the optimized fusion operator.

[0034] Optionally, obtaining the cache size, data flow loop polyhedron structure, and computational flow corresponding to the optimized fusion operator includes: extracting the partitioning size and transpose information from the externally transmitted tuning parameter set, and obtaining the cache size according to the input and output information, partitioning size, and transpose information in the optimized fusion operator; obtaining the data flow loop polyhedron structure according to the cache size, input and output information in the optimized fusion operator, and memory access mapping relationship; obtaining the computational flow according to the sub-operator information in the optimized fusion operator.

[0035] Among them, as Figure 2 shown is the method flow chart implemented by the fusion operator, which mainly expands and explains step S103 in detail. Step S103 mainly includes steps S1031 - S1033. The following will specifically explain each of steps S1031 - S1033:

[0036] Step S1031 , extracting the segmentation size and transposition information from the externally transmitted tuning parameter set, and obtaining the cache size according to the input and output information, segmentation size and transposition information in the optimized fusion operator.

[0037] Among them, in this embodiment, due to the multi-level storage structure of the chip, the implementation of the operator usually needs to move data between different caches, so it is necessary to obtain the cache size allocated for moving data. In this embodiment, an externally transmitted tuning parameter set will be received, and when obtaining the cache size, it is necessary to extract the pre-set segmentation size and transposition information from the tuning parameter set. In addition, the input and output information will be extracted from the optimized fusion operator. When the segmentation size, input and output information, and transposition information are known, the cache size can be obtained by calculating according to the specified algorithm based on the above three parameters. The specific type of the specified algorithm is not limited in this embodiment. As long as the cache size can be determined based on the above known information, it is within the scope of protection of this application.

[0038] It should be noted that, when the cache size is determined, subsequent data transfer can be performed based on the determined cache size, and the calculated cache size also affects the determination of the subsequent data flow cycle polyhedron structure.

[0039] Step S1032, obtaining a data flow loop polyhedron structure according to the cache size, the input and output information in the optimized fusion operator, and the memory access mapping relationship.

[0040] Optionally, a data flow loop polyhedron structure is obtained according to the cache size, input and output information in the optimized fusion operator, and the memory access mapping relationship, including: determining the number of loop axes according to the number of dimensions of the input and output information, determining the boundaries of each loop axis according to the dimensional size of the input and output information, and determining the step size of each loop axis according to the segmentation size; constructing a loop body according to the number of loop axes, the boundaries of the loop axes, and the step size of the loop axes; determining the cache handling method according to the memory access mapping relationship, the cache size, and the variables of the loop body; and determining the data flow loop polyhedron structure based on the loop body and the cache handling method.

[0041] Specifically, in the multi-level storage structure of the chip, the cache close to the computing core is called a high-level cache, and the cache far away from the computing core is called a low-level cache. Similar to other common single operators, a data flow loop polyhedron structure can be generated based on the memory access mapping relationship and input and output information in the fusion operator, as well as the cache size determined above, so that the input of the fusion operator can be moved into the high-level cache or the output can be moved out of the high-level cache multiple times. The data flow loop polyhedron structure mainly includes a loop body and a cache handling method. An example of a data flow loop polyhedron structure is shown below:

[0042]

[0043] Among them, in the implementation manner, the number of loop axes is determined according to the number of dimensions of the input and output information, such as the number of fors; the boundaries of each loop axis are determined according to the dimensions of the input and output information, such as dim0; the step size of each loop is determined according to the partitioning size extracted from the tuning parameter set, such as length1. Based on the confirmed number of loop axes, loop axis boundaries, and loop axis step sizes, a loop body can be constructed. Additionally, after the loop body is constructed, variables of the loop body can be extracted, and the cache transfer method can be determined according to the memory access mapping relationship in the optimized fusion operator, the cache size determined above, and the variables of the extracted loop body. Thus, a data flow loop polyhedron structure is determined based on the constructed loop body and the cache transfer method.

[0044] It should be noted that the partitioning size for determining the loop axis step size in this implementation manner can be obtained from the tuning parameter set input externally, or can be determined by calculation according to the cache size of the chip and the input and output information. In practical applications, the specific acquisition method of the partitioning size can be selected according to the user's needs, and this implementation manner does not limit it.

[0045] Optionally, after obtaining the data flow loop polyhedron structure according to the cache size, input and output information, and memory access mapping relationship in the optimized fusion operator, it further includes: identifying loop invariants for each data transfer operation in the cache transfer method, and extracting the identified loop invariants that are irrelevant to the loop; additionally applying for a memory on the local cache for each data transfer operation to satisfy the calculation of the previously transferred data in synchronization when transferring data to the memory; obtaining the number of running threads on the chip, and performing parallel processing on the input and output information based on the number of running threads.

[0046] Specifically, in this embodiment, after constructing the data flow loop polyhedron structure, various general optimizations of the data flow can be performed based on the data flow loop polyhedron structure of the fusion operator to significantly improve the operator performance. When performing general data flow optimization, loop invariant extraction, double buffering optimization, and parallel optimization can be carried out. Among them, loop invariant extraction refers to identifying loop invariants for each data transfer operation in the cache transfer method and extracting the loop invariants that are independent of the loop. In this embodiment, by extracting the operations independent of the loop invariants from the loop body, the number of operation executions can be reduced. Double buffering optimization refers to applying for an additional memory on the local cache of the multi-level storage chip for each data transfer operation to enable synchronous calculation for the previously transferred data when transferring data to the memory of the multi-level storage chip. Therefore, the double buffering optimization technology can cache data transfer and data calculation separately without affecting each other. Parallel optimization refers to obtaining the number of running threads on the chip, that is, the number of processing units on the chip, and performing parallel processing on the input and output information based on the number of running threads. Through parallel optimization, the efficiency of the fusion operator implementation can be significantly improved. Of course, in this embodiment, only loop invariant extraction, double buffering optimization, and parallel optimization are used as examples, and the specific optimization methods for the general data flow optimization of the data flow loop polyhedron structure are not limited in actual applications.

[0047] Step S1033, obtain the computation flow according to the sub-operator information in the optimized fusion operator.

[0048] Optionally, obtaining the computation flow according to the sub-operator information in the optimized fusion operator includes: obtaining a predefined computation flow template and a packaged vector instruction set matching the multi-level storage chip, and obtaining the computation flow according to the computation flow template, the vector instruction set, and the sub-operator information, or obtaining a code module database matching the multi-level storage chip, and obtaining the computation flow according to the code module database and the sub-operator information, where the code module database includes handwritten computation flow code modules corresponding to each operator.

[0049] Optionally, obtaining the computation flow according to the computation flow template, the vector instruction set, and the sub-operator information includes: obtaining the information of the broadcast sub-operator in the optimized fusion operator, and screening out the target computation flow template from the computation flow template according to the information of the broadcast sub-operator, where the information of the broadcast sub-operator includes the existence situation and the type of the broadcast sub-operator; screening out the target vector instructions matching each sub-operator from the vector instruction set according to the sub-operator type in the sub-operator information; obtaining the sequence relationship of the target vector instructions according to the dependency relationship between the sub-operators in the sub-operator information; and splicing the target vector instructions in the target computation flow template according to the sequence relationship to generate the computation flow.

[0050] Optionally, obtaining a computation flow according to the code module database and sub-operator information, including: obtaining, from the code module database, target handwritten computation flow code modules corresponding to each sub-operator according to the sub-operator type in the sub-operator information; obtaining the call order of each target handwritten computation flow code module according to the dependence relationship between sub-operators in the sub-operator information; and calling each target handwritten computation flow code module according to the call order to generate a computation flow.

[0051] Specifically, in this embodiment, a computation flow corresponding to the optimized fusion operator can also be obtained, and specifically, the computation flow is determined according to the sub-operator information in the optimized fusion operator, that is, the sub-operator type and the computational dependence relationship between sub-operators. In this embodiment, there are two ways to obtain the computation flow: the first way is to automatically generate the code related to the computation flow during the code generation phase according to the predefined computation flow template and the encapsulated vector instruction set; the second way is to directly call the handwritten computation flow code module in the code module database to automatically generate the code related to the computation flow.

[0052] In a specific implementation, the above first method is applicable to fusion operators with simple computational flows, such as elementwise operators. This generation method supports the use of vector instructions for data broadcasting in the generation of computational flows. Therefore, the operators generated in this way achieve register-level fusion. Among them, in this embodiment, information about the broadcast sub-operator in the optimized fusion operator is obtained. The information of the broadcast sub-operator includes the existence situation and the type of the broadcast sub-operator, and the target computational flow template is selected from the computational flow templates according to the information of the broadcast sub-operator. Three types of computational flow templates are provided in this embodiment. The first type of computational flow template can handle fusion operators that do not perform data broadcasting or only perform scalar data broadcasting; the second type of computational flow template can handle fusion operators that need to perform [N, 1]->[N, M] type broadcasting on data; the third type of computational flow template can handle fusion operators that need to perform [1, M]->[N, M] type broadcasting on data. Such classification is to generate broadcast code at different levels of the loop polyhedron of the computational flow according to the type of broadcasting, improving the performance of the computational flow code. Therefore, when it is determined that there is no broadcast sub-operator in the fusion operator, the first type of computational flow template is directly used as the target computational flow template; when it is determined that there is a broadcast sub-operator in the fusion operator and the type of the broadcast sub-operator is [N, 1]->[N, M], the second type of computational flow template is used as the target computational flow template; when it is determined that there is a broadcast sub-operator in the fusion operator and the type of the broadcast sub-operator is [1, M]->[N, M], the third type of computational flow template is used as the target computational flow template. In addition, different vector instruction sets are encapsulated for different chips in this embodiment. Therefore, after determining the multi-level storage chip where the fusion operator is located in this embodiment, the encapsulated vector instruction set that matches the multi-level storage chip can be obtained. Among them, the vector instructions corresponding to each sub-operator are included in the vector instruction set. Therefore, in this embodiment, the target vector instructions that match each sub-operator are selected from the vector instruction set according to the sub-operator type. For example, in the fusion operator, there are sub-operator a, sub-operator b, and sub-operator c, and the target vector instruction 1 that matches sub-operator a, the target vector instruction 2 that matches sub-operator b, and the target vector instruction 3 that matches sub-operator c are selected from the vector instruction set according to the types of each sub-operator. In addition, in this embodiment, the dependency relationship between sub-operators is also obtained. For example, if the output of sub-operator a is passed to sub-operator b and the output of sub-operator b is passed to sub-operator c, it is determined that sub-operator b depends on sub-operator a and sub-operator c depends on sub-operator b. Correspondingly, the order of the above three target instructions can be obtained as: target vector instruction 1, target vector instruction 2, target vector instruction 3.Therefore, when the target computation flow template, target vector instructions, and the order relationship are determined, in the target computation flow template, target vector instruction 1, target vector instruction 2, and target vector instruction 3 can be concatenated in sequence to generate a computation flow. Of course, in this embodiment, only the case where the fusion operator includes three sub-operators is used as an example for illustration, and the specific number of sub-operators in the fusion operator is not limited in this embodiment.

[0053] In another specific implementation, the above second method is applicable to scenarios where the computation flow is complex and the automatically generated operators may not meet the requirements. After the data stream finishes transporting the currently participating data block to the local cache, the handwritten sub-operator interface is called according to the order of the sub-operators in the fusion operator. At this time, the calculation results of each step of the sub-operator are stored in the local cache. Therefore, the fusion operator generated in this way is a local cache-level fusion operator. Obviously, in terms of the efficiency of using intermediate results, since the register-level fusion operator can read the intermediate calculation results from the register to achieve faster access efficiency, however, for complex sub-operators, the performance improvement of an efficient handwritten implementation can cover the performance degradation caused by accessing the intermediate results in memory. In this embodiment, different code module databases are created for different chips. Therefore, in this embodiment, after determining the multi-level storage chip where the fusion operator is located, the code module database matching the multi-level storage chip can be obtained. Among them, the code module database includes the handwritten computation flow code modules corresponding to each operator. In this embodiment, the target handwritten computation flow code modules corresponding to each sub-operator are filtered out from the code module database according to the sub-operator type. For example, in the fusion operator, there are sub-operator a, sub-operator b, and sub-operator c, and the target handwritten computation flow code module 1 matching sub-operator a, the target handwritten computation flow code module 2 matching sub-operator b, and the target handwritten computation flow code module 3 matching sub-operator c are filtered out from the code module database according to the type of each sub-operator. In addition, in this embodiment, the dependency relationship between sub-operators is also obtained. For example, the output of sub-operator a is passed to sub-operator b, and the output of sub-operator b is passed to sub-operator c. Then it is determined that sub-operator b depends on sub-operator a, and sub-operator c depends on sub-operator b. Therefore, the call order of the above three target handwritten computation flow code modules can be obtained as follows: target handwritten computation flow code module 1, target handwritten computation flow code module 2, target handwritten computation flow code module 3. Then, the above three target handwritten computation flow code modules are called according to the call order to generate a computation flow. Of course, in this embodiment, only the case where the fusion operator includes three sub-operators is used as an example for illustration, and the specific number of sub-operators in the fusion operator is not limited in this embodiment.

[0054] Step S104: Combine the cache size, the data flow loop polyhedron structure, and the computation flow to obtain the underlying code, and implement the optimized fusion operator by running the underlying code.

[0055] Specifically, in this embodiment, after obtaining the cache size, the data flow loop polyhedron structure, and the computation flow, since the above three parts are mainly presented in the nature of code, the above three parts can be combined in a specified order to obtain the code. For example, they can be concatenated in the order of cache size, data flow loop polyhedron structure, and computation flow, so as to obtain the underlying code. Moreover, the underlying code obtained in this embodiment matches the optimized fusion operator. Therefore, the optimized fusion operator can be implemented by running the underlying code.

[0056] It should be noted that, before running the underlying code in this embodiment, the underlying code also needs to be detected. Specifically, it is to detect whether there are obvious errors in the underlying code, such as the case where the parameter value exceeds the hardware support ability, etc. Of course, only examples are given in this embodiment, and the specific scenarios for detecting errors are not limited. And in the case of detecting errors, a prompt message can be generated to facilitate the user to adjust the underlying code in time, so as to ensure the accuracy of the underlying code.

[0057] In this embodiment, by converting the semantics of the high-level intermediate representation fusion operator into a high-level intermediate representation fusion operator, and performing algebraic optimization to obtain the optimized fusion operator, based on the information in the optimized fusion operator, the obtained cache size, data flow loop polyhedron structure, and computation flow are combined to automatically determine the underlying code corresponding to the fusion operator, and the fusion operator is implemented by running the underlying code without the need for the user to manually write, thereby improving the efficiency and accuracy of implementing the fusion operator.

[0058] Embodiment 2

[0059] Figure 3 The flowchart of a method for implementing a fusion operator provided in Embodiment 2 of the present invention. This embodiment is based on the above embodiment. After obtaining the data flow loop polyhedron structure according to the cache size, the input / output information in the optimized fusion operator, and the memory access mapping relationship, it further includes: when the optimized fusion operator includes a broadcast sub-operator and the tuning parameter set includes a user broadcast selection instruction, extract the cache transfer method in the data flow loop polyhedron structure, and modify the cache transfer method to broadcast the input / output data. As Figure 3 shown, the method includes:

[0060] Step S201: Perform semantic conversion on the high-level intermediate representation fusion operator on the multi-level storage chip to obtain a self-owned semantic fusion operator.

[0061] Optionally, perform semantic transformation on the high-level intermediate representation fusion operator on the multi-level storage chip to obtain a self-owned semantic fusion operator, including: extracting the input and output information, sub-operator information, and memory access mapping relationship of the high-level intermediate representation fusion operator on the multi-level storage chip, where the sub-operator information includes sub-operator types and the computational dependency relationship between sub-operators; obtaining a self-owned semantic fusion operator based on the input and output information, sub-operator information, and memory access mapping relationship.

[0062] Optionally, extract the input and output information, sub-operator information, and memory access mapping relationship of the high-level intermediate representation fusion operator on the multi-level storage chip, including: directly extracting the input and output information and sub-operator information from the high-level intermediate representation fusion operator; determining whether the high-level intermediate representation fusion operator contains non-one-to-one type sub-operators, if so, using the non-one-to-one type sub-operator as a typical operator, otherwise, using any one-to-one type sub-operator in the fusion operator as a typical operator; using the memory access mapping relationship of the typical operator as the memory access mapping relationship of the high-level intermediate representation fusion operator.

[0063] Step S202, perform algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator.

[0064] Optionally, perform algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator, including: obtaining the typical operator in the self-owned semantic fusion operator, and determining the dimension adjustment method of the input and output information according to the typical operator, where the dimension adjustment method includes the adjustment of the number of dimensions and the adjustment of the size of dimensions; adjusting the dimension of the input and output information of the self-owned semantic fusion operator according to the dimension adjustment method.

[0065] Step S203, extract the segmentation size and transpose information from the externally transmitted tuning parameter set, and obtain the cache size according to the input and output information, segmentation size, and transpose information in the optimized fusion operator.

[0066] Step S204, obtain the data flow loop polyhedron structure according to the cache size, input and output information, and memory access mapping relationship in the optimized fusion operator.

[0067] Optionally, obtain the data flow loop polyhedron structure according to the cache size, input and output information, and memory access mapping relationship in the optimized fusion operator, including: determining the number of loop axes according to the number of dimensions of the input and output information, determining the boundary of each loop axis according to the size of the dimensions of the input and output information, and determining the step size of each loop axis according to the segmentation size; constructing a loop body according to the number of loop axes, loop axis boundaries, and loop axis step sizes; determining the cache transfer method according to the memory access mapping relationship, cache size, and variables of the loop body; determining the data flow loop polyhedron structure based on the loop body and the cache transfer method.

[0068] Step S205: When the optimized fusion operator includes a broadcast sub-operator and the tuning parameter set includes a user broadcast selection instruction, extract the cache transfer method in the data flow loop polyhedron structure, and modify the cache transfer method to broadcast the input data.

[0069] Optionally, modifying the cache transfer method to broadcast the input data includes: determining the data transfer operation in the cache transfer method, splitting the input data to be broadcast in the data transfer operation to obtain local slices; obtaining the broadcast type of the broadcast sub-operator, broadcasting the local slices according to the broadcast type to obtain broadcast data, and extracting the data required in the current loop from the broadcast data.

[0070] Specifically, in this embodiment, when the optimized fusion operator includes a broadcast operator and the tuning parameter set includes a user broadcast selection instruction, the input data is also broadcast through Direct Memory Access (DMA). Among them, for the broadcast of scalar data, it can always be considered to be completed using instructions in the computational flow. For the broadcast of vector data, if it is completed in the computational flow, data alignment also needs to be considered. When the shape to be broadcast is not an integer multiple of the shape that needs to be aligned, there may be situations where it cannot be processed or a large amount of local memory needs to be occupied. On the other hand, completing data broadcast in the computational flow will result in more vector instructions in the core computational code. When this fusion operator is a computationally intensive operator, more vector instructions will further reduce the performance of the operator. This patent proposes a set of methods for using DMA for broadcast for two common data broadcast types: the first is to broadcast from shape [N, 1] to [N, M], and the second is to broadcast from shape [1, M] to [N, M]. The splitting size of the fusion operator is denoted as tilingSize, and the loop variable of the loop polyhedron is denoted as loopIndex. Then, the data parameters to be broadcast each time can be calculated using the above four parameters (N, M, tilingSize, loopIndex).

[0071] Among them, each broadcast of data needs to be completed in three steps: the first step is to split the input data to be broadcast in the data transfer operation to obtain local slices. The calculation parameters required for the local slice slice are the slice offset and the slice size; the second step is to obtain the broadcast type of the broadcast sub-operator, broadcast the local slices according to the broadcast type to obtain broadcast data, and the calculation parameters required to perform the local slice broadcast are the shape of the data obtained by the broadcast, local buffer shape; the third step is to extract the data required in the current loop from the broadcast data, and the calculation parameters required to perform the extraction are the extraction offset, subview offset.

[0072] Step S206: Obtain a computation flow according to the sub-operator information in the optimized fusion operator.

[0073] Optionally, obtaining a computation flow according to the sub-operator information in the optimized fusion operator includes: obtaining a predefined computation flow template and a packaged vector instruction set matching the multi-level storage chip, and obtaining the computation flow according to the computation flow template, the vector instruction set, and the sub-operator information; or obtaining a code module database matching the multi-level storage chip, and obtaining the computation flow according to the code module database and the sub-operator information, where the code module database includes handwritten computation flow code modules corresponding to each operator.

[0074] Step S207: Combine the cache size, the data flow loop polyhedron structure, and the computation flow to obtain low-level code, and implement the optimized fusion operator by running the low-level code.

[0075] In this embodiment, by converting the semantics of the high-level intermediate representation fusion operator into the semantics of the self-owned fusion operator, and performing algebraic optimization to obtain the optimized fusion operator, and combining the obtained cache size, data flow loop polyhedron structure, and computation flow based on the information in the optimized fusion operator, the low-level code corresponding to the fusion operator is automatically determined, and the fusion operator is implemented by running the low-level code without the need for manual writing by the user, thereby improving the efficiency and accuracy of the implementation of the fusion operator.

[0076] Embodiment III

[0077] Figure 4 FIG. is a schematic structural diagram of a device for implementing a fusion operator provided in Embodiment III of the present invention. As Figure 4 shown, the device includes: a semantic conversion module 310, an algebraic optimization module 320, a data flow loop polyhedron structure, a computation flow acquisition module 330, and a fusion operator implementation module 340.

[0078] Among them, the semantic conversion module 310 is configured to perform semantic conversion on the high-level intermediate representation fusion operator on the multi-level storage chip to obtain a self-owned semantic fusion operator;

[0079] The algebraic optimization module 320 is configured to perform algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator;

[0080] The information acquisition module 330 is configured to obtain the cache size, the data flow loop polyhedron structure, and the computation flow corresponding to the optimized fusion operator;

[0081] The fusion operator implementation module 340 is configured to combine the cache size, the data flow loop polyhedron structure, and the computation flow to obtain low-level code, and implement the optimized fusion operator by running the low-level code.

[0082] Optionally, the semantic conversion module includes an information extraction unit for extracting input and output information, sub-operator information, and memory access mapping relationships of the high-level intermediate expression fusion operator on the multi-level storage chip, wherein the sub-operator information includes sub-operator types and computational dependencies between sub-operators;

[0083] The free semantic fusion operator acquisition unit is used to acquire the free semantic fusion operator based on input and output information, sub-operator information and memory access mapping relationship.

[0084] Optionally, the semantic conversion module includes an information extraction unit for directly extracting input and output information and sub-operator information from the high-level intermediate expression fusion operator;

[0085] Determine whether the high-level intermediate expression fusion operator contains a non-one-to-one type sub-operator. If so, take the non-one-to-one type sub-operator as a typical operator. Otherwise, take any one-to-one type sub-operator in the fusion operator as a typical operator.

[0086] The memory access mapping relationship of the typical operator is used as the memory access mapping relationship of the high-level intermediate expression fusion operator.

[0087] Optionally, the algebraic optimization module includes a dimension adjustment method acquisition unit, which is used to acquire a typical operator in the own semantic fusion operator, and determine the dimension adjustment method of the input and output information according to the typical operator, wherein the dimension adjustment method includes dimension quantity adjustment and dimension size adjustment;

[0088] The dimension adjustment unit is used to adjust the dimension of the input and output information of the own semantic fusion operator according to the dimension adjustment method.

[0089] Optionally, the information acquisition module includes a cache size acquisition unit, which is used to extract the segmentation size and transposition information from the externally input tuning parameter set, and acquire the cache size according to the input and output information, segmentation size and transposition information in the optimized fusion operator;

[0090] A data stream loop polyhedron structure acquisition unit, used for acquiring a data stream loop polyhedron structure according to a cache size, input and output information in an optimized fusion operator, and a memory access mapping relationship;

[0091] The computing flow acquisition unit is used to acquire the computing flow according to the sub-operator information in the optimized fusion operator.

[0092] Optionally, a data stream cyclic polyhedron structure acquisition unit is used to determine the number of cyclic axes according to the number of dimensions of the input and output information, determine the boundaries of each cyclic axis according to the size of the dimensions of the input and output information, and determine the step size of each cyclic axis according to the segmentation size;

[0093] Construct the loop body according to the number of loop axes, loop axis boundaries and loop axis steps;

[0094] Determine the cache transfer method according to the memory access mapping relationship, cache size, and variables of the loop body;

[0095] Determine the data flow loop polyhedron structure based on the loop body and the cache transfer method.

[0096] Optionally, the device further includes a broadcast module, which is configured to extract the cache transfer method in the data flow loop polyhedron structure when the optimized fusion operator includes a broadcast sub-operator and the tuning parameter set includes a user broadcast selection instruction;

[0097] Modify the cache transfer method to broadcast the input and output data.

[0098] Optionally, the broadcast module is further configured to determine the data transfer operations in the cache transfer method, and split the input data that needs to be broadcast in the data transfer operations to obtain local slices;

[0099] Obtain the broadcast type of the broadcast sub-operator, broadcast the local slices according to the broadcast type to obtain broadcast data, and extract the data required in the current loop from the broadcast data.

[0100] Optionally, the device further includes a general data flow optimization module, which is configured to identify loop invariants in each data transfer operation in the cache transfer method, and extract the loop invariants that are irrelevant to the loop;

[0101] Allocate an additional memory on the local cache for each data transfer operation to satisfy the synchronous calculation of the previously transferred data when transferring data to the memory;

[0102] Obtain the number of running threads on the chip, and perform parallel processing on the input and output information based on the number of running threads.

[0103] The calculation flow acquisition unit is configured to obtain a predefined calculation flow template and a packaged vector instruction set that matches the multi-level storage chip, and obtain a calculation flow according to the calculation flow template, the vector instruction set, and the sub-operator information,

[0104] Alternatively, obtain a code module database that matches the multi-level storage chip, and obtain a calculation flow according to the code module database and the sub-operator information, where the code module database includes handwritten calculation flow code modules corresponding to each operator.

[0105] Optionally, the calculation flow acquisition unit is configured to obtain the information of the broadcast sub-operator in the optimized fusion operator, and filter out the target calculation flow template from the calculation flow template according to the information of the broadcast sub-operator, where the information of the broadcast sub-operator includes the existence situation and the type of the broadcast sub-operator;

[0106] Filter out the target vector instructions matching each sub-operator from the vector instruction set according to the sub-operator type in the sub-operator information;

[0107] Obtain the sequence relationship of each target vector instruction according to the dependency relationship between sub-operators in the sub-operator information;

[0108] In the target calculation flow template, splice each target vector instruction according to the sequence relationship to generate a calculation flow.

[0109] Optionally, a calculation flow acquisition unit is used to obtain the target handwritten calculation flow code module corresponding to each sub-operator from the code module database according to the sub-operator type in the sub-operator information;

[0110] Obtain the call sequence of each target handwritten calculation flow code module according to the dependency relationship between sub-operators in the sub-operator information;

[0111] Call each target handwritten calculation flow code module according to the call sequence to generate a calculation flow.

[0112] The device for implementing the fusion operator provided by the embodiments of the present invention can execute the method for implementing the fusion operator provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0113] Embodiment 4

[0114] Figure 5 It is a schematic structural diagram of a computer device provided for Embodiment 4 of the present invention. As Figure 5 shown, the computer device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the computer device can be one or more. Figure 5 Taking one processor 610 as an example; the processor 610, the memory 620, the input device 630, and the output device 640 in the computer device can be connected through a bus or other means. Figure 5 Taking the connection through the bus as an example.

[0115] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for implementing the fusion operator in the embodiments of the present invention. The processor 610 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 620, that is, implements the above-mentioned method for implementing the fusion operator.

[0116] The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely provided with respect to the processor 610, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0117] The input device 630 may be used to receive input digital or character information, and generate key signal inputs related to user settings and function controls of the computer device. The output device 640 may include display devices such as a display screen.

[0118] Embodiment Five

[0119] Embodiment Five of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a method implemented by a fusion operator when executed by a computer processor, including:

[0120] Performing semantic conversion on the high-level intermediate expression fusion operator on the multi-level storage chip to obtain a self-owned semantic fusion operator;

[0121] Performing algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator;

[0122] Obtaining the cache size, data stream loop polyhedron structure, and computation stream corresponding to the optimized fusion operator;

[0123] Combining the cache size, data stream loop polyhedron structure, and computation stream to obtain low-level code, and implementing the optimized fusion operator by running the low-level code.

[0124] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the above method operations, and may also execute related operations in the method for implementing a fusion operator provided by any embodiment of the present invention.

[0125] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0126] It should be noted that in the embodiments of the above compilation optimization device for the loop calculation graph, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0127] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for implementing a fusion operator, characterized in that: include: Perform semantic conversion on high-level intermediate expression fusion operators on multi-level storage chips to obtain self-owned semantic fusion operators; Performing algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator; Obtaining the cache size, data flow loop polyhedron structure and calculation flow corresponding to the optimized fusion operator; Combining the cache size, the data flow loop polyhedron structure and the calculation flow to obtain underlying code, and implementing the optimized fusion operator by running the underlying code; Acquiring the computation flow corresponding to the optimized fusion operator includes: acquiring a predefined computation flow template and a packaged vector instruction set matching the multi-level storage chip, and acquiring the computation flow according to the computation flow template, the vector instruction set and sub-operator information, wherein the sub-operator information is extracted from the high-level intermediate expression fusion operator on the multi-level storage chip, and the sub-operator information includes sub-operator types and computational dependencies between sub-operators; The obtaining the calculation flow according to the calculation flow template, the vector instruction set and the sub-operator information includes: Acquire information of the broadcast sub-operator in the optimized fusion operator, and filter out a target computing flow template from the computing flow template according to the information of the broadcast sub-operator, wherein the information of the broadcast sub-operator includes existence and type of the broadcast sub-operator; Filtering target vector instructions matching each sub-operator from the vector instruction set according to the sub-operator type in the sub-operator information; Acquire the sequence relationship of each of the target vector instructions according to the dependency relationship between the sub-operators in the sub-operator information; In the target computing flow template, each of the target vector instructions is spliced ​​according to the sequence relationship to generate the computing flow.

2. The method according to claim 1, characterized in that The step of performing semantic conversion on the high-level intermediate expression fusion operator on the multi-level storage chip to obtain the own semantic fusion operator includes: Extracting input and output information, sub-operator information, and memory access mapping relationship of the high-level intermediate expression fusion operator on the multi-level storage chip, wherein the sub-operator information includes sub-operator type and computational dependency relationship between sub-operators; The own semantic fusion operator is obtained based on the input and output information, the sub-operator information and the memory access mapping relationship.

3. The method according to claim 2, characterized in that The extracting the input and output information, sub-operator information and memory access mapping relationship of the high-level intermediate expression fusion operator on the multi-level storage chip includes: Directly extracting the input and output information and the sub-operator information from the high-level intermediate expression fusion operator; Determine whether the high-level intermediate expression fusion operator contains a non-one-to-one type sub-operator. If so, take the non-one-to-one type sub-operator as a typical operator. Otherwise, take any one-to-one type sub-operator in the fusion operator as a typical operator. The memory access mapping relationship of the typical operator is used as the memory access mapping relationship of the high-level intermediate expression fusion operator.

4. The method according to claim 2, characterized in that: The step of performing algebraic optimization on the own semantic fusion operator to obtain an optimized fusion operator includes: Acquire a typical operator in the self-owned semantic fusion operator, and determine a dimension adjustment method of the input and output information according to the typical operator, wherein the dimension adjustment method includes dimension quantity adjustment and dimension size adjustment; The input and output information of the own semantic fusion operator is dimensionally adjusted according to the dimension adjustment method.

5. The method according to claim 2, characterized in that: The obtaining and optimizing the cache size, data flow loop polyhedron structure and calculation flow corresponding to the fusion operator includes: Extracting the split size and transposition information from the externally input tuning parameter set, and obtaining the cache size according to the input and output information, the split size and the transposition information in the optimized fusion operator; Acquire the data flow loop polyhedron structure according to the cache size, the optimized input and output information in the fusion operator, and the memory access mapping relationship; The calculation flow is obtained according to the sub-operator information in the optimized fusion operator.

6. The method according to claim 5, characterized in that The obtaining the data flow cycle polyhedron structure according to the cache size, the optimized input and output information in the fusion operator, and the memory access mapping relationship includes: Determine the number of circulation axes according to the number of dimensions of the input and output information, determine the boundaries of each circulation axis according to the size of the dimensions of the input and output information, and determine the step size of each circulation axis according to the segmentation size; Construct a loop body according to the number of loop axes, the loop axis boundaries and the loop axis step length; Determine a cache transfer method according to the memory access mapping relationship, the cache size and the variables of the loop body; The data flow loop polyhedron structure is determined based on the loop body and the cache handling method.

7. The method according to claim 6, characterized in that After acquiring the data flow cycle polyhedron structure according to the cache size, the optimized input and output information in the fusion operator and the memory access mapping relationship, the method further includes: When the optimized fusion operator includes a broadcast sub-operator and the tuning parameter set includes a user broadcast selection instruction, extracting the cache transport mode in the data stream cycle polyhedron structure; The cache handling method is modified to broadcast the input data.

8. The method according to claim 7, characterized in that The modifying the cache handling method to broadcast the input data includes: Determine a data transfer operation in the cache transfer mode, and divide the input data that needs to be broadcast in the data transfer operation into local slices; The broadcast type of the broadcast sub-operator is obtained, the local slice is broadcasted according to the broadcast type to obtain broadcast data, and data required in the current cycle is extracted from the broadcast data.

9. The method according to claim 6, characterized in that After acquiring the data flow cycle polyhedron structure according to the cache size, the optimized input and output information in the fusion operator and the memory access mapping relationship, the method further includes: Identifying loop invariants for each data transfer operation in the cache transfer mode, and extracting the identified loop invariants that are not related to the loop; For each of the data transfer operations, a memory is additionally applied for in the local cache, so as to meet the requirement of synchronously calculating the historically transferred data when transferring data to the memory; The number of running threads on the chip is obtained, and the input and output information is processed in parallel based on the number of running threads.

10. The method according to claim 5, characterized in that The obtaining the calculation flow according to the sub-operator information in the optimized fusion operator includes: Acquire a predefined calculation flow template and a packaged vector instruction set matching the multi-level storage chip, and acquire the calculation flow according to the calculation flow template, the vector instruction set and the sub-operator information, Alternatively, a code module database matching the multi-level storage chip is obtained, and the calculation flow is obtained according to the code module database and the sub-operator information, wherein the code module database includes a handwritten calculation flow code module corresponding to each operator.

11. The method according to claim 10, characterized in that The obtaining the calculation flow according to the code module database and the sub-operator information includes: Acquire a target handwriting calculation flow code module corresponding to each sub-operator from the code module database according to the sub-operator type in the sub-operator information; Acquire the calling order of each of the target handwritten calculation flow code modules according to the inter-sub-operator dependency relationship in the sub-operator information; Each of the target handwritten calculation flow code modules is called according to the calling sequence to generate the calculation flow.

12. A device for implementing a fusion operator, characterized in that: include: The semantic conversion module is used to perform semantic conversion on the high-level intermediate expression fusion operator on the multi-level storage chip to obtain its own semantic fusion operator; an algebraic optimization module, used for performing algebraic optimization on the self-owned semantic fusion operator to obtain an optimized fusion operator; An information acquisition module, used to acquire the cache size, data flow cycle polyhedron structure and calculation flow corresponding to the optimized fusion operator; A fusion operator implementation module is used to combine the cache size, the data flow loop polyhedron structure and the calculation flow to obtain the underlying code, and implement the optimized fusion operator by running the underlying code. The information acquisition module is further used to acquire a predefined calculation flow template and a packaged vector instruction set matching the multi-level storage chip, and acquire the calculation flow according to the calculation flow template, the vector instruction set and the sub-operator information, wherein the sub-operator information is extracted from the high-level intermediate expression fusion operator on the multi-level storage chip, and the sub-operator information includes the sub-operator type and the calculation dependency relationship between the sub-operators; The obtaining the calculation flow according to the calculation flow template, the vector instruction set and the sub-operator information includes: Acquire information of the broadcast sub-operator in the optimized fusion operator, and filter out a target computing flow template from the computing flow template according to the information of the broadcast sub-operator, wherein the information of the broadcast sub-operator includes existence and type of the broadcast sub-operator; Filtering target vector instructions matching each sub-operator from the vector instruction set according to the sub-operator type in the sub-operator information; Acquire the sequence relationship of each of the target vector instructions according to the dependency relationship between the sub-operators in the sub-operator information; In the target computing flow template, each of the target vector instructions is spliced ​​according to the sequence relationship to generate the computing flow.

13. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 11 is implemented.

14. A computer executable instruction storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Automatic operator generation method and device, equipment and medium

    CN118151906A