Operator optimization method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202311193617.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-09-14
AI Technical Summary
这种固定范式缺少灵活性,无法支撑灵活多变的tiling策略,是一种非常低效的调优手段
[0029]本发明提供的算子优化方法、装置、电子设备及存储介质,基于获取到的切分策略,生成切分策略对应的第一schedule tree,其中,切分策略中的for循环语句与第一schedule tree中的节点一一对应;由于第一schedule tree是可配置的,可以支撑灵活多变的切分策略,因此利用第一schedule tree对神经网络模型中的各算子进行优化,避免了将切分策略以硬编码的形式添加到代码工程中,进而可以加速各算子的调优过程。
Smart Images

Figure CN117291259B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an operator optimization method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the rapid development of Artificial Intelligence (AI), more and more practical applications are beginning to utilize deep learning technology, such as speech recognition, machine translation, and autonomous driving. Deep learning technology has enormous potential in real-world applications, which has led to its increasing attention. The computational efficiency of the deep neural network (DNN) models used in deep learning technology directly affects the effectiveness of practical applications. For example, the computation time of algorithms such as object detection, object recognition, and motion prediction in autonomous driving determines the usability and safety of the algorithms. Therefore, how to perform AI computing with high performance is an urgent need in current AI chip development.
[0003] High-performance operators are the foundation of efficient computing in AI chips, and related technologies typically use tiling strategies to optimize operators. However, current tiling is performed using a fixed paradigm. This fixed paradigm requires expanding the support for each new tiling strategy and hard-coding it into the codebase for testing and verification. This fixed paradigm lacks flexibility and cannot support flexible and varied tiling strategies, making it a highly inefficient optimization method.
[0004] Therefore, how to optimize each operator and accelerate the tuning process is an urgent problem to be solved. Summary of the Invention
[0005] To address the problems existing in the prior art, embodiments of the present invention provide an operator optimization method, apparatus, electronic device, and storage medium.
[0006] This invention provides an operator optimization method, comprising:
[0007] Based on the obtained segmentation strategy, a first scheduling tree corresponding to the segmentation strategy is generated; the segmentation strategy is used to segment the data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the nodes in the first scheduling tree;
[0008] Based on the first schedule tree, each operator in the neural network model is optimized.
[0009] Optionally, optimizing each operator in the neural network model based on the first schedule tree includes:
[0010] The first schedule tree is deserialized to generate the target object corresponding to the first schedule tree.
[0011] Based on the target object, each operator is optimized.
[0012] Optionally, optimizing each operator based on the target object includes:
[0013] Perform a depth-first traversal operation on the target object to obtain the depth-first traversal result;
[0014] Based on the depth-first traversal results, the data associated with each operator is segmented to obtain at least one data block associated with each operator;
[0015] The operators are optimized based on the data blocks.
[0016] Optionally, the splitting strategy is a set of for loop statements pre-configured in a yaml file.
[0017] Optionally, the method further includes:
[0018] If an update to the segmentation strategy is detected, the first schedule tree is updated based on the updated segmentation strategy to generate a second schedule tree.
[0019] Based on the second schedule tree, each of the operators in the neural network model is optimized.
[0020] Optionally, the method further includes:
[0021] If the aforementioned splitting strategy is not obtained, obtain the pre-configured third schedule tree;
[0022] Based on the third schedule tree, each operator in the neural network model is optimized.
[0023] The present invention also provides an operator optimization apparatus, comprising:
[0024] The first acquisition module is used to generate a first schedule tree corresponding to the acquired segmentation strategy; the segmentation strategy is used to segment the data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the node in the first schedule tree;
[0025] The first optimization module is used to optimize each of the operators in the neural network model based on the first schedule tree.
[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the operator optimization method as described above.
[0027] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the operator optimization method as described above.
[0028] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the operator optimization method as described above.
[0029] The operator optimization method, apparatus, electronic device, and storage medium provided by this invention generate a first schedule tree corresponding to the obtained segmentation strategy. The for loop statement in the segmentation strategy corresponds one-to-one with the node in the first schedule tree. Since the first schedule tree is configurable and can support flexible and varied segmentation strategies, the first schedule tree is used to optimize each operator in the neural network model, avoiding the need to hard-code the segmentation strategy into the code project, thereby accelerating the tuning process of each operator. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0031] Figure 1 This is one of the flowcharts illustrating the operator optimization method provided by the present invention;
[0032] Figure 2This is a schematic diagram of the first schedule tree provided by the present invention;
[0033] Figure 3 This is a schematic diagram of the second schedule tree provided by the present invention;
[0034] Figure 4 This is the second flowchart of the operator optimization method provided by the present invention;
[0035] Figure 5 This is a logical schematic diagram of the operator optimization method provided by the present invention;
[0036] Figure 6 This is a schematic diagram of the operator optimization device provided by the present invention;
[0037] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0039] The following is combined with Figures 1 to 5 The operator optimization method provided by this invention will be described in detail. Figure 1 This is one of the flowcharts illustrating the operator optimization method provided by this invention. See [link / reference]. Figure 1 As shown, the method includes steps 101-102, wherein:
[0040] Step 101: Based on the obtained segmentation strategy, generate a first scheduling tree (scheduletree) corresponding to the segmentation strategy; the segmentation strategy is used to segment the data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the nodes in the first scheduling tree.
[0041] First, it should be noted that the subject of this invention can be any electronic device capable of operator optimization, such as any smartphone, smartwatch, desktop computer, laptop, etc.
[0042] This invention is applied to operator tuning scenarios in the software stack of artificial intelligence chips. High-performance operators are the foundation of efficient computation in artificial intelligence chips. In this embodiment, the segmentation strategy, also known as the tiling strategy, is a technique that utilizes shared memory on the Graphics Processing Unit (GPU) to reduce access to global memory, thereby improving the execution efficiency of kernel functions. The tiling strategy can segment the data associated with each operator in a neural network model, thereby reducing the amount of data input to each operator and effectively improving the performance of the neural network model.
[0043] Currently, tiling operators is performed using a fixed paradigm. Each time a new tiling strategy is introduced, the support needs to be expanded and hard-coded into the codebase for testing and verification. This fixed paradigm lacks flexibility and cannot support diverse tiling strategies.
[0044] Therefore, in this embodiment of the invention, the tiling strategy is converted into a tree structure, namely a schedule tree. A schedule tree is a tree-shaped representation of the execution order of a schedule, consisting of nodes and edges.
[0045] Based on the schedule tree, tiling strategies can be flexibly configured. When there is a new tiling, only the schedule tree needs to be updated, without having to hard-code the tiling strategy. This speeds up the iteration of performance optimization for developers and lays the foundation for intelligent tiling in the future.
[0046] Optionally, the splitting strategy is a set of for loop statements pre-configured in a yaml file.
[0047] Specifically, the segmentation strategy can be expressed as follows:
[0048]
[0049]
[0050] In this embodiment of the invention, the user can configure the tiling strategy arbitrarily in the yaml file. After obtaining the tiling strategy configured by the user, the tiling strategy needs to be abstracted into a first schedule tree.
[0051] For example, the tiling policy configured by the user is:
[0052]
[0053]
[0054] Therefore, after obtaining the tiling strategy from the YAML file, the tiling strategy needs to be converted into the first schedule tree. Figure 2 This is a schematic diagram of the first schedule tree provided by the present invention.
[0055] Step 102: Based on the first schedule tree, optimize each operator in the neural network model.
[0056] In this embodiment of the invention, after the segmentation strategy is converted into a first schedule tree, the first schedule tree represents the segmentation strategy for segmenting the data associated with each operator in the neural network model.
[0057] Based on the first schedule tree, the data associated with each operator is segmented, thereby reducing the amount of data processing for each operator and optimizing each operator.
[0058] The operator optimization method provided by this invention generates a first schedule tree corresponding to the obtained segmentation strategy. Since the first schedule tree is configurable and can support flexible and varied segmentation strategies, the first schedule tree is used to optimize each operator in the neural network model, avoiding the need to add the segmentation strategy to the code project in the form of hard coding, thereby accelerating the tuning process of each operator.
[0059] Optionally, the optimization of each operator in the neural network model based on the first schedule tree can be achieved through the following steps 1)-2):
[0060] Step 1) Deserialize the first schedule tree to generate the target object corresponding to the first schedule tree;
[0061] Step 2) Optimize each operator based on the target object.
[0062] In this embodiment of the invention, for example, if a user predefines a tiling strategy in a YAML file, the user-configured tiling strategy is first obtained from the YAML file. Then, after converting the tiling strategy into a first schedule tree, the first schedule tree needs to be deserialized to generate the target object corresponding to the computer-executable first schedule tree.
[0063] Then, the data associated with each operator is segmented based on the target object, thereby optimizing each operator.
[0064] Optionally, the optimization of each operator based on the target object can be achieved through the following steps 1)-3):
[0065] Step 1) Perform a depth-first traversal operation on the target object to obtain the depth-first traversal result;
[0066] Step 2) Based on the depth-first traversal results, the data associated with each operator is segmented to obtain at least one data block associated with each operator;
[0067] Step 3) Optimize each operator based on each of the data blocks.
[0068] In this embodiment of the invention, depth-first traversal means starting from the initial access node. The initial access node may have multiple adjacent nodes. The strategy of depth-first traversal is to first access the first adjacent node, and then use this accessed adjacent node as the initial node to access its first adjacent node. It can be understood as follows: after accessing the current node, the first adjacent node of the current node is accessed first.
[0069] After obtaining the depth-first traversal results, code can be generated to segment the data associated with each operator. Then, the data associated with each operator is segmented to obtain at least one data block associated with each operator. Finally, each data block is input into each operator so that each operator processes each data block to generate the final assembly code.
[0070] Optionally, if an update to the segmentation strategy is detected, the following steps need to be performed:
[0071] Step 1) If the partitioning strategy is updated, update the first schedule tree based on the updated partitioning strategy to generate the second schedule tree;
[0072] Step 2) Based on the second schedule tree, optimize each operator in the neural network model.
[0073] In this embodiment of the invention, since each node in the first schedule tree can be flexibly configured, when an update to the tiling strategy is detected, the first schedule tree can be adaptively adjusted based on the updated tiling strategy to generate the second schedule tree. This supports flexible and varied tiling strategies, avoids hard-coding the tiling strategy, and accelerates the tuning process of each operator.
[0074] In practical applications, with Figure 2 For example, Figure 2 This is the first schedule tree before the tiling strategy update. After detecting an update to the tiling strategy, the updated tiling strategy is as follows:
[0075]
[0076] Then, based on the updated tiling, the first schedule tree is updated to generate the second schedule tree. Figure 3 This is a schematic diagram of the second schedule tree provided by the present invention.
[0077] Optionally, if the segmentation strategy is not obtained, the following steps need to be performed:
[0078] Step 1) If the splitting strategy is not obtained, obtain the pre-configured third schedule tree;
[0079] Step 2) Optimize each operator in the neural network model based on the third schedule tree.
[0080] In this embodiment of the invention, if no segmentation strategy is obtained, for example, if the user has not predefined a segmentation strategy in the yaml file, then a pre-configured third schedule tree is obtained, and the operators in the neural network model are optimized.
[0081] Figure 4 This is the second flowchart illustrating the operator optimization method provided by this invention. See also... Figure 4 As shown, the method includes steps 401-410, wherein:
[0082] Step 401: Obtain the tiling strategy, where the tiling strategy is a for loop statement pre-configured in the yaml file.
[0083] It should be noted that if the tiling strategy is not obtained, proceed with steps 409-410.
[0084] Step 402: Based on the obtained tiling strategy, generate the first schedule tree corresponding to the tiling strategy;
[0085] Step 403: Deserialize the first schedule tree to generate the target object corresponding to the first schedule tree.
[0086] Step 404: Perform a depth-first traversal on the target object to obtain the depth-first traversal result.
[0087] Step 405: Based on the depth-first traversal results, the data associated with each operator in the neural network model is segmented to obtain at least one data block associated with each operator.
[0088] Step 406: Optimize each operator based on each data block.
[0089] Step 407: If an update to the tiling strategy is detected, update the first schedule tree based on the updated tiling strategy to generate the second schedule tree.
[0090] Step 408: Optimize each operator in the neural network model based on the second schedule tree.
[0091] Step 409: If the splitting strategy is not obtained, obtain the pre-configured third schedule tree.
[0092] Step 410: Optimize each operator in the neural network model based on the third schedule tree.
[0093] The operator optimization method provided by this invention generates a first schedule tree corresponding to the obtained segmentation strategy. Since the first schedule tree is configurable and can support flexible and varied segmentation strategies, the first schedule tree is used to optimize each operator in the neural network model, avoiding the need to add the segmentation strategy to the code project in the form of hard coding, thereby accelerating the tuning process of each operator.
[0094] Figure 5 This is a logical schematic diagram of the operator optimization method provided by this invention. See also... Figure 5 As shown, (a) represents a logical diagram of optimizing operators using the tiling strategy in the prior art.
[0095] Specifically, the tiling algorithm is obtained from the Open Pluggable Specification (OPS) file, and then the execution code of the tiling strategy is generated using the tiling generation module. The execution code of the tiling strategy is used to optimize the operators stored in llapis, and finally an Asm file is generated.
[0096] In (a), the tiling strategy needs to be hard-coded. This fixed paradigm lacks flexibility and is a very inefficient tuning method.
[0097] (b) illustrates the logic of optimizing operators using a tiling strategy provided by the present invention. In (b), the Schedule Manager first determines whether predefined schedules can be obtained from the yaml file.
[0098] If so, select the corresponding schedule tree. It should be noted that the schedule tree is automatically generated by YAML.
[0099] If not, the default schedule tree is obtained from the Tiling Algorithm module. It should be noted that the Tiling Algorithm is implemented in C++.
[0100] Then, the Generation System and Generators modules are used to perform a depth-first traversal of the schedule tree, generating the execution code for the tiling strategy to optimize the operators stored in llapis, and finally generating an Asm file.
[0101] The operator optimization apparatus provided by the present invention is described below. The operator optimization apparatus described below and the operator optimization method described above can be referred to in correspondence. Figure 6 This is a schematic diagram of the operator optimization device provided by the present invention, as shown below. Figure 6 As shown, the operator optimization device 600 includes: a first acquisition module 601 and a first optimization module 602, wherein:
[0102] The first acquisition module 601 is used to generate a first scheduling tree corresponding to the acquired segmentation strategy based on the segmentation strategy; the segmentation strategy is used to segment the data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the nodes in the first scheduling tree;
[0103] The first optimization module 602 is used to optimize each of the operators in the neural network model based on the first schedule tree.
[0104] The operator optimization device provided by the present invention generates a first schedule tree corresponding to the obtained segmentation strategy. Since the first schedule tree is configurable and can support flexible and varied segmentation strategies, the first schedule tree is used to optimize each operator in the neural network model, avoiding the need to add the segmentation strategy to the code project in the form of hard coding, thereby accelerating the tuning process of each operator.
[0105] Optionally, the optimization module 602 is further configured to:
[0106] The first schedule tree is deserialized to generate the target object corresponding to the first schedule tree.
[0107] Based on the target object, each operator is optimized.
[0108] Optionally, the optimization module 602 is further configured to:
[0109] Perform a depth-first traversal operation on the target object to obtain the depth-first traversal result;
[0110] Based on the depth-first traversal results, the data associated with each operator is segmented to obtain at least one data block associated with each operator;
[0111] The operators are optimized based on the data blocks.
[0112] Optionally, the splitting strategy is a set of for loop statements pre-configured in a yaml file.
[0113] Optionally, the device further includes:
[0114] The update module is used to update the first schedule tree based on the updated schedule tree when the update of the segmentation strategy is detected, and generate a second schedule tree.
[0115] The second optimization module is used to optimize each of the operators in the neural network model based on the second schedule tree.
[0116] Optionally, the device further includes:
[0117] The second acquisition module is used to acquire a pre-configured third schedule tree if the segmentation strategy is not acquired.
[0118] The third optimization module is used to optimize each operator in the neural network model based on the third schedule tree.
[0119] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute an operator optimization method. This method includes: generating a first scheduling tree corresponding to the acquired segmentation strategy; the segmentation strategy is used to segment data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the nodes in the first scheduling tree; and optimizing each operator in the neural network model based on the first scheduling tree.
[0120] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0121] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the operator optimization method provided by the above methods. The method includes: generating a first scheduling tree corresponding to the obtained segmentation strategy; the segmentation strategy is used to segment the data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the nodes in the first scheduling tree; and optimizing each operator in the neural network model based on the first scheduling tree.
[0122] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the operator optimization method provided by the above methods. The method includes: generating a first scheduling tree corresponding to the obtained segmentation strategy; the segmentation strategy is used to segment the data associated with at least one operator in the neural network model; the for loop statement in the segmentation strategy corresponds one-to-one with the nodes in the first scheduling tree; and optimizing each of the operators in the neural network model based on the first scheduling tree.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An operator optimization method applied to the software stack of an artificial intelligence chip, characterized in that, include: A segmentation strategy is obtained for a target operator, which is at least one operator in a neural network model. The segmentation strategy is configured in the form of a for loop statement to segment the data associated with the target operator into at least one data block and to reduce access to global storage by utilizing shared memory on the graphics processor. Based on the for loop statement configured in the splitting strategy, a first scheduling tree corresponding to the splitting strategy is generated; The for loop statement in the segmentation strategy corresponds one-to-one with the node in the first scheduling tree; The first scheduling tree is deserialized to generate the target object corresponding to the first scheduling tree, and the target object is subjected to a depth-first traversal operation to obtain the depth-first traversal result. Based on the depth-first traversal result, the data associated with the target operator is segmented to obtain at least one data block associated with the target operator, and each data block is input into the target operator to generate the assembly code of the target operator; The assembly code is used to control the graphics processor to perform the following operations: load at least one data block associated with the target operator from global storage to shared storage, and perform calculations using the data block in the shared storage.
2. The operator optimization method applied to the software stack of an artificial intelligence chip according to claim 1, characterized in that, The segmentation strategy consists of multiple for loop statements pre-configured in a yaml file.
3. The operator optimization method applied to the software stack of an artificial intelligence chip according to claim 1, characterized in that, The method further includes: If an update to the partitioning strategy is detected, the first scheduling tree is updated based on the updated partitioning strategy to generate a second scheduling tree; Based on the second scheduling tree, each operator in the neural network model is optimized.
4. The operator optimization method applied to the software stack of an artificial intelligence chip according to claim 1, characterized in that, The method further includes: If the aforementioned splitting strategy is not obtained, obtain the pre-configured third scheduling tree; Based on the third scheduling tree, each operator in the neural network model is optimized.
5. An operator optimization device applied in an artificial intelligence chip software stack, characterized in that, include: The first acquisition module is used to acquire a segmentation strategy for a target operator, wherein the target operator is at least one operator in a neural network model. The segmentation strategy is configured in the form of a for loop statement to segment the data associated with the target operator into at least one data block and to reduce access to global storage by utilizing shared storage on the graphics processor. Based on the for loop statement configured in the splitting strategy, a first scheduling tree corresponding to the splitting strategy is generated; The for loop statement in the segmentation strategy corresponds one-to-one with the node in the first scheduling tree; The first optimization module is used to deserialize the first scheduling tree to generate a target object corresponding to the first scheduling tree, and perform a depth-first traversal operation on the target object to obtain a depth-first traversal result; based on the depth-first traversal result, the data associated with the target operator is segmented to obtain at least one data block associated with the target operator, and each data block is input into the target operator to generate the assembly code of the target operator; The assembly code is used to control the graphics processor to perform the following operations: load at least one data block associated with the target operator from global storage to shared storage, and perform calculations using the data block in the shared storage.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the operator optimization method applied to the software stack of an artificial intelligence chip as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the operator optimization method applied to the software stack of an artificial intelligence chip as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the operator optimization method applied to the software stack of an artificial intelligence chip as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Graph neural network optimization method and graph neural network inference system
CN115860061A
Operator tuning method and device, electronic equipment and storage medium
CN116244059A