Neural network operator fusion method and device, electronic equipment and storage medium

By dividing the target fusion module in the neural network into multiple submodules that perform calculations in different operator libraries and perform video memory management optimization, the problem of limited optimization of operator fusion performance in neural network is solved, and more efficient computing performance and video memory resource utilization is achieved.

CN119939512APending Publication Date: 2025-05-06广州壁仞智能科技有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510075450.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, neural network operator fusion relies on designated operator databases, resulting in limited performance optimization and it is difficult to improve the performance optimization level of neural networks.

Method used

By determining the target fusion module in the neural network, dividing it into multiple submodules that perform calculations in different operator libraries, the memory usage life cycle analysis is performed on the operators in each submodule, the memory multiplexing and release information is generated, and the video memory is allocated based on this, so that the target fusion module performs calculations in each operator library.

Benefits of technology

The comprehensive utilization of multiple operator libraries is realized, the computing performance of the target fusion module is improved, the dependence on designated operator libraries is avoided, the performance optimization level of neural networks is improved, and the video memory multiplexing and timely release is maximized, reducing the demand for video memory resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939512A_ABST
    Figure CN119939512A_ABST
Patent Text Reader

Abstract

The invention provides a neural network operator fusion method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining a target fusion module in a neural network; the target fusion module comprises a plurality of operators; segmenting the target fusion module to obtain a plurality of sub-modules for executing calculation in different operator libraries; analyzing the video memory occupancy life cycle of the operator in each sub-module in the corresponding operator library, and determining operator video memory multiplexing information and operator video memory release information; and based on the operator video memory multiplexing information and the operator video memory release information, allocating a video memory to the target fusion module, so that the target fusion module executes calculation in each operator library. According to the method and the device provided by the invention, the performance optimization level of the neural network is improved, and the demand of the neural network for video memory resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a neural network operator fusion method, device, electronic device and storage medium. Background Art

[0002] Neural network operator fusion refers to replacing a sub-network that takes a long time or conforms to certain rules in a neural network with a fusion operator. By processing the fusion operator through a specified operator library, the performance of the neural network can be optimized. This makes neural network operator fusion very dependent on the specified operator library. The performance of the specified operator library limits the performance optimization level of the neural network.

[0003] Therefore, how to improve the performance optimization level of neural networks has become a technical problem that needs to be solved urgently in the industry. Summary of the invention

[0004] The present invention provides a neural network operator fusion method, device, electronic device and storage medium, which are used to solve the technical problem of how to improve the performance optimization level of a neural network.

[0005] The present invention provides a neural network operator fusion method, comprising: Determining a target fusion module in a neural network; the target fusion module includes a plurality of operators; The target fusion module is divided into multiple submodules for performing calculations in different operator libraries; Analyze the memory usage life cycle of operators in each submodule in the corresponding operator library to determine operator memory reuse information and operator memory release information; Based on the operator memory reuse information and the operator memory release information, memory is allocated to the target fusion module, so that the target fusion module performs calculations in each operator library.

[0006] In some embodiments, the target fusion module is segmented to obtain a plurality of submodules that perform calculations in different operator libraries, including: Determining a current operator and a next operator in the target fusion module; When the operator library corresponding to the current operator is the same as the operator library corresponding to the next operator, the current operator and the next operator are divided into the same submodule; When the operator library corresponding to the current operator is different from the operator library corresponding to the next operator, the current operator and the next operator are divided into different sub-modules.

[0007] In some embodiments, the operator library corresponding to each operator is determined based on the following steps: When the number of target operator libraries supporting the current operator is one, determining the target operator library as the operator library corresponding to the current operator; In the case that there are multiple target operator libraries supporting the current operator, the computational cost of the current operator in each target operator library is determined based on the computational characteristics of the current operator; and the target operator library corresponding to the minimum computational cost is determined as the operator library corresponding to the current operator.

[0008] In some embodiments, determining the target operator library corresponding to the minimum computing cost as the operator library corresponding to the current operator includes: When there are multiple target operator libraries corresponding to the minimum computing cost, determine the computing performance of each target operator library; The target operator library corresponding to the optimal computing performance is determined as the operator library corresponding to the current operator.

[0009] In some embodiments, analyzing the memory occupation life cycle of operators in each submodule in the corresponding operator library to determine operator memory reuse information and operator memory release information includes: Determine the memory usage lifecycle of each operator in each submodule when performing calculations in the corresponding operator library; When it is determined that the video memory occupation life cycle of any operator does not overlap with the video memory occupation life cycle of other operators and there is no data dependency relationship between the any operator and the other operators, it is determined that there is a video memory multiplexing relationship between the any operator and the other operators; Based on the memory reuse relationship between the operators, the operator memory reuse information is generated.

[0010] In some embodiments, after determining the video memory occupation life cycle of each operator in each submodule when performing calculation in the corresponding operator library, the method further includes: When it is determined that the output of any operator has no dependency relationship with the input of other operators, based on the video memory occupation life cycle of the any operator, determining the video memory release information of the any operator; The operator video memory release information is generated based on the video memory release information of each operator.

[0011] The present invention provides a neural network operator fusion device, comprising: An operator determination unit, used to determine a target fusion module in a neural network; the target fusion module includes a plurality of operators; An operator segmentation unit, used to segment the target fusion module to obtain multiple submodules that perform calculations in different operator libraries; A memory analysis unit is used to analyze the memory occupancy life cycle of operators in each submodule in the corresponding operator library, and determine operator memory reuse information and operator memory release information; A video memory allocation unit is used to allocate video memory to the target fusion module based on the operator video memory reuse information and the operator video memory release information, so that the target fusion module performs calculations in each operator library.

[0012] The present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the neural network operator fusion method when executing the computer program.

[0013] The present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program implements the neural network operator fusion method when executed by a processor.

[0014] The present invention provides a computer program product, comprising a computer program, wherein the computer program implements the neural network operator fusion method when executed by a processor.

[0015] The neural network operator fusion method, device, electronic device and storage medium provided by the present invention determine a target fusion module in a neural network; the target fusion module includes multiple operators; the target fusion module is segmented to obtain multiple sub-modules that perform calculations in different operator libraries; the video memory occupancy life cycle of the operators in each sub-module in the corresponding operator library is analyzed to determine the operator video memory reuse information and the operator video memory release information; based on the operator video memory reuse information and the operator video memory release information, video memory is allocated to the target fusion module so that the target fusion module performs calculations in each operator library; since the target fusion module is segmented to obtain multiple sub-modules, each sub-module can perform calculations in different operator libraries, and the comprehensive use of multiple operator libraries is realized to optimize the calculation performance of the target fusion module without being limited to a specified operator library, thereby improving the performance optimization level of the neural network; since the video memory occupancy life cycle of each operator is analyzed, the video memory reuse can be maximized in the video memory management process of the target fusion module, and the video memory release can be realized in the most timely manner, thereby reducing the demand of the neural network for video memory resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a flow chart of the neural network operator fusion method provided by the present invention.

[0019] Figure 2 It is a schematic diagram of the architecture of the neural network operator fusion provided by the present invention.

[0020] Figure 3 It is a calculation diagram of the fusion operator provided by the present invention.

[0021] Figure 4 These are effect diagrams before and after the video memory allocation optimization provided by the present invention.

[0022] Figure 5 It is a structural schematic diagram of the neural network operator fusion device provided by the present invention.

[0023] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units or modules is not necessarily limited to those steps or units or modules that are clearly listed, but may include other steps or units or modules that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] The related technology determines the fusion operator in the neural network through the existing operator fusion algorithm. The fusion operator is usually calculated only in a specified operator library, which makes the implementation of the fusion operator highly dependent on the specified operator library.

[0027] In order to solve the shortcomings of related technologies, Figure 1 This is one of the flow charts of the neural network operator fusion method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 , step 130 and step 140 .

[0028] Step 110, determining a target fusion module in the neural network; the target fusion module includes multiple operators.

[0029] Specifically, the execution subject of the neural network operator fusion method provided in the embodiment of the present invention is a neural network operator fusion device. The device can be implemented by software, such as a neural network operator fusion program; or by hardware, such as a computer or server that executes the neural network operator fusion method.

[0030] The application scenario of the method provided in the embodiment of the present invention is to fuse multiple operators in a neural network. The embodiment of the present invention does not specifically limit the model structure of the neural network.

[0031] Operators refer to the basic mathematical operations or operations performed during the neural network calculation process. They are the basic building blocks for neural network construction and training. For example, operators can include matrix multiplication (MatMul), convolution (Conv), activation functions (ReLU, Sigmoid, etc.), pooling, normalization (BatchNorm), etc. Each operator usually involves some transformation or calculation of the input data to generate output results.

[0032] The target fusion module combines multiple operators in the neural network into a more efficient single operation, thereby optimizing the calculation process, reducing the consumption of memory and computing resources, and improving the calculation efficiency, thereby increasing the model reasoning speed and reducing latency. For example, convolution can be fused with normalization, convolution can be fused with activation function, and matrix multiplication can be fused with addition.

[0033] According to the model structure of the neural network, the sub-network or module in it can be determined as the target fusion module. For example, for a neural network, the sub-network composed of operator A, operator B, operator C and operator D in the model structure (A->B->C->A->B->D, "->" indicates the direction of data flow between operators) can be used as the target fusion module.

[0034] Step 120: Split the target fusion module into multiple submodules that perform calculations in different operator libraries.

[0035] Specifically, an operator library is a collection of pre-implemented and optimized operators that contain efficient implementations of various common computing operations. The operator library provides infrastructure for deep learning frameworks, allowing developers to call these operations directly without having to manually implement the details of each operation. The operator library provides a large number of implemented and optimized operators that are used to perform various neural network tasks. They are implemented through efficient code and can be accelerated when running on different hardware (such as CPUs and GPUs), thereby improving computing efficiency.

[0036] The calculation characteristics of each operator in the target fusion module can be analyzed to determine the operator library corresponding to each operator. The corresponding operator library is an operator library that supports the operator to perform corresponding calculations. According to the correspondence between the operator and the operator library, the operators in the target fusion module are grouped to achieve segmentation of the target fusion module to obtain multiple sub-modules. At least two of these sub-modules correspond to different operator libraries.

[0037] For example, by analyzing the computing characteristics of each operator in the target fusion module (A->B->C->A->B->D), it is determined that the target fusion module can be divided into three sub-modules, namely (A->B->C), (A), and (B->D). The computing performance of sub-module (A->B->C) is the best in operator library 1, sub-module (A) is only supported in operator library 2, and sub-module (B->D) has the highest utilization rate of video memory when implemented in operator library 3.

[0038] By dividing the target fusion module, different operator libraries can be assigned to each operator, thereby improving the computing performance of each operator and further improving the computing performance of the target fusion module.

[0039] Step 130: Analyze the memory occupancy life cycle of the operators in each submodule in the corresponding operator library to determine operator memory reuse information and operator memory release information.

[0040] Specifically, video memory refers to the memory allocated by the graphics processing unit (GPU) for operators when processing neural network computing tasks. The video memory occupancy lifecycle refers to the entire process of video memory allocation, use, and release when using a graphics processor for neural network training and reasoning.

[0041] Memory reuse refers to avoiding allocating independent memory areas to each operator during the training and reasoning of the neural network model. Instead, memory is shared between different operators, thereby reducing memory usage and improving computing efficiency. Operator memory reuse information refers to information related to memory reuse between operators, which may include operator name, memory address, memory size, and memory occupancy time of each operator.

[0042] Video memory release refers to the recycling of video memory resources for other calculations when they are no longer needed during the training and reasoning of the neural network model. Reasonable video memory release can improve resource utilization efficiency and prevent memory overflow. Operator video memory release information refers to information related to the video memory release of each operator, which can include operator name, video memory address, video memory size, release time and other information.

[0043] By analyzing the memory occupancy life cycle of operators in each sub-module in the corresponding operator library, it is possible to determine whether each operator can reuse the memory and whether the memory can be released in time, and then generate operator memory reuse information and operator memory release information.

[0044] Step 140: Allocate video memory to the target fusion module based on the operator video memory reuse information and the operator video memory release information, so that the target fusion module performs calculations in each operator library.

[0045] Specifically, based on the operator memory reuse information and operator memory release information, the memory management information such as memory allocation, memory occupancy, memory reuse and memory release of each operator can be determined, and overall optimization management can be performed to maximize memory reuse and realize memory release in the most timely manner.

[0046] After determining the memory management information (including memory allocation, memory occupancy, memory reuse, and memory release) and operator allocation information (including operators and operator libraries corresponding to submodules) of the target fusion module during execution, a computational execution plan for the target fusion module can be generated.

[0047] The neural network operator fusion method provided by the embodiment of the present invention determines a target fusion module in a neural network; the target fusion module includes multiple operators; the target fusion module is segmented to obtain multiple sub-modules that perform calculations in different operator libraries; the video memory occupancy life cycle of the operators in each sub-module in the corresponding operator library is analyzed to determine the operator video memory reuse information and the operator video memory release information; based on the operator video memory reuse information and the operator video memory release information, video memory is allocated to the target fusion module so that the target fusion module performs calculations in each operator library; since the target fusion module is segmented to obtain multiple sub-modules, each sub-module can perform calculations in different operator libraries, and the comprehensive use of multiple operator libraries is realized to optimize the computing performance of the target fusion module without being limited to a specified operator library, thereby improving the performance optimization level of the neural network; since the video memory occupancy life cycle of each operator is analyzed, video memory reuse can be maximized in the video memory management process of the target fusion module, and video memory release can be realized in the most timely manner, thereby reducing the demand for video memory resources of the neural network.

[0048] It should be noted that each implementation of the present invention can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.

[0049] In some embodiments, the target fusion module is segmented to obtain multiple submodules that perform calculations in different operator libraries, including: Determine the current operator and the next operator in the target fusion module; When the operator library corresponding to the current operator is the same as the operator library corresponding to the next operator, the current operator and the next operator are divided into the same submodule; When the operator library corresponding to the current operator is different from the operator library corresponding to the next operator, the current operator and the next operator are divided into different sub-modules.

[0050] Specifically, data dependency is the data dependency between different operators. For example, in a neural network, the output of one operator is the input of the next operator, so the latter depends on the calculation result of the former.

[0051] When the target fusion module is segmented, the current operator and the next operator can be determined according to the data dependency relationship between the operators. According to the calculation characteristics of the operators, the operator library corresponding to the current operator and the operator library corresponding to the next operator can be determined.

[0052] If the operator library corresponding to the current operator is the same as the operator library corresponding to the next operator, the current operator and the next operator can be divided into the same sub-module; if the operator library corresponding to the current operator is different from the operator library corresponding to the next operator, the current operator and the next operator are divided into different sub-modules.

[0053] The neural network operator fusion method provided in the embodiment of the present invention divides the target fusion module according to the different operator libraries corresponding to each operator, so that each sub-module can perform calculations in different operator libraries, so as to utilize multiple operator libraries to optimize the computing performance of the target fusion module.

[0054] In some embodiments, the operator library corresponding to each operator is determined based on the following steps: When the number of target operator libraries supporting the current operator is one, the target operator library is determined to be the operator library corresponding to the current operator; When there are multiple target operator libraries supporting the current operator, the computational cost of the current operator in each target operator library is determined based on the computational characteristics of the current operator; and the target operator library corresponding to the minimum computational cost is determined as the operator library corresponding to the current operator.

[0055] Specifically, in the deep learning framework, different operator libraries can support different operators for optimized accelerated computing. For example, the SUDNN (SUPA Deep Neural Network library) operator library can support operators such as convolution, pooling, normalization, activation functions, and matrix multiplication in deep neural networks for accelerated computing; the SUBLAS (SUPA Basic Linear Algebra Subprograms) operator library can support operators related to matrices and vectors for accelerated computing; the Supa operator library can support operators related to convolution, matrix operations, and parallel computing for accelerated computing. You can also customize related custom operator libraries according to user needs to support special operators.

[0056] The following takes the current operator as an example to illustrate the method for determining the operator library. When determining the operator library corresponding to the current operator, the support information of each operator library can be queried, and the operator library that supports the current operator is determined as the target operator library.

[0057] If the number of the target operator library is one, the target operator library may be determined as the operator library corresponding to the current operator.

[0058] If there are multiple target operator libraries, you can determine them based on the computing characteristics of the current operator. The computing characteristics of an operator refer to the characteristics or properties related to the calculation process of the operator, which may include shape, layout, architecture, data type, and memory usage.

[0059] Shape refers to the dimension or shape of a tensor. Operators usually need to process tensors of matching shapes, and different shapes will affect the construction and execution of the computational graph.

[0060] Layout refers to the way tensors are stored in memory. Different memory layouts affect the efficiency of calculations, especially when using graphics processors. Choosing a suitable layout can improve the efficiency of memory access.

[0061] Architecture refers to the architecture of the computing platform. The execution mode of operators may depend on the hardware architecture. The optimization of certain operators may be different on different architectures.

[0062] Data type refers to the data type of a tensor. The data type determines the precision and memory usage of the data. For example, float32 is often used in deep learning because it strikes a good balance between speed and precision, while float64 provides higher precision but takes up more memory and is slower to compute.

[0063] Memory usage refers to the memory resources used by tensors during computation. Memory characteristics include the size of the tensor and whether additional memory allocation is required (for example, whether additional temporary storage is required to save intermediate results).

[0064] In addition, computing characteristics can also include reorder and conditional processing. Reorder means that input data or intermediate data in the computing process needs to be reordered. For example, some hardware may be sensitive to the order of data storage (such as row priority and column priority), and data may need to be rearranged to improve computing efficiency and memory access performance. Conditional processing means that different execution paths may need to be selected according to the different characteristics of the operator library. For example, if an operator library is optimized for a certain data type (such as low-precision calculations), it may be necessary to adjust according to the data type when selecting the operator library.

[0065] The computational cost of the current operator in each target operator library can be determined based on the computational characteristics of the current operator. The computational cost refers to the resource consumption required to complete a certain operator computation task, and is usually used to measure the computational efficiency of the operator, such as computational time, memory resource requirements, storage resource requirements, etc. The target operator library corresponding to the minimum computational cost can be determined as the operator library corresponding to the current operator.

[0066] The neural network operator fusion method provided in the embodiment of the present invention determines the operator library corresponding to each operator according to the computing characteristics of each operator, which is conducive to achieving the best computing performance of each operator in the corresponding operator library.

[0067] In some embodiments, determining the target operator library corresponding to the minimum computational cost as the operator library corresponding to the current operator includes: When there are multiple target operator libraries corresponding to the minimum computing cost, determine the computing performance of each target operator library; The target operator library corresponding to the optimal computing performance is determined as the operator library corresponding to the current operator.

[0068] Specifically, if there are multiple target operator libraries corresponding to the minimum computing cost, the target operator libraries may be further screened according to the computing performance of each target operator library.

[0069] The computing performance of an operator library refers to the execution efficiency of each operator in the operator library and the efficiency of computing resource usage when performing deep learning or computing tasks. It can be measured by parameters such as computing speed, resource utilization, memory usage, parallel computing capability, hardware adaptability, etc. It can also include whether broadcast, reorder, and memory area (workspace) allocation are required.

[0070] Broadcasting refers to automatically expanding smaller tensors when performing arithmetic operations between tensors of different shapes so that they can operate on the same dimension. Rearrangement refers to adjusting the memory layout or storage order of data before performing calculations. In convolution operations, some intermediate results may be generated (such as intermediate caches of convolution kernels and input data), which need to be stored in a temporary memory space. This part of memory is called workspace.

[0071] After the computing performance of each target operator library is determined, the target operator library corresponding to the optimal computing performance may be determined as the operator library corresponding to the current operator.

[0072] The neural network operator fusion method provided by the embodiment of the present invention selects the operator library corresponding to the current operator according to the computing performance when there are multiple target operator libraries corresponding to the minimum computing cost. This is conducive to selecting the operator library with the best computing performance for the operator, and realizes the comprehensive use of multiple operator libraries to optimize the computing performance of the target fusion module without being limited to the specified operator library.

[0073] In some embodiments, the memory occupation life cycle of operators in each submodule in the corresponding operator library is analyzed to determine operator memory reuse information and operator memory release information, including: Determine the memory usage lifecycle of each operator in each submodule when performing calculations in the corresponding operator library; When it is determined that the video memory occupation life cycle of any operator does not overlap with the video memory occupation life cycle of other operators and there is no data dependency relationship between any operator and other operators, it is determined that there is a video memory reuse relationship between any operator and other operators; Based on the memory reuse relationship between operators, operator memory reuse information is generated.

[0074] Specifically, the video memory required to be allocated for each operator in each submodule to perform calculation in the corresponding operator library and the video memory occupancy life cycle corresponding to the video memory can be analyzed and determined.

[0075] If the video memory occupation lifecycle of any operator does not overlap with the video memory occupation lifecycle of other operators and there is no data dependency relationship between the operator and other operators, it can be determined that other operators can reuse the video memory of the operator or the operator can reuse the video memory of other operators. In other words, there is a video memory reuse relationship between the operator and other operators.

[0076] Finally, the memory reuse relationship between operators in each sub-module is summarized to generate operator memory reuse information.

[0077] The neural network operator fusion method provided in the embodiment of the present invention generates operator memory reuse information according to the memory reuse relationship between each operator, which is conducive to maximizing the memory reuse in the memory management process of the target fusion module and reducing the demand of the neural network for memory resources.

[0078] In some embodiments, after determining the video memory required to be allocated for each operator in each submodule to perform calculations in the corresponding operator library and the video memory occupancy life cycle, the method further includes: When it is determined that the output of any operator has no dependency relationship with the input of other operators, based on the video memory occupation life cycle of any operator, the video memory release information of any operator is determined; Based on the video memory release information of each operator, the operator video memory release information is generated.

[0079] Specifically, if the output of any operator has no dependency on the input of other operators, it means that the video memory of the operator can be released in time after use. The video memory release information of the operator can be determined according to the video memory occupation life cycle of the operator, and used to recycle it for use in other calculations. The video memory release information may include video memory location, video memory size, and release time.

[0080] Finally, the video memory release information of each operator is summarized to generate operator video memory release information.

[0081] The neural network operator fusion method provided by the embodiment of the present invention generates operator memory release information according to the memory release information of each operator, realizes the most timely memory release in the memory management process of the target fusion module, and reduces the demand of the neural network for memory resources.

[0082] Figure 2 is a schematic diagram of the architecture of the neural network operator fusion provided by the present invention, such as Figure 2As shown in the figure, the architecture includes the front-end and the back-end. The front-end refers to the deep learning framework, and the back-end refers to multiple operator libraries. The deep learning framework can be used to process the neural network, determine the target fusion module (fusion operator), split the target fusion operator, and call different operator libraries to perform calculations.

[0083] The specific execution is realized by the graphics processor (such as an accelerator card) in the computing terminal. The graphics processor manages the memory resources according to the operator memory reuse information and operator memory release information in the deep learning framework.

[0084] Figure 3 is a calculation diagram of the fusion operator provided by the present invention, such as Figure 3 As shown in the figure, for a neural network, the sub-network composed of operators A, B, C, and D in the model structure (A->B->C->A->B->D, "->" indicates the direction of data flow between operators) can be used as the target fusion module. The dotted line indicates the memory reuse relationship. The following steps can quickly realize a fusion operator with optimal performance and memory usage: Step 1: Determine whether the four basic operators (operator A, operator B, operator C, and operator D) in the target fusion module are implemented in the deep learning framework; Step 2: If the conditions in the previous step are met, a fusion operator can be registered in the framework for this target fusion module; Step 3: Analyze the computational characteristics of each operator in the target fusion module, and determine the segmentation scheme of the target fusion module based on the implementation of the basic operators in the framework.

[0085] After analysis, it is known that (A->B->C) has the best performance on backend 1 (operator library 1), the input scale of (A) is only supported on backend 2, and (B->D) has the best memory utilization on backend 3. Therefore, the target fusion module can be split into sub-modules (A->B->C), (A), and (B->D) that are executed on three backends (operator libraries) respectively.

[0086] Step 4: Further analyze the memory occupancy life cycle of all tensors in the target fusion module, and find that in the process of implementing (A->B->C), the output memory of C (C_output) can reuse the output memory of A (A_output), and after this submodule is executed, the output memory of B (B_output) can be released. At this time, the peak value of the memory is: all input memory + A_output + B_output.

[0087] After (A) is executed on backend 2, the video memory of B's ​​output memory (B_output) can be released, and the peak video memory is: all input video memory + C_output (multiplexed A_output) + D_output.

[0088] After (B->D) is executed on backend 3, the peak memory usage is: all input memory + F_output (reused D_output) + E_output; if the existing implementation method is used (without considering memory reuse and memory release), the peak memory usage of the entire target fusion module is: all input memory + A_output + B_output + C_output + D_output + E_output.

[0089] All input memory includes input_1, input_2, input_3 and input_4.

[0090] Step 5: Based on the analysis of the previous four steps, implement the target fusion module algorithm.

[0091] Figure 4 : is the effect diagram before and after the video memory allocation optimization provided by the present invention, such as Figure 4 As shown in the figure, free means that the video memory is not occupied, and each square represents a video memory unit. Before the video memory allocation is optimized, the maximum video memory peak requires 10 video memory units; after the video memory allocation is optimized, the maximum video memory peak requires 6 video memory units.

[0092] The following describes a system provided by an embodiment of the present invention. The system described below and the method described above can be referenced to each other.

[0093] Figure 5 Schematic diagram of the structure of the neural network operator fusion device provided by the present invention. Figure 5 As shown, the device comprises: An operator determination unit 510 is used to determine a target fusion module in a neural network; the target fusion module includes a plurality of operators; An operator segmentation unit 520 is used to segment the target fusion module to obtain multiple submodules that perform calculations in different operator libraries; The video memory analysis unit 530 is used to analyze the video memory occupation life cycle of the operators in each submodule in the corresponding operator library, and determine the operator video memory reuse information and operator video memory release information; The video memory allocation unit 540 is used to allocate video memory to the target fusion module based on the operator video memory reuse information and the operator video memory release information, so that the target fusion module performs calculations in each operator library.

[0094] The neural network operator fusion device provided by the embodiment of the present invention determines a target fusion module in a neural network; the target fusion module includes multiple operators; the target fusion module is divided to obtain multiple sub-modules that perform calculations in different operator libraries; the video memory occupancy life cycle of the operators in each sub-module in the corresponding operator library is analyzed to determine the operator video memory reuse information and the operator video memory release information; based on the operator video memory reuse information and the operator video memory release information, video memory is allocated to the target fusion module so that the target fusion module performs calculations in each operator library; since the target fusion module is divided to obtain multiple sub-modules, each sub-module can perform calculations in different operator libraries, and the comprehensive use of multiple operator libraries is realized to optimize the computing performance of the target fusion module without being limited to a specified operator library, thereby improving the performance optimization level of the neural network; since the video memory occupancy life cycle of each operator is analyzed, video memory reuse can be maximized in the video memory management process of the target fusion module, and video memory release can be realized in the most timely manner, thereby reducing the demand for video memory resources of the neural network.

[0095] In some embodiments, the operator segmentation unit is used to: Determine the current operator and the next operator in the target fusion module; When the operator library corresponding to the current operator is the same as the operator library corresponding to the next operator, the current operator and the next operator are divided into the same submodule; When the operator library corresponding to the current operator is different from the operator library corresponding to the next operator, the current operator and the next operator are divided into different sub-modules.

[0096] In some embodiments, the operator segmentation unit is further configured to: When the number of target operator libraries supporting the current operator is one, the target operator library is determined to be the operator library corresponding to the current operator; When there are multiple target operator libraries supporting the current operator, the computational cost of the current operator in each target operator library is determined based on the computational characteristics of the current operator; and the target operator library corresponding to the minimum computational cost is determined as the operator library corresponding to the current operator.

[0097] In some embodiments, the operator segmentation unit is used to: When there are multiple target operator libraries corresponding to the minimum computing cost, determine the computing performance of each target operator library; The target operator library corresponding to the optimal computing performance is determined as the operator library corresponding to the current operator.

[0098] In some embodiments, the video memory analysis unit is used to: Determine the video memory required to be allocated and the video memory occupancy life cycle for each operator in each submodule to perform calculations in the corresponding operator library; When it is determined that the video memory occupation life cycle of any operator does not overlap with the video memory occupation life cycle of other operators and there is no data dependency relationship between any operator and other operators, it is determined that there is a video memory reuse relationship between any operator and other operators; Based on the memory reuse relationship between operators, operator memory reuse information is generated.

[0099] In some embodiments, the video memory analysis unit is used to: When it is determined that the output of any operator has no dependency relationship with the input of other operators, based on the video memory occupation life cycle of any operator, the video memory release information of any operator is determined; Based on the video memory release information of each operator, the operator video memory release information is generated.

[0100] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor (Processor) 610, a communication interface (Communications Interface) 620, a memory (Memory) 630 and a communication bus (Communications Bus) 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic command in the memory 630 to execute the method described in the above embodiment, for example: Determine a target fusion module in a neural network; the target fusion module includes multiple operators; divide the target fusion module to obtain multiple sub-modules that perform calculations in different operator libraries; analyze the memory occupancy life cycle of operators in each sub-module in the corresponding operator library, and determine operator memory reuse information and operator memory release information; based on the operator memory reuse information and operator memory release information, allocate memory to the target fusion module so that the target fusion module performs calculations in each operator library.

[0101] In addition, the logic commands in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several commands to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0102] The processor in the electronic device provided in the embodiment of the present invention can call the logic instructions in the memory to implement the above method. Its specific implementation method is consistent with the implementation method of the aforementioned method and can achieve the same beneficial effects, which will not be repeated here.

[0103] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.

[0104] Its specific implementation is consistent with the aforementioned method implementation and can achieve the same beneficial effects, so it will not be repeated here.

[0105] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described above is implemented.

[0106] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A neural network operator fusion method, characterized in that: include: Determine the target fusion module in the neural network; The target fusion module includes multiple operators; The target fusion module is divided into multiple submodules for performing calculations in different operator libraries; Analyze the memory usage life cycle of operators in each submodule in the corresponding operator library to determine operator memory reuse information and operator memory release information; Based on the operator memory reuse information and the operator memory release information, memory is allocated to the target fusion module, so that the target fusion module performs calculations in each operator library.

2. The neural network operator fusion method according to claim 1, characterized in that: The target fusion module is divided into multiple submodules for performing calculations in different operator libraries, including: Determining a current operator and a next operator in the target fusion module; When the operator library corresponding to the current operator is the same as the operator library corresponding to the next operator, the current operator and the next operator are divided into the same submodule; When the operator library corresponding to the current operator is different from the operator library corresponding to the next operator, the current operator and the next operator are divided into different sub-modules.

3. The neural network operator fusion method according to claim 1, characterized in that: The operator library corresponding to each operator is determined based on the following steps: When the number of target operator libraries supporting the current operator is one, determining the target operator library as the operator library corresponding to the current operator; In the case that there are multiple target operator libraries supporting the current operator, determining the computation cost of the current operator in each target operator library based on the computation characteristics of the current operator; The target operator library corresponding to the minimum computing cost is determined as the operator library corresponding to the current operator.

4. The neural network operator fusion method according to claim 3, characterized in that: The step of determining the target operator library corresponding to the minimum computing cost as the operator library corresponding to the current operator includes: When there are multiple target operator libraries corresponding to the minimum computing cost, determine the computing performance of each target operator library; The target operator library corresponding to the optimal computing performance is determined as the operator library corresponding to the current operator.

5. The neural network operator fusion method according to claim 1, characterized in that: The analyzing the memory occupation life cycle of the operators in each submodule in the corresponding operator library to determine the operator memory reuse information and operator memory release information includes: Determine the memory usage lifecycle of each operator in each submodule when performing calculations in the corresponding operator library; When it is determined that the video memory occupation life cycle of any operator does not overlap with the video memory occupation life cycle of other operators and there is no data dependency relationship between the any operator and the other operators, it is determined that there is a video memory multiplexing relationship between the any operator and the other operators; Based on the memory reuse relationship between the operators, the operator memory reuse information is generated.

6. The neural network operator fusion method according to claim 5, characterized in that: After determining the video memory occupation life cycle of each operator in each submodule when executing calculation in the corresponding operator library, the method further includes: When it is determined that the output of any operator has no dependency relationship with the input of other operators, based on the video memory occupation life cycle of the any operator, determining the video memory release information of the any operator; The operator video memory release information is generated based on the video memory release information of each operator.

7. A neural network operator fusion device, characterized in that: include: An operator determination unit, used to determine a target fusion module in a neural network; The target fusion module includes multiple operators; An operator segmentation unit, used to segment the target fusion module to obtain multiple submodules that perform calculations in different operator libraries; A memory analysis unit is used to analyze the memory occupancy life cycle of operators in each submodule in the corresponding operator library, and determine operator memory reuse information and operator memory release information; A video memory allocation unit is used to allocate video memory to the target fusion module based on the operator video memory reuse information and the operator video memory release information, so that the target fusion module performs calculations in each operator library.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the neural network operator fusion method described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the neural network operator fusion method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the neural network operator fusion method according to any one of claims 1 to 6 is implemented.