Compilation Method, Device, Storage Medium and Electronic Device of Neural Network Model

By transforming the pooling operation of the pooling layer in the neural network model into operations that support neural network accelerator, the problem that pooling operation cannot be implemented normally when there is no dedicated hardware in the accelerator is solved, and the effect of reducing hardware overhead and ensuring normal use of the model is achieved.

CN115756493BActive Publication Date: 2025-06-10BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211558282.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2025-06-10
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

When special hardware is not set up for the pooling layer in the neural network accelerator, the pooling operation of the pooling layer cannot be implemented normally, resulting in an increase in hardware overhead.

Method used

By obtaining the pooling method information of the pooling layer in the neural network model to be compiled, the pooling operation of the input feature map is transformed into the target operation of the transposed feature map, and the operation supporting the neural network accelerator is obtained, thereby generating the target neural network model in the compilation process.

Benefits of technology

Even if there is no dedicated hardware for the pooling layer, pooling operations can be indirectly implemented through operations equivalent to pooling operations, reducing hardware overhead, and ensuring the normal use of neural network models including pooling layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115756493B_ABST
    Figure CN115756493B_ABST
Patent Text Reader

Abstract

Disclosed are a compilation method, device, storage medium and electronic device for a neural network model. The method includes: obtaining a neural network model to be compiled; based on the pooling method information adopted by a first pooling layer in the neural network model to be compiled, transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map, to obtain a second pooling layer, where the transposed feature map is obtained by performing a dimension transposition operation on the input feature map, and the target operation is an operation supported by a neural network accelerator; based on the network layers other than the first pooling layer in the neural network model to be compiled, and the second pooling layer, generating a target neural network model through compilation processing. The present disclosure can normally implement the pooling operation on the premise that no dedicated hardware is set for the pooling layer in the neural network accelerator, which is beneficial to reducing the hardware overhead and ensuring the normal use of the neural network model including the pooling layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to artificial intelligence technology, and in particular to a method, apparatus, storage medium, and electronic device for compiling a neural network model. Background Art

[0002] In some cases, there is a pooling layer in a neural network model. To ensure the normal implementation of the pooling operation of the pooling layer, a dedicated hardware needs to be set for the pooling layer in a neural network accelerator. If no dedicated hardware is set for the pooling layer in the neural network accelerator, the pooling operation of the pooling layer cannot be normally implemented, and setting dedicated hardware for the pooling layer will increase the hardware cost. Summary of the Invention

[0003] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a method, apparatus, storage medium, and electronic device for compiling a neural network model.

[0004] According to one aspect of the embodiments of the present disclosure, a method for compiling a neural network model is provided, including:

[0005] Obtaining a neural network model to be compiled;

[0006] Based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled, transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map, obtaining a second pooling layer, where the transposed feature map is obtained by performing a dimension transposition operation on the input feature map, the channel dimension of the transposed feature map corresponds to the non-channel dimension of the input feature map, the non-channel dimension of the transposed feature map corresponds to the channel dimension of the input feature map, and the target operation is an operation supported by the neural network accelerator;

[0007] Based on the network layers other than the first pooling layer in the neural network model to be compiled and the second pooling layer, generating a target neural network model through compilation processing.

[0008] According to another aspect of the embodiments of the present disclosure, a device for compiling a neural network model is provided, including:

[0009] An obtaining module, configured to obtain a neural network model to be compiled;

[0010] A transformation module, configured to transform a pooling operation on an input feature map of the first pooling layer into a target operation on a transposed feature map based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled obtained by the acquisition module, so as to obtain a second pooling layer. The transposed feature map is obtained by performing a dimension transposition operation on the input feature map. The channel dimension of the transposed feature map corresponds to the non-channel dimension of the input feature map, and the non-channel dimension of the transposed feature map corresponds to the channel dimension of the input feature map. The target operation is an operation supported by a neural network accelerator.

[0011] A generation module, configured to generate a target neural network model through compilation processing based on the network layers other than the first pooling layer in the neural network model to be compiled obtained by the acquisition module and the second pooling layer obtained by the transformation module.

[0012] According to still another aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the above-mentioned neural network model compilation method.

[0013] According to yet another aspect of the present disclosure, there is provided an electronic device, including:

[0014] A processor;

[0015] A memory for storing executable instructions of the processor;

[0016] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-mentioned neural network model compilation method.

[0017] Based on the neural network model compilation method, device, storage medium and electronic device provided in the above embodiments of the present disclosure, during the compilation of the neural network model, the pooling operation can be transformed into an operation supported by a neural network accelerator. In this way, even if there is no dedicated hardware for the pooling layer in the neural network accelerator, the pooling operation can be indirectly implemented through an operation equivalent to the pooling operation. That is, the embodiments of the present disclosure can normally implement the pooling operation on the premise that there is no dedicated hardware for the pooling layer in the neural network accelerator, which is beneficial to reducing the hardware cost and ensuring the normal use of the neural network model including the pooling layer.

[0018] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0019] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 is a schematic diagram of the calculation method of the convolutional layer in the related art.

[0021] Figure 2-1 is one of the schematic diagrams of the calculation method of the pooling layer in the related art.

[0022] Figure 2-2 is another schematic diagram of the calculation method of the pooling layer in the related art.

[0023] Figure 2-3 is the third schematic diagram of the calculation method of the pooling layer in the related art.

[0024] Figure 3 is the schematic diagram of the principle for implementing the pooling operation in the embodiment of the present disclosure.

[0025] Figure 4 is the schematic flowchart of the compilation method of the neural network model provided by an exemplary embodiment of the present disclosure.

[0026] Figure 5 is the schematic flowchart of the compilation method of the neural network model provided by another exemplary embodiment of the present disclosure.

[0027] Figure 6 is the schematic diagram of the input feature map in the embodiment of the present disclosure.

[0028] Figure 7 is the schematic diagram of the transposed feature map in the embodiment of the present disclosure.

[0029] Figure 8 is the schematic diagram of the first set of convolutional kernels in the embodiment of the present disclosure.

[0030] Figure 9 is the schematic diagram of the target convolutional kernel group in the embodiment of the present disclosure.

[0031] Figure 10 is the schematic diagram of the grouping result of grouping the second feature map along the channel direction in the embodiment of the present disclosure.

[0032] Figure 11 is the schematic diagram of the third set of convolutional kernels in the embodiment of the present disclosure.

[0033] Figure 12 is the schematic diagram of the fifth feature map in the embodiment of the present disclosure.

[0034] Figure 13 It is a schematic structural diagram of a compilation device for a neural network model provided by an exemplary embodiment of the present disclosure.

[0035] Figure 14 It is a schematic structural diagram of a compilation device for a neural network model provided by another exemplary embodiment of the present disclosure.

[0036] Figure 15 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0037] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0038] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0039] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and do not represent any specific technical meaning, nor do they indicate an inevitable logical order between them.

[0040] It should also be understood that in the embodiments of the present disclosure, "a plurality" may refer to two or more, and "at least one" may refer to one, two or more.

[0041] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, unless otherwise clearly defined or given a contrary indication in the context, it is generally understood to be one or more.

[0042] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0043] It should also be understood that the present disclosure emphasizes the differences between the various embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be described one by one here.

[0044] At the same time, it should be understood that for the sake of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0045] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present disclosure, its application, or use.

[0046] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.

[0047] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0048] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0049] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logic, data structures, and so on, which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0050] Application Overview

[0051] Some chips can be deployed with neural network accelerators. For example, a brain processing unit (BPU) can be deployed on an artificial intelligence (AI) chip. It should be noted that neural network accelerators can be used to implement neural network models, such as neural network models for object detection.

[0052] In the process of implementing the present disclosure, the inventors found that the calculation methods of different types of network layers in a neural network model often vary, and the hardware used to implement the calculations of different types of network layers may also be different. For example, there are significant differences in the calculation methods of the convolutional layer and the pooling layer in a neural network model. Among them, the calculation method of the convolutional layer can be referred to Figure 1 , and the calculation method of the pooling layer can be referred to Figure 2-1 , Figure 2-2 or Figure 2-3 . Obviously, the calculation of the convolutional layer is performed in three dimensions: the width dimension, the height dimension, and the channel dimension, while the calculation of the pooling layer is performed in two dimensions: the width dimension and the height dimension.

[0053] Currently, in order to ensure the normal implementation of the pooling operation of the pooling layer, dedicated hardware needs to be set for the pooling layer in a neural network accelerator. Conversely, if dedicated hardware is not set for the pooling layer in the neural network accelerator, the pooling operation of the pooling layer cannot be normally implemented, thus affecting the normal use of the neural network model. And setting dedicated hardware for the pooling layer will increase the hardware cost. Therefore, it is necessary to adopt a more reasonable method to implement the pooling operation of the pooling layer.

[0054] Exemplary System

[0055] It should be noted that even if dedicated hardware is not set for the pooling layer in the neural network accelerator deployed on the chip, dedicated hardware is often set for some common network layers. For example, dedicated hardware is set for the convolutional layer, the rectified linear unit (ReLU) layer, etc. respectively. In this case, the neural network accelerator supports common operation types such as convolutional operation and ReLU operation.

[0056] In view of this, in the embodiments of the present disclosure, as Figure 3 shown, the neural network model involves two stages, namely the compilation stage and the execution stage; among them, in the compilation stage, the pooling operation can be transformed into an operation supported by the neural network accelerator; in the execution stage, the pooling operation can be indirectly implemented through the actual execution of the operation supported by the neural network accelerator. In this way, even if dedicated hardware is not set for the pooling layer in the neural network accelerator, the pooling operation can be implemented, which is beneficial to reducing the hardware cost and at the same time can ensure the normal use of the neural network model including the pooling layer.

[0057] Exemplary Method

[0058] Figure 4 is a schematic flowchart of a method for compiling a neural network model provided by an exemplary embodiment of the present disclosure. Figure 4The method shown can be applied to a compiler. Figure 4 The method shown includes step 410, step 420, and step 430, which will be described separately below.

[0059] Step 410: Obtain the neural network model to be compiled.

[0060] It should be noted that the neural network model to be compiled refers to the neural network model that needs to be compiled. The neural network model to be compiled may include multiple network layers, and the multiple network layers include but are not limited to convolutional layers, pooling layers, ReLU layers, etc.

[0061] Step 420: Based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled, transform the pooling operation of the input feature map of the first pooling layer into the target operation of the transposed feature map to obtain the second pooling layer. The transposed feature map is obtained by performing a dimension transposition operation on the input feature map. The channel dimension of the transposed feature map corresponds to the non-channel dimension of the input feature map, and the non-channel dimension of the transposed feature map corresponds to the channel dimension of the input feature map. The target operation is an operation supported by the neural network accelerator.

[0062] It should be noted that at least one pooling layer may be included in the multiple network layers included in the neural network model to be compiled, and each pooling layer in the at least one pooling layer may be used as the first pooling layer.

[0063] In step 420, the pooling method information adopted by the first pooling layer can be determined first. The pooling method information refers to the information required for the first pooling layer to implement the pooling operation. Then, referring to the pooling method information adopted by the first pooling layer, the pooling operation of the input feature map of the first pooling layer can be transformed into the target operation of the transposed feature map to obtain the second pooling layer.

[0064] It should be noted that any feature map involved in the embodiments of the present disclosure may have a channel dimension and a non-channel dimension; among them, the channel dimension may also be referred to as the C dimension, and the size of the feature map in the C dimension is the number of channels; the non-channel dimension may include a height dimension and a width dimension. The height dimension may also be referred to as the H dimension, the width dimension may also be referred to as the W dimension, and the non-channel dimension may also be referred to as the HW dimension. The size of the feature map in the H dimension is the height, and the size of the feature map in the W dimension is the width. In addition, the Batch (batch processing amount) during feature map processing can be represented as N.

[0065] Optionally, the dimension transpose operation on the input feature map can essentially be to rearrange the elements in the input feature map to change the data layout, so as to obtain a transposed feature map. The relationship between the transposed feature map and the input feature map can be as follows: all elements in the same channel of the input feature map are located in different channels of the transposed feature map after the dimension transpose operation, and the element positions of all elements in the same channel of the input feature map in the transposed feature map are corresponding; each element with corresponding element positions and located in different channels in the input feature map is located in the same channel of the transposed feature map after the dimension transpose operation. In this way, the C dimension of the transposed feature map can correspond to the HW dimension of the input feature map, the HW dimension of the transposed feature map can correspond to the C dimension of the input feature map, the sizes of the input feature map in the H dimension and the W dimension can jointly determine the size of the transposed feature map in the C dimension, and the size of the input feature map in the C dimension can determine the sizes of the transposed feature map in the H dimension and the W dimension. If the data arrangement format of the input feature map is regarded as the NHWC format, the data arrangement format of the transposed feature map can be regarded as the NCHW format.

[0066] Optionally, the target operation includes but is not limited to common operation types such as convolution operation, ReLU operation, channel maximum operation, element-wise operation, etc.

[0067] It should be noted that transforming the pooling operation into the target operation in step 420 is essentially transforming the pooling operation into an equivalent operation, that is, through the execution of the equivalent operation, the result obtained is the same as or substantially the same as the result obtained by executing the pooling operation.

[0068] Step 430: Based on the network layers other than the first pooling layer in the neural network model to be compiled, and the second pooling layer, generate a target neural network model through compilation processing.

[0069] In step 430, the compiler backend can perform compilation processing based on the network layers other than the first pooling layer in the neural network model to be compiled, and the second pooling layer obtained through operation transformation, so as to generate a binary target neural network model. The specific compilation processing method can adopt any implementable method according to actual needs, and the present disclosure will not elaborate on this.

[0070] Based on the neural network model compilation method provided in the above embodiments of the present disclosure, during the compilation of the neural network model, the pooling operation can be transformed into an operation supported by the neural network accelerator. In this way, even if there is no dedicated hardware for the pooling layer in the neural network accelerator, the pooling operation can be indirectly implemented through an operation equivalent to the pooling operation. That is, the embodiments of the present disclosure can normally implement the pooling operation on the premise that there is no dedicated hardware for the pooling layer in the neural network accelerator, which is beneficial to reducing the hardware cost and ensuring the normal use of the neural network model including the pooling layer.

[0071] Based on Figure 4 the embodiments shown, as Figure 5 shown, step 420 includes step 4201 and step 4203.

[0072] Step 4201, based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled, obtain the pooling processing type.

[0073] Optionally, the pooling method information adopted by the first pooling layer may include the pooling processing type; among them, the pooling processing type includes but is not limited to local average pooling (any Average Pooling) type, local maximum pooling (any MaxPooling) type, global average pooling (Global Average Pooling) type, global maximum pooling (Global MaxPooling) type, etc.

[0074] Optionally, when the pooling processing type is the local average pooling type or the local maximum pooling type, the pooling method information adopted by the first pooling layer may further include pooling parameters, where the pooling parameters refer to the parameters required to implement the local average pooling operation or the local maximum operation, and the pooling parameters include but are not limited to the pooling window size, padding parameter (padding), stride parameter (stride), etc.

[0075] In step 4201, the pooling processing type can be directly extracted from the pooling method information adopted by the first pooling layer.

[0076] Step 4203, according to the transformation method corresponding to the pooling processing type, transform the pooling operation of the input feature map of the first pooling layer into the target operation of the transposed feature map.

[0077] It should be noted that for the four cases where the pooling processing type is the local average pooling type, the local maximum pooling type, the global average pooling type, and the global maximum pooling type, step 4203 has different implementation manners, and these implementation manners of step 4203 all meet the following premise: the height, width, and number of channels of the input feature map are H 1, W 1 , C 1 , the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 * W 2 = C 1 , H 1 * W 1 = C 2 .

[0078] In this way, if the input feature map is a feature map with a shape of 4×5×16 as shown in Figure 6 , the transposed feature map can be a feature map with a shape of 4×4×20 as shown in Figure 7 ; if the input feature map is a feature map with a shape of 16×16×128, the transposed feature map can be a feature map with a shape of 1×128×256, or a feature map with a shape of 2×64×256, or a feature map with a shape of 4×32×256.

[0079] The following introduces these implementation manners of step 4203.

[0080] In an alternative implementation manner, step 4203 includes:

[0081] When the pooling processing type is the local average pooling type, based on the pooling method information, obtain the pooling parameters;

[0082] Based on the pooling parameters, H 1 , W 1 , determine the width W 3 and height H 3 of the output feature map of the first pooling layer;

[0083] Based on the pooling parameters, H 3 , W 3 , C 2 , the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, construct the first convolution kernel set;

[0084] Transform the pooling operation of the input feature map into a target operation, where the target operation includes: performing a convolution operation on the transposed feature map and the first convolution kernel set to obtain a first feature map; performing a dimension transpose operation on the first feature map to obtain the output feature map of the second pooling layer.

[0085] For the case where the pooling processing type is the local average pooling type, the pooling parameters can be directly obtained from the pooling method information, and then refer to the pooling parameters, H 1 , W1 , determine the width W of the output feature map of the first pooling layer 3 and height H 3 .

[0086] Assume that the pooling parameters are expressed as: Kernel = 3×3, Padding = 1×1, Stride = 2×2. Here, the 3 before the "×" in Kernel = 3×3 can represent the pooling window height pooling_kernel.h = 3, the 3 after the "×" in Kernel = 3×3 can represent the pooling window width pooling_kernel.w = 3, the 1 before the "×" in Padding = 1×1 can represent the left padding value padding.left = right padding value padding.right = 1, the 1 after the "×" in Padding = 1×1 can represent the top padding value padding.top = bottom padding value padding.down = 1, the 2 before the "×" in Stride = 2×2 can represent the stride in the height direction Stride.h = 2, and the 2 after the "×" in Stride = 2×2 can represent the stride in the width direction Stride.w = 2. The parameter meanings of pooling_kernel.h, pooling_kernel.w, padding.left, padding.right, padding.top, padding.down, Stride.h, and Stride.w appearing in the following text can be referred to the relevant introductions in this paragraph, and will not be elaborated later.

[0087] In this way, the following formula can be used to calculate W 3 and H 3 :

[0088] W 3 =(W 1 +padding.left+padding.right) / Stride.w+1

[0089] H 3 =(H 1 +padding.top+padding.down) / Stride.h+1

[0090] Next, based on the pooling parameters, H 3 , W 3 , C 2 , the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, construct the first set of convolution kernels.

[0091] Optionally, the correspondence between points in the output feature map and points in the input feature map can be determined based on the operation rules of the pooling operation. Assume the input feature map is as shown in Figure 6 shown, and the pooling parameters are expressed as: Kernel = 3×3, Padding = 1×1, Stride = 2×2. Then, the point at the first row and first column in the first channel of the output feature map (width is 2, height is 3) can correspond to the 4 points where 000, 001, 004, and 005 are located in the input feature map. The point at the first row and second column in the first channel of the output feature map can correspond to the 6 points where 001, 002, 003, 005, 006, and 007 are located in the input feature map. The point at the second row and first column in the first channel of the output feature map can correspond to the 6 points where 004, 005, 008, 009, 012, and 013 are located in the input feature map. The point at the second row and second column in the first channel of the output feature map can correspond to the 9 points where 005, 006, 007, 009, 010, 011, 013, 014, and 015 are located in the input feature map. The point at the third row and first column in the first channel of the output feature map can correspond to the 4 points where 012, 013, 016, and 017 are located in the input feature map. The point at the third row and second column in the first channel of the output feature map can correspond to the 6 points where 013, 014, 015, 017, 018, and 019 are located in the input feature map.

[0092] For any point P in the output feature map, assuming the coordinates of point P are (pooling_output.w, pooling_output.h), the following formula can be used to determine the point in the input feature map corresponding to point P:

[0093] point.x = pooling_output.h * Stride.h – padding.top

[0094] point.y = pooling_output.w * Stride.w – padding.left

[0095] x.size = pooling_kernel.h

[0096] y.size = pooling_kernel.w

[0097] Among them, (point.x, point.y) represents the coordinates of the point at the upper left corner among the points corresponding to point P, and x.size and y.size respectively represent the width and height of the image area occupied by the points corresponding to point P.

[0098] Assume that point P is the point at the first row and first column in the first channel of the output feature map. The coordinates of point P can be (0, 0), that is, pooling_output.w = pooling_output.h = 0. Continuing the example in the above text, we have:

[0099] Stride.h = Stride.w = 2

[0100] padding.left = padding.top = 1

[0101] pooling_kernel.h = pooling_kernel.w = 3

[0102] By substituting pooling_kernel.h, pooling_kernel.w, Stride.h, Stride.w, padding.left, and padding.top into the above formula for calculation, it can be determined that point.x = -1, point.y = -1, x.size = 3, and y.size = 3. In this way, the point with coordinates (-1, -1) can be used as the upper-left corner point to find a region with a height of 3 and a width of 3 in the input feature map. Obviously, this region includes 4 points, which are the 4 points where 000, 001, 004, and 005 are located in the input feature map. The 4 points where 000, 001, 004, and 005 are located in the input feature map can be used as the points corresponding to the point at the first row and first column in the first channel of the output feature map. Among them, the coordinates of the point where 000 is located can be considered as (0, 0), the coordinates of the point where 001 is located can be considered as (0, 1), the coordinates of the point where 004 is located can be considered as (1, 0), and the coordinates of the point where 005 is located can be considered as (1, 1).

[0103] Optionally, the correspondence between the points in the input feature map and the points in the transposed feature map can be determined based on the transpose rule of the dimension transpose operation. Assume the input feature map is as Figure 6 shown, and the transposed feature map is as Figure 7 shown. Then, the point at the first row and first column in the first channel of the input feature map corresponds to the point at the first row and first column in the first channel of the transposed feature map. The point at the first row and second column in the first channel of the input feature map corresponds to the point at the first row and first column in the second channel of the transposed feature map. The point at the first row and third column in the first channel of the input feature map corresponds to the point at the first row and first column in the third channel of the transposed feature map, ……, the point at the last row and last column in the first channel of the input feature map corresponds to the point at the first row and first column in the last channel of the transposed feature map.

[0104] Optionally, the pooling parameters include a pooling window width W 4 and a pooling window height H 4 ;

[0105] The first set of convolution kernels includes H 3 *W 3 convolution kernels, each convolution kernel having a width of 1, a height of 1, and a number of channels of C2;

[0106] For any one of the H 3 *W 3 points in the output feature map as the first target point, the convolution kernel in the first set of convolution kernels corresponding to the first target point is the first target convolution kernel, the first set of points in the input feature map corresponds to the first target point, and when the second set of points in the transposed feature map corresponds to the first set of points in the input feature map, the element at the element position in the first target convolution kernel that satisfies the preset condition is 1 / (H 4 *W 4 ), and the elements at the remaining element positions in the first target convolution kernel are all 0. That any element position satisfies the preset condition means that the relative position of the element position in the first target convolution kernel is the same as the relative position of a point in the second set of points in the channel direction of the transposed feature map.

[0107] In an example, the pooling parameters are expressed as: Kernel = 3×3, Padding = 1×1, Stride = 2×2, then W 4 = H 4 = 3, H 4 *W 4 = 9.

[0108] Assume the input feature map is as shown in Figure 6 , the transposed feature map is as shown in Figure 7 . Taking the point at the first row and first column in the first channel of the output feature map as the first target point, the first set of points in the input feature map corresponding to the first target point includes the 4 points where 000, 001, 004, and 005 are located in the input feature map. The second set of points in the transposed feature map corresponding to the first set of points includes: the point at the first row and first column in the first channel of the transposed feature map, the point at the first row and first column in the second channel of the transposed feature map, the point at the first row and first column in the fifth channel of the transposed feature map, and the point at the first row and first column in the sixth channel of the transposed feature map. Then, in the first target convolution kernel corresponding to the first target point, the elements at the first channel, the second channel, the fifth channel, and the sixth channel are all 1 / 9, while the elements at the remaining channels are all 0. In this way, the first set of target convolution kernels can be seen in Figure 8 as K1.

[0109] Assume the input feature map is as shown inFigure 6 As shown, the transposed feature map is as Figure 7 shown. The point at the second column of the first row in the first channel of the output feature map is used as the first target point. Then, the set of the first points in the input feature map corresponding to the first target point includes: 6 points where 001, 002, 003, 005, 006, and 007 are located in the input feature map. The set of the second points in the transposed feature map corresponding to the set of the first points includes: the point at the first column of the first row in the second channel of the transposed feature map, the point at the first column of the first row in the third channel of the transposed feature map, the point at the first column of the first row in the fourth channel of the transposed feature map, the point at the first column of the first row in the sixth channel of the transposed feature map, the point at the first column of the first row in the seventh channel of the transposed feature map, and the point at the first column of the first row in the eighth channel of the transposed feature map. Then, in the first target convolution kernel corresponding to the first target point, the elements at the second channel, the third channel, the fourth channel, the sixth channel, the seventh channel, and the eighth channel are all 1 / 9, while the elements at the remaining channels are all 0. In this way, the first target convolution kernel set can be referred to Figure 8 K2 in

[0110] In a similar way to the above two paragraphs, a total of Figure 8 6 convolution kernels K1 to K6 shown in

[0111] can be obtained. These 6 convolution kernels can form the first convolution kernel set. 2 In this way, referring to C 2 , the shape of the convolution kernel can be determined efficiently and reliably. Referring to the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, the positions of the elements in the convolution kernel that meet the preset conditions can be determined efficiently and reliably. Referring to H 4 and W 4 , the corresponding elements can be filled in the positions of the elements that meet the preset conditions, and then by filling 0 in the positions of the remaining elements, the first convolution kernel set with a shape of 6×1×1×20 can be constructed efficiently and reliably.

[0112] By performing a convolution operation on the transposed feature map and the first convolution kernel set, a first feature map can be obtained. The convolution operation here can adopt the following convolution parameters: Padding = 0×0, Stride = 1×1. Combining Figure 6 , Figure 7 , Figure 8It can be seen that the element at the first row and first column in the first channel of the first feature map is the same as the element at the first row and first column in the first channel of the output feature map of the first pooling layer. The element at the first row and first column in the second channel of the first feature map is the same as the element at the first row and second column in the first channel of the output feature map of the first pooling layer. The element at the first row and first column in the third channel of the first feature map is the same as the element at the second row and first column in the first channel of the output feature map of the first pooling layer, and so on. Therefore, by performing a dimension transposition operation on the first feature map in the same way as the dimension transposition operation on the input feature map in the above text, the output feature map of the second pooling layer can be obtained, and the output feature map of the second pooling layer is the same as the output feature map of the first pooling layer.

[0113] In this implementation, when the pooling processing type is the local average pooling type, through the construction of the first convolution kernel set, the pooling operation of the input feature map can be transformed into a target operation including a convolution operation and a dimension transposition operation. Since the neural network accelerator usually supports the convolution operation and the dimension transposition operation, this can ensure the normal implementation of the target operation, which is beneficial to the normal implementation of the pooling operation on the premise that no dedicated hardware is set for the pooling layer in the neural network accelerator, and at the same time ensure the normal use of the neural network model including the pooling layer.

[0114] In another alternative implementation, step 4203 includes:

[0115] When the pooling processing type is the local maximum pooling type, based on the pooling method information, obtain the pooling parameters;

[0116] Based on the pooling parameters, H 1 、W 1 , determine the width W 3 and height H 3 of the output feature map of the first pooling layer;

[0117] Based on the pooling parameters, H 3 、W 3 , the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, construct the second convolution kernel set;

[0118] Transform the pooling operation of the input feature map into a target operation, where the target operation includes: performing a convolution operation on the transposed feature map and the second convolution kernel set to obtain a second feature map; performing a maximum value operation on the second feature map in groups along the channel direction to obtain a third feature map; performing a dimension transposition operation on the third feature map to obtain the output feature map of the second pooling layer.

[0119] It should be noted that W 3 and H 3 The specific determination methods of, the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map can all refer to the relevant introductions in the previous embodiment, and will not be elaborated here.

[0120] Optionally, the pooling parameters include the pooling window width W 4 and the pooling window height H 4 ;

[0121] The second set of convolutional kernels includes H 3 *W 3 convolutional kernel groups, and each convolutional kernel group includes H 4 *W 4 convolutional kernels. The width of each convolutional kernel is 1, the height is 1, and the number of channels is C2;

[0122] For any one of the H 3 *W 3 points in the output feature map as the first target point, the convolutional kernel group in the second set of convolutional kernels corresponding to the first target point is the target convolutional kernel group, and any one of the convolutional kernels in the target convolutional kernel group is the second target convolutional kernel. When the first point set in the input feature map corresponds to the first target point and the second point set in the transposed feature map corresponds to the first point set in the input feature map, the element at one element position in the second target convolutional kernel is 1, and the elements at the remaining element positions in the second target convolutional kernel are all 0. Moreover, the relative position of the element with value 1 in the second target convolutional kernel is consistent with the relative position of a point in the second point set in the channel direction of the transposed feature map.

[0123] In an example, the pooling parameters are expressed as: Kernel = 3×3, Padding = 0×0, Stride = 1×1, then W 4 = H 4 = 3, H 4 *W 4 = 9.

[0124] Assume the input feature map is as Figure 6 shown, and the transposed feature map is as Figure 7As shown, the point at the first row and first column in the first channel of the output feature map is used as the first target point. Then, the first point set corresponding to the first target point in the input feature map includes 9 points where 000, 001, 002, 004, 005, 006, 008, 009, and 010 are located in the input feature map. The second point set corresponding to the first point set in the transposed feature map includes: the point at the first row and first column in the first channel of the transposed feature map, the point at the first row and first column in the second channel of the transposed feature map, the point at the first row and first column in the third channel of the transposed feature map, the point at the first row and first column in the fifth channel of the transposed feature map, the point at the first row and first column in the sixth channel of the transposed feature map, the point at the first row and first column in the seventh channel of the transposed feature map, the point at the first row and first column in the ninth channel of the transposed feature map, the point at the first row and first column in the tenth channel of the transposed feature map, and the point at the first row and first column in the eleventh channel of the transposed feature map. Then, among the 9 convolution kernels included in the target convolution kernel group corresponding to the first target point, the elements at the first channel in the first convolution kernel, the elements at the second channel in the second convolution kernel, the elements at the third channel in the third convolution kernel, the elements at the fifth channel in the fourth convolution kernel, the elements at the sixth channel in the fifth convolution kernel, the elements at the seventh channel in the sixth convolution kernel, the elements at the ninth channel in the seventh convolution kernel, the elements at the tenth channel in the eighth convolution kernel, and the elements at the eleventh channel in the ninth convolution kernel are all 1, and the remaining elements in these 9 convolution kernels are all 0. In this way, these 9 convolution kernels in the target convolution kernel group can be referred to Figure 9 K0 to K8 in

[0125] Assume the input feature map is as Figure 6 shown, and the transposed feature map is as Figure 7As shown, the point at the second column of the first row in the first channel of the output feature map is used as the first target point. Then, the first point set corresponding to the first target point in the output feature map includes: 9 points where 001, 002, 003, 005, 006, 007, 009, 010, and 011 are located in the input feature map. The second point set corresponding to the first point set in the transposed feature map includes: the point at the first column of the first row in the second channel of the transposed feature map, the point at the first column of the first row in the third channel of the transposed feature map, the point at the first column of the first row in the fourth channel of the transposed feature map, the point at the first column of the first row in the sixth channel of the transposed feature map, the point at the first column of the first row in the seventh channel of the transposed feature map, the point at the first column of the first row in the eighth channel of the transposed feature map, the point at the first column of the first row in the tenth channel of the transposed feature map, the point at the first column of the first row in the eleventh channel of the transposed feature map, and the point at the first column of the first row in the twelfth channel of the transposed feature map. Then, among the 9 convolution kernels included in the target convolution kernel group corresponding to the first target point, the elements at the second channel in the first convolution kernel, the elements at the third channel in the second convolution kernel, the elements at the fourth channel in the third convolution kernel, the elements at the sixth channel in the fourth convolution kernel, the elements at the seventh channel in the fifth convolution kernel, the elements at the eighth channel in the sixth convolution kernel, the elements at the tenth channel in the seventh convolution kernel, the elements at the eleventh channel in the eighth convolution kernel, and the elements at the twelfth channel in the ninth convolution kernel are all 1, and the remaining elements in these 9 convolution kernels are all 0.

[0126] In the above two hypothesis methods, the method of obtaining the corresponding 9 convolution kernels for the first target point in the output feature map is introduced in detail. These 9 convolution kernels can form a convolution kernel group corresponding to the first target point. For each point in the output feature map, this method can be used to obtain the corresponding convolution kernel group, and the set of these convolution kernel groups can be used as the second convolution kernel set.

[0127] In this way, referring to C 2 , the shape of the convolution kernel can be determined efficiently and reliably. By referring to the correspondence between the points in the output feature map and the points in the input feature map, as well as the correspondence between the points in the input feature map and the points in the transposed feature map, the positions of the elements that need to be filled with 1 in the convolution kernel can be determined efficiently and reliably. Then, by filling the remaining element positions in the convolution kernel with 0, a second convolution kernel set of 54×1×1×20 can be constructed efficiently and reliably.

[0128] By performing a convolution operation on the transposed feature map and the second convolution kernel set, a second feature map can be obtained. The convolution operation here can use the following convolution parameters: Padding = 0×0, Stride = 1×1. CombiningFigure 6 , Figure 7 , Figure 9 It can be seen that the element at the first row and first column in the first channel of the second feature map is 000, the element at the first row and first column in the second channel of the second feature map is 001, the element at the first row and first column in the third channel of the second feature map is 002, the element at the first row and first column in the fourth channel of the second feature map is 004, the element at the first row and first column in the fifth channel of the second feature map is 005, the element at the first row and first column in the sixth channel of the second feature map is 006, the element at the first row and first column in the seventh channel of the second feature map is 008, the element at the first row and first column in the eighth channel of the second feature map is 009, the element at the first row and first column in the ninth channel of the second feature map is 010, the element at the first row and first column in the tenth channel of the second feature map is 001, the element at the first row and first column in the eleventh channel of the second feature map is 002, the element at the first row and first column in the twelfth channel of the second feature map is 003, the element at the first row and first column in the twelfth channel of the second feature map is 005, and so on. In this way, 54 elements at the first row and first column can be obtained.

[0129] Next, the 54 elements at the first row and first column can be grouped. Specifically, as Figure 10 shown, these 54 elements can be divided into 6 groups, with each group including 9 elements. For the 9 elements in each group, the maximum value can be calculated. Thus, 6 maximum values can be obtained. These 6 maximum values can be arranged along the channel direction at the first row and first column, thereby obtaining a third feature map with 6 channels. Then, only by performing a dimension transposition operation on the third feature map in the same specific manner as the dimension transposition operation on the input feature map in the above text, so that the above 6 maximum values are all arranged on the first channel, all the elements at the first channel in the output feature map of the second pooling layer can be obtained. In this way, the complete output feature map of the second pooling layer can be obtained. The output feature map of the second pooling layer can be the same as the output feature map of the first pooling layer.

[0130] In this implementation manner, in the case where the pooling processing type is the local maximum pooling type, through the construction of the second convolution kernel set, the pooling operation of the input feature map can be transformed into a target operation including convolution operation, channel maximum value operation, and dimension transposition operation. Since the neural network accelerator usually supports convolution operation, channel maximum value operation, and dimension transposition operation, this can ensure the normal implementation of the target operation, thereby facilitating the normal implementation of the pooling operation on the premise that no dedicated hardware is set for the pooling layer in the neural network accelerator, and at the same time ensuring the normal use of the neural network model including the pooling layer.

[0131] In yet another alternative embodiment, step 4203 includes:

[0132] In the case where the pooling processing type is the global average pooling type, a third set of convolutional kernels is constructed, and the height, width, and number of channels of the third set of convolutional kernels are H 2 , W 2 , C 2 , respectively, and all elements at all element positions in the third set of convolutional kernels are 1 / C 2 ;

[0133] Transform the pooling operation of the input feature map into a target operation, where the target operation includes: performing a convolutional operation on the transposed feature map and the third set of convolutional kernels to obtain a fourth feature map; performing a dimensional transposition operation on the fourth feature map to obtain the output feature map of the second pooling layer.

[0134] Assume that the input feature map is as Figure 6 shown, and the transposed feature map is as Figure 7 shown. Then H 2 = W 2 = 4, C 2 = 20. The specific form of the third set of convolutional kernels can be referred to Figure 11 . By performing a depthwise convolutional operation on the transposed feature map and the third set of convolutional kernels, a fourth feature map can be obtained. The fourth feature map includes only one channel. The element at the first row and first column in the fourth feature map is the average value of all elements in the first channel of the input feature map. The element at the first row and second column in the fourth feature map is the average value of all elements in the second channel of the input feature map. The element at the first row and third column in the fourth feature map is the average value of all elements in the third channel of the input feature map, and so on. Therefore, only by performing a dimensional transposition operation on the fourth feature map in the same specific manner as the dimensional transposition operation on the input feature map described above to arrange all elements in the fourth feature map to different channels, the output feature map of the second pooling layer can be obtained, and the output feature map of the second pooling layer may be the same as the output feature map of the first pooling layer.

[0135] In this embodiment, in the case where the pooling processing type is the global average pooling type, through the construction of the third set of convolutional kernels, the pooling operation of the input feature map can be transformed into a target operation including a convolutional operation and a dimensional transposition operation. Since neural network accelerators generally support convolutional operations and dimensional transposition operations, this can ensure the normal implementation of the target operation, which is beneficial to the normal implementation of the pooling operation on the premise that no dedicated hardware is set for the pooling layer in the neural network accelerator, and at the same time ensure the normal use of the neural network model including the pooling layer.

[0136] In yet another alternative embodiment, step 4203 includes:

[0137] In the case where the pooling processing type is the global maximum pooling type, transform the pooling operation of the input feature map into a target operation, where the target operation includes: performing a channel maximum operation on the transposed feature map to obtain a fifth feature map; performing a dimension transpose operation on the fifth feature map to obtain the output feature map of the second pooling layer.

[0138] Assume the input feature map is as Figure 6 shown, and the transposed feature map is as Figure 7 shown. A channel maximum operation can be performed on the transposed feature map to obtain a fifth feature map. The specific form of the fifth feature map can be referred to Figure 12 . In this way, only by performing a dimension transpose operation on the fifth feature map in the same specific manner as the dimension transpose operation on the input feature map described above, so that all elements in the fifth feature map are arranged in different channels in sequence, the output feature map of the second pooling layer can be obtained. The output feature map of the second pooling layer may be the same as the output feature map of the first pooling layer.

[0139] In this embodiment, in the case where the pooling processing type is the global maximum pooling type, the pooling operation of the input feature map can be transformed into a target operation including a channel maximum operation and a dimension transpose operation. Since neural network accelerators generally support channel maximum operations and dimension transpose operations, this can ensure the normal implementation of the target operation, which is beneficial to the normal implementation of the pooling operation without specifically setting dedicated hardware for the pooling layer in the neural network accelerator, while ensuring the normal use of the neural network model including the pooling layer.

[0140] In the embodiments of the present disclosure, referring to the pooling processing type adopted by the first pooling layer, an adapted transformation method can be used to implement the transformation of the pooling operation to the target operation, so as to ensure that the target operation and the pooling operation are equivalent.

[0141] In summary, in the embodiments of the present disclosure, through data layout conversion (implemented by dimension transpose operation), and then by utilizing convolution operations, channel maximum operations, etc., flexible pooling operations (such as pooling windows of any size, arbitrary Padding, Stride, etc.) can be achieved, so as to ensure the normal use of the neural network model with relatively small hardware overhead.

[0142] Any compilation method of a neural network model provided by an embodiment of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices, servers, etc. Alternatively, any compilation method of a neural network model provided by an embodiment of the present disclosure can be executed by a processor. For example, the processor executes any compilation method of a neural network model mentioned in an embodiment of the present disclosure by calling corresponding instructions stored in a memory. This will not be elaborated below.

[0143] Exemplary Device

[0144] Figure 13 It is a schematic structural diagram of a compilation device for a neural network model provided by an exemplary embodiment of the present disclosure. Figure 13 The shown device includes an acquisition module 1310, a transformation module 1320, and a generation module 1330.

[0145] The acquisition module 1310 is configured to acquire a neural network model to be compiled;

[0146] The transformation module 1320 is configured to, based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled acquired by the acquisition module 1310, transform the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map, obtain a second pooling layer. The transposed feature map is obtained by performing a dimension transposition operation on the input feature map. The channel dimension of the transposed feature map corresponds to the non-channel dimension of the input feature map, and the non-channel dimension of the transposed feature map corresponds to the channel dimension of the input feature map. The target operation is an operation supported by a neural network accelerator;

[0147] The generation module 1330 is configured to, based on the network layers other than the first pooling layer in the neural network model to be compiled acquired by the acquisition module 1310, and the second pooling layer obtained by the transformation module 1320, generate a target neural network model through compilation processing.

[0148] In an optional example, as Figure 14 shown, the transformation module 1320 includes:

[0149] An acquisition sub-module 13201 is configured to acquire a pooling processing type based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled;

[0150] A transformation sub-module 13203 is configured to transform the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map according to a transformation method corresponding to the pooling processing type acquired by the acquisition sub-module 13201.

[0151] In an optional example, the height, width, and number of channels of the input feature map are H 1 、W 1 、C1 The height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 * W 2 = C 1 , H 1 * W 1 = C 2 ;

[0152] The transformation sub-module 13203 includes:

[0153] The first acquisition unit is used to obtain pooling parameters based on the pooling method information when the pooling processing type is local average pooling type;

[0154] The first determination unit is used to determine the width W 1 and height H 1 of the output feature map of the first pooling layer based on the pooling parameters, H 3 ; 3 ;

[0155] The first construction unit is used to output the correspondence between the points in the feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, and construct the first convolution kernel set based on the pooling parameters, H 3 , W 3 , C 2 ;

[0156] The first transformation unit is used to transform the pooling operation of the input feature map into a target operation, and the target operation includes: performing a convolution operation on the transposed feature map and the first convolution kernel set to obtain a first feature map; performing a dimension transposition operation on the first feature map to obtain the output feature map of the second pooling layer.

[0157] In an optional example,

[0158] The pooling parameters include the pooling window width W 4 and the pooling window height H 4 ;

[0159] The first convolution kernel set includes H 3 * W 3 convolution kernels, each convolution kernel has a width of 1, a height of 1, and a number of channels of C2;

[0160] Any point in the output feature map is a first target point, the convolutional kernel corresponding to the first target point in the first convolutional kernel set is a first target convolutional kernel, the first point set in the input feature map corresponds to the first target point, and when the second point set in the transposed feature map corresponds to the first point set in the input feature map, the element at the position of the element in the first target convolutional kernel that satisfies the preset condition is 1 / (H 4 *W 4 ), and the elements at the positions of the remaining elements in the first target convolutional kernel are all 0. That any element position satisfies the preset condition means that the relative position of this element position in the first target convolutional kernel is the same as the relative position of a point in the second point set in the channel direction of the transposed feature map.

[0161] In an optional example, the height, width, and number of channels of the input feature map are H 1 、W 1 、C 1 , the height, width, and number of channels of the transposed feature map are H 2 、W 2 、C 2 , H 2 *W 2 =C 1 , H 1 *W 1 =C 2 ;

[0162] The transformation sub-module 13203 includes:

[0163] A second acquisition unit, configured to, when the pooling processing type is the local maximum pooling type, acquire pooling parameters based on the pooling method information;

[0164] A second determination unit, configured to determine the width W 1 、W 1 and height H 3 of the output feature map of the first pooling layer based on the pooling parameters, H 3 ;

[0165] A second construction unit, configured to output the correspondence between the points in the feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map based on the pooling parameters, H 3 、W 3 , and construct a second convolutional kernel set;

[0166] A second transformation unit for transforming the pooling operation of the input feature map into a target operation, where the target operation includes: performing a convolution operation on the transposed feature map and a second set of convolution kernels to obtain a second feature map; grouping the second feature map along the channel dimension and performing a maximum operation on each group to obtain a third feature map; and performing a dimension transposition operation on the third feature map to obtain the output feature map of the second pooling layer.

[0167] In an optional example,

[0168] The pooling parameters include the pooling window width W 4 and the pooling window height H 4 ;

[0169] The second set of convolution kernels includes H 3 *W 3 convolution kernel groups, and each convolution kernel group includes H 4 *W 4 convolution kernels. Each convolution kernel has a width of 1, a height of 1, and a number of channels of C2;

[0170] For any one of the H 3 *W 3 points in the output feature map as the first target point, the convolution kernel group in the second set of convolution kernels corresponding to the first target point is the target convolution kernel group, and any one of the convolution kernels in the target convolution kernel group is the second target convolution kernel. When the first point set in the input feature map corresponds to the first target point and the second point set in the transposed feature map corresponds to the first point set in the input feature map, the element at one element position in the second target convolution kernel is 1, and the elements at the remaining element positions in the second target convolution kernel are all 0. Moreover, the relative position of the element with value 1 in the second target convolution kernel is consistent with the relative position of a point in the second point set in the channel dimension of the transposed feature map.

[0171] In an optional example, the height, width, and number of channels of the input feature map are H 1 , W 1 , C 1 in sequence, and the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 *W 2 = C 1 , H 1 *W 1 = C 2 ;

[0172] The transformation sub-module 13203 includes:

[0173] The third construction unit is used to construct a third set of convolution kernels when the pooling processing type is the global average pooling type. The height, width, and number of channels of the third set of convolution kernels are H 2 , W 2 , C 2 , and all elements at all element positions in the third set of convolution kernels are 1 / C 2 ;

[0174] The third transformation unit is used to transform the pooling operation of the input feature map into a target operation. The target operation includes: performing a convolution operation on the transposed feature map and the third set of convolution kernels to obtain a fourth feature map; performing a dimension transpose operation on the fourth feature map to obtain the output feature map of the second pooling layer.

[0175] In an optional example, the height, width, and number of channels of the input feature map are H 1 , W 1 , C 1 , the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 *W 2 = C 1 , H 1 *W 1 = C 2 ;

[0176] The transformation sub-module 13203 includes:

[0177] The fourth transformation unit is used to transform the pooling operation of the input feature map into a target operation when the pooling processing type is the global maximum pooling type. The target operation includes: performing a channel maximum operation on the transposed feature map to obtain a fifth feature map; performing a dimension transpose operation on the fifth feature map to obtain the output feature map of the second pooling layer.

[0178] Exemplary Electronic Device

[0179] Next, refer to Figure 15 to describe the electronic device according to the embodiments of the present disclosure. The electronic device can be any one or both of the first device and the second device, or a stand-alone device independent of them. The stand-alone device can communicate with the first device and the second device to receive the input signals collected by them.

[0180] Figure 15 illustrates a block diagram of an electronic device according to an embodiment of the present disclosure.

[0181] As Figure 15 shown, the electronic device 1500 includes one or more processors 1501 and a memory 1502.

[0182] The processor 1501 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 1500 to perform desired functions.

[0183] The memory 1502 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1501 can run the program instructions to implement the compilation method of the neural network model of various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. can also be stored in the computer-readable storage media.

[0184] In one example, the electronic device 1500 can further include: an input device 1503 and an output device 1504, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0185] For example, when the electronic device is the first device or the second device, the input device 1503 can be the above-mentioned microphone or microphone array for capturing the input signal of the sound source. When the electronic device is a stand-alone device, the input device 1503 can be a communication network connector for receiving the collected input signals from the first device and the second device.

[0186] In addition, the input device 1503 can further include, for example, a keyboard, a mouse, and so on.

[0187] The output device 1504 can output various information to the outside, including the determined distance information, direction information, etc. The output device 1504 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0188] Of course, for simplicity, Figure 15 only some of the components in the electronic device 1500 related to the present disclosure are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1500 can further include any other appropriate components.

[0189] Exemplary Computer Program Product and Computer Readable Storage Medium

[0190] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the compilation method of the neural network model according to various embodiments of the present disclosure described in the "Exemplary Method" section above in this specification.

[0191] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0192] An embodiment of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored, and the program instructions, when run by a processor, cause the processor to execute the steps in the compilation method of the neural network model according to various embodiments of the present disclosure described in the "Exemplary Method" section above in this specification.

[0193] The computer-readable storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0194] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for the purposes of illustration and easy understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0195] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms that mean "including but not limited to" and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0196] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0197] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0198] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A method for compiling a neural network model, comprising: obtaining a neural network model to be compiled; based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled, transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map, obtaining a second pooling layer, where the transposed feature map is obtained by performing a dimension transposition operation on the input feature map, the channel dimension of the transposed feature map corresponds to the non-channel dimension of the input feature map, the non-channel dimension of the transposed feature map corresponds to the channel dimension of the input feature map, the target operation is an operation supported by a neural network accelerator, and no dedicated hardware is set for the pooling layer in the neural network accelerator; based on the network layers other than the first pooling layer in the neural network model to be compiled and the second pooling layer, generating a target neural network model through compilation processing, so that during the execution stage of the neural network model, the pooling operation is implemented through the hardware set for the target operation in the neural network accelerator.

2. The method according to claim 1, wherein, the transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled includes: obtaining a pooling processing type based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled; transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map according to a transformation method corresponding to the pooling processing type.

3. The method according to claim 2, wherein, The height, width, and number of channels of the input feature map are H 1 , W 1 , C 1 , and the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 * W 2 = C 1 , H 1 * W 1 = C 2 ; the transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map according to a transformation method corresponding to the pooling processing type includes: when the pooling processing type is a local average pooling type, obtaining pooling parameters based on the pooling method information; Based on the pooling parameters, H 1 , W 1 , determine the width W 3 and height H 3 of the output feature map of the first pooling layer; Based on the pooling parameters, H 3 , W 3 , C 2 , the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, construct a first set of convolutional kernels; transforming the pooling operation of the input feature map into the target operation, where the target operation includes: performing a convolution operation on the transposed feature map and the first convolution kernel set to obtain a first feature map; performing a dimension transposition operation on the first feature map to obtain the output feature map of the second pooling layer.

4. The method according to claim 3, wherein, The pooling parameters include the pooling window width W 4 and the pooling window height H 4 ; The first set of convolutional kernels includes H 3 *W 3 convolutional kernels. Each convolutional kernel has a width of 1, a height of 1, and a number of channels of C2; Any point in the output feature map is a first target point. The convolution kernel corresponding to the first target point in the first convolution kernel set is the first target convolution kernel. When the first point set in the input feature map corresponds to the first target point and the second point set in the transposed feature map corresponds to the first point set in the input feature map, the element at the element position in the first target convolution kernel that satisfies the preset condition is 1 / (H 4 *W 4 ), and the elements at the remaining element positions in the first target convolution kernel are all 0. That any element position satisfies the preset condition means that the relative position of the element position in the first target convolution kernel is the same as the relative position of a point in the second point set in the channel direction of the transposed feature map.

5. The method according to claim 2, wherein, The height, width, and number of channels of the input feature map are H 1 , W 1 , C 1 , and the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 * W 2 = C 1 , H 1 * W 1 = C 2 ; the transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map according to a transformation method corresponding to the pooling processing type includes: when the pooling processing type is a local maximum pooling type, obtaining pooling parameters based on the pooling method information; Based on the pooling parameters, H 1 , W 1 , determine the width W 3 and height H 3 of the output feature map of the first pooling layer; Based on the pooling parameters, H 3 , W 3 , the correspondence between the points in the output feature map and the points in the input feature map, and the correspondence between the points in the input feature map and the points in the transposed feature map, construct a second set of convolution kernels; transforming the pooling operation of the input feature map into the target operation, where the target operation includes: performing a convolution operation on the transposed feature map and the second convolution kernel set to obtain a second feature map; grouping the second feature map along the channel direction and performing a maximum value operation on each group to obtain a third feature map; performing a dimension transposition operation on the third feature map to obtain the output feature map of the second pooling layer.

6. The method according to claim 5, wherein, The pooling parameters include the pooling window width W 4 and the pooling window height H 4 ; The second set of convolutional kernels includes H 3 *W 3 convolutional kernel groups, and each convolutional kernel group includes H 4 *W 4 convolutional kernels. The width of each convolutional kernel is 1, the height is 1, and the number of channels is C2; Any one of the H 3 *W 3 points in the output feature map is a first target point, the convolution kernel group corresponding to the first target point in the second convolution kernel set is a target convolution kernel group, any convolution kernel in the target convolution kernel group is a second target convolution kernel, the first point set in the input feature map corresponds to the first target point, and when the second point set in the transposed feature map corresponds to the first point set in the input feature map, the element at one element position in the second target convolution kernel is 1, the elements at the remaining element positions in the second target convolution kernel are all 0, and the relative position of the element position where the element is 1 in the second target convolution kernel is consistent with the relative position of a point in the second point set in the channel direction of the transposed feature map.

7. The method according to claim 2, wherein, The height, width, and number of channels of the input feature map are H 1 , W 1 , C 1 , and the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 * W 2 = C 1 , H 1 * W 1 = C 2 ; The step of transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map according to a transformation method corresponding to the pooling processing type includes: In the case where the pooling processing type is the global average pooling type, a third convolution kernel set is constructed, and the height, width, and number of channels of the third convolution kernel set are H 2 , W 2 , C 2 , and the elements at all element positions in the third convolution kernel set are all 1 / C 2 ; transforming the pooling operation of the input feature map into the target operation, where the target operation includes: performing a convolution operation on the transposed feature map and the third convolution kernel set to obtain a fourth feature map; performing a dimension transposition operation on the fourth feature map to obtain the output feature map of the second pooling layer.

8. The method according to claim 2, wherein, The height, width, and number of channels of the input feature map are H 1 , W 1 , C 1 , and the height, width, and number of channels of the transposed feature map are H 2 , W 2 , C 2 , H 2 * W 2 = C 1 , H 1 * W 1 = C 2 ; The step of transforming the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map according to a transformation method corresponding to the pooling processing type includes: in the case where the pooling processing type is the global maximum pooling type, transforming the pooling operation of the input feature map into the target operation, where the target operation includes: performing a channel maximum operation on the transposed feature map to obtain a fifth feature map; performing a dimension transposition operation on the fifth feature map to obtain the output feature map of the second pooling layer.

9. A compilation device for a neural network model, comprising: an acquisition module configured to acquire a neural network model to be compiled; a transformation module configured to, based on the pooling method information adopted by the first pooling layer in the neural network model to be compiled acquired by the acquisition module, transform the pooling operation of the input feature map of the first pooling layer into a target operation of a transposed feature map, to obtain a second pooling layer, where the transposed feature map is obtained by performing a dimension transposition operation on the input feature map, the channel dimension of the transposed feature map corresponds to the non-channel dimension of the input feature map, the non-channel dimension of the transposed feature map corresponds to the channel dimension of the input feature map, the target operation is an operation supported by a neural network accelerator, and no dedicated hardware is provided for the pooling layer in the neural network accelerator; a generation module configured to, based on the network layers other than the first pooling layer in the neural network model to be compiled acquired by the acquisition module and the second pooling layer obtained by the transformation module, generate a target neural network model through compilation processing, so that during the execution phase of the neural network model, the pooling operation is implemented by the hardware provided for the target operation in the neural network accelerator.

10. A computer-readable storage medium storing a computer program for executing the compilation method of the neural network model according to any one of claims 1-8 above.

11. An electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the compilation method of the neural network model according to any one of claims 1-8 above.

Citation Information

Patent Citations

  • Neural network model-based reasoning and compiling method and related products thereof

    CN113469365A

  • Instruction sequence generation method and device for neural network

    CN113762472A