Neural network model reasoning method and device

By inserting auxiliary operators into slices of neural network models and linking slices, the performance degradation of neural network processors under limited resources is solved, and more efficient computing and performance improvement is achieved.

CN120218249APending Publication Date: 2025-06-27ARM TECH CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344555.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Under the limited computing resources and internal cache of neural network processors, as the size specifications of neural network models increase, the performance of processors decreases significantly. In the prior art, parallel computing of graph structures of slicing neural network models cannot effectively improve processor performance.

Method used

By inserting auxiliary operators, including split operators and splicing operators in each slice, and linking the slices through these operators, a new graph structure is formed, and the repeated calculations when the operators in the slice are mapped to the operation unit.

Benefits of technology

Improves the computing efficiency and performance of neural network processors, and avoids performance degradation due to repeated calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218249A_ABST
    Figure CN120218249A_ABST
Patent Text Reader

Abstract

The invention provides an inference method and device of a neural network model, and relates to the technical field of artificial intelligence. The method comprises the following steps: segmenting a first graph structure of a neural network model into a plurality of slices in a preset dimension; inserting an auxiliary operator into each slice and linking each slice through the auxiliary operator to obtain a second graph structure; and loading the second graph structure to an arithmetic unit of a neural network processor for reasoning. According to the reasoning method and device of the neural network model, the auxiliary operator is inserted into each slice, each slice is linked through the auxiliary operator, and then the processed graph structure is loaded to the arithmetic unit of the neural network processor for reasoning, so that the phenomenon that the graph structure of the neural network model is directly segmented, and the reasoning efficiency of the neural network model is greatly improved is avoided. The calculation efficiency of the processor is improved and the performance of the processor is improved through repeated calculation when the operators in the slices are mapped to the operation units.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology. Specifically, the present disclosure relates to a method and apparatus for inferring a neural network model. Background Art

[0002] The size specifications of neural network models are increasing day by day. However, the computing resources and internal caches of neural network processors (NPUs) are very limited. With the computing resources and internal caches of the NPU remaining unchanged, the larger the size specifications of the neural network model, the worse the performance of the NPU.

[0003] In related technologies, when performing neural network model inference, the graph structure of the neural network model is often sliced and parallel computing is performed using the NPU to improve the performance of the processor.

[0004] However, how to slice the graph structure of the neural network model is crucial for improving the performance of the processor. In related technologies, the graph structure of the neural network model is sliced into multiple slices, and each slice is computed in parallel, but the performance of the processor is not effectively improved. Summary of the Invention

[0005] The present disclosure provides a method and apparatus for inferring a neural network model, which can solve the technical problems of low computing efficiency and poor performance of the processor in related technologies.

[0006] The technical solutions provided by the present disclosure are as follows: In a first aspect, the present disclosure provides a method for inferring a neural network model, including: Slicing a first graph structure of a neural network model into multiple slices in a preset dimension; Inserting auxiliary operators into each slice and linking each slice through the auxiliary operators to obtain a second graph structure; Loading the second graph structure into an arithmetic unit of a neural network processor for inference.

[0007] In some embodiments, the auxiliary operators include a splitting operator and a splicing operator.

[0008] In some embodiments, the inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: Inserting a splitting operator between the first layer of each slice and the input layer of the neural network model; Inserting a splicing operator between the last layer of all slices and the output layer of the neural network model.

[0009] In some embodiments, inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: When the outputs of all operators in the first graph structure of the neural network model are all one edge: For the first slice, traverse the operators of each layer sequentially from the last layer upwards, and determine the feature map input to the operator of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; For the nth slice, traverse the operators of each layer sequentially from the last layer upwards, and determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the layer above the (n - 1)th slice, and the feature map output by the operator of the layer above the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer. When it is determined that there is an overlap with the (n - 1)th slice, insert a splitting operator after the output of the operator of the layer above the (n - 1)th slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the layer above the current layer. n is an integer greater than 1.

[0010] In some embodiments, inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: When there is an operator in the first graph structure of the neural network model whose output is multiple edges: For the first slice, traverse the operators of each layer of each path sequentially from the last layer upwards, and determine the feature map input to the operator of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; when it is determined that the feature maps output by the target operator are different according to different paths, insert a splitting operator after a part of the outputs of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths and the offset values corresponding to different paths are different, insert splitting operators after all the outputs of the target operator; For the nth slice, traverse each operator layer by layer from the last layer upwards for each path. According to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the layer above the (n - 1)th slice, and the feature map output by the operator of the layer above the current layer. And when it is determined that there is an overlap with the (n - 1)th slice, insert a splitting operator after the output of the operator of the layer above the (n - 1)th slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the layer above the current layer. n is an integer greater than 1; when it is determined that the feature maps output by the target operator are different according to different paths, insert a splitting operator after a part of the output of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths and the offset values corresponding to different paths are different, insert a splitting operator after all the outputs of the target operator.

[0011] In some embodiments, the preset dimensions include at least one of the following: Batch size; Height; Number of channels; Width.

[0012] In a second aspect, the present disclosure also provides an inference device for a neural network model, including: A slicing module, configured to slice the first graph structure of the neural network model into multiple slices in a preset dimension; A processing module, configured to insert auxiliary operators in each slice and link each slice through the auxiliary operators to obtain a second graph structure; An inference module, which loads the second graph structure into the arithmetic unit of the neural network processor for inference.

[0013] In a third aspect, the present disclosure also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the inference method of the neural network model according to any one of the above first aspects.

[0014] In a fourth aspect, the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the inference method of the neural network model according to any one of the above first aspects are implemented.

[0015] In a fifth aspect, the present disclosure also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the inference method of the neural network model according to any one of the above first aspects are implemented.

[0016] The inference method and device for a neural network model provided by the present disclosure first insert auxiliary operators in each slice and link each slice through the auxiliary operators, and then load the processed graph structure into the arithmetic unit of the neural network processor for inference, avoiding the repeated calculation when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly sliced, improving the computing efficiency of the processor and enhancing the performance of the processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for description in the embodiments of the present disclosure.

[0018] Figure 1 It is a schematic flowchart of an inference method for a neural network model provided by an embodiment of the present disclosure; Figure 2 It is one of the schematic diagrams of the slice structure of the graph structure of a neural network model provided by an embodiment of the present disclosure; Figure 3 It is the second schematic diagram of the slice structure of the graph structure of a neural network model provided by an embodiment of the present disclosure; Figure 4 It is one of the schematic diagrams of the slicing process of the graph structure of a neural network model provided by an embodiment of the present disclosure; Figure 5 It is the second schematic diagram of the slicing process of the graph structure of a neural network model provided by an embodiment of the present disclosure; Figure 6 It is the third schematic diagram of the slicing process of the graph structure of a neural network model provided by an embodiment of the present disclosure; Figure 7 It is the fourth schematic diagram of the slicing process of the graph structure of a neural network model provided by an embodiment of the present disclosure; Figure 8 It is a schematic structural diagram of an inference device for a neural network model provided by an embodiment of the present disclosure; Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The embodiments of the present disclosure will be described below with reference to the drawings in the present disclosure. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure, and do not constitute limitations on the technical solutions of the embodiments of the present disclosure.

[0020] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present disclosure. It should be understood that when we say an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates that it is implemented as "A", or implemented as "B", or implemented as "A and B".

[0021] In the embodiments of the present disclosure, the terms "first", "second", etc. are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and do not limit the number of objects. For example, the first object can be one or multiple.

[0022] To make the purpose, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0023] The size specifications of neural network models are increasing day by day, but the computing resources and internal caches of NPUs are very limited. When the computing resources and internal caches of NPUs remain unchanged, the larger the size specifications of neural network models, the worse the performance of NPUs. In related technologies, the graph structure of neural network models is often sliced and parallel computing is performed using NPUs to improve the performance of the processor. The slicing of the graph structure of neural network models is limited by the local cache of NPUs. After the graph structure of the neural network model is sliced, each slice performs parallel computing, and many operators (such as Convolution, Pooling, etc.) will bring repeated computations, and the repeated computations will gradually accumulate as the neural network deepens. The overhead brought by this repeated computation reduces the benefits brought by slicing and does not effectively improve the performance of the processor.

[0024] Based on the above technical problems, embodiments of the present disclosure provide a method and device for inferring a neural network model. First, auxiliary operators are inserted into each slice and each slice is linked through the auxiliary operators, and then the processed graph structure is loaded into the arithmetic unit of the neural network processor for inference, avoiding repeated calculations when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly sliced, improving the computing efficiency of the processor and enhancing the performance of the processor.

[0025] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure are described below by describing several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be described repeatedly.

[0026] Figure 1 It is a schematic flowchart of a method for inferring a neural network model provided by an embodiment of the present disclosure. As Figure 1 shown, embodiments of the present disclosure provide a method for inferring a neural network model, including: Step 101: Slice the first graph structure of the neural network model into multiple slices in a preset dimension.

[0027] Specifically, in embodiments of the present disclosure, first, it is determined that the first graph structure of the neural network model needs to be sliced into multiple slices, and then the first graph structure of the neural network model is sliced into multiple slices in a preset dimension. The first graph structure can be understood as the original graph structure of the neural network model, that is, the graph structure before slicing.

[0028] In some embodiments, the number of slices into which the graph structure of the neural network model is sliced is determined according to the needs of the actual application scenario.

[0029] Figure 2 It is one of the schematic diagrams of the slice structure of the graph structure of a neural network model provided by an embodiment of the present disclosure. As Figure 2 shown, the graph structure of a neural network composed of Convolution and Pooling has a total of 4 layers. The input feature map is (1, 224, 224, 3) in the dimension order of N, H, W, C (all embodiments of the present disclosure represent in this order), and the output feature map is (1, 56, 56, 64). Among them, N represents the batch quantity dimension, H represents the height dimension, W represents the channel number dimension, and C represents the width dimension. The graph structure of this neural network model is sliced into 4 slices.

[0030] Figure 3 It is another schematic diagram of the slice structure of the graph structure of a neural network model provided by an embodiment of the present disclosure. As Figure 3As shown, a graph structure of a neural network composed of Eltwise, Convolution, and Pooling, with a total of 10 layers. The input feature map is (1, 224, 224, 3), and the output feature map is (1, 56, 56, 256). The graph structure of this neural network model is sliced into 3 slices.

[0031] In some embodiments, the preset dimensions include at least one of the following: Batch size; Height; Number of channels; Width.

[0032] Specifically, in the embodiments of the present disclosure, the dimensions considered when slicing the graph structure of the neural network model include at least one of the batch size, height, number of channels, and width.

[0033] For example, Figure 2 in [example context], the graph structure of this neural network model is evenly sliced into 4 slices along the H dimension, and the output feature map of the 4th layer of each slice is (1, 14, 56, 64).

[0034] For another example, Figure 3 in [example context], the graph structure of this neural network model is sliced into 3 slices along the H dimension, and the output feature maps of the 10th layer of the 3 slices are (1, 18, 56, 256), (1, 19, 56, 256), and (1, 19, 56, 256) respectively.

[0035] Step 102: Insert auxiliary operators into each slice and link each slice through the auxiliary operators to obtain a second graph structure.

[0036] Specifically, in the embodiments of the present disclosure, after slicing the first graph structure of the neural network model into multiple slices in the preset dimension, auxiliary operators are inserted into each slice and each slice is linked through the auxiliary operators to obtain a second graph structure.

[0037] In some embodiments, the auxiliary operator is used to avoid duplicate calculations when the operators in the slices are mapped to the arithmetic units after the graph structure of the neural network model is directly sliced.

[0038] The second graph structure of the neural network model can be understood as the graph structure obtained by slicing the original graph structure of the neural network model, inserting auxiliary operators into each slice, and linking each slice through the auxiliary operators.

[0039] Step 103: Load the second graph structure into the arithmetic unit of the neural network processor for inference.

[0040] Specifically, in the embodiments of the present disclosure, after determining the second graph structure, the second graph structure is loaded into the arithmetic unit of the neural network processor for inference to obtain the inference result of the neural network model.

[0041] For the inference method of the neural network model provided by the present disclosure, auxiliary operators are first inserted into each slice and each slice is linked through the auxiliary operators, and then the processed graph structure is loaded into the arithmetic unit of the neural network processor for inference, avoiding the repeated calculation when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly sliced, improving the computing efficiency of the processor and enhancing the performance of the processor.

[0042] In some embodiments, the auxiliary operators include a Crop operator and a Concat operator.

[0043] Specifically, in the embodiments of the present disclosure, the auxiliary operator itself belongs to a data transfer operation and can be offset (implemented) by a dedicated hardware unit on the NPU platform for adding data to the input and output stages of the computing unit respectively. The present disclosure will not elaborate on the calculation process of the auxiliary operator in the related art.

[0044] For the inference method of the neural network model provided by the present disclosure, a Crop operator and a Concat operator are first inserted into each slice and each slice is linked through the Crop operator and the Concat operator, and then the processed graph structure is loaded into the arithmetic unit of the neural network processor for inference, avoiding the repeated calculation when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly sliced, improving the computing efficiency of the processor and enhancing the performance of the processor.

[0045] In some embodiments, inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: Inserting a Crop operator between the first layer of each slice and the input layer of the neural network model respectively; Inserting a Concat operator between the last layer of all slices and the output layer of the neural network model.

[0046] Specifically, in the embodiments of the present disclosure, in the process of inserting auxiliary operators into each slice and linking each slice through the auxiliary operators, a Crop operator needs to be inserted between the first layer of each slice and the input layer of the neural network model respectively.

[0047] For example, for Figure 2 the graph structure of the neural network model in is evenly sliced into 4 slices along the H dimension, and the graph structure after inserting auxiliary operators into each slice and linking each slice through the auxiliary operators is as shown in Figure 6As shown, a splitting operator is inserted between the first layer (layer 1) of each slice and the input layer of the neural network model respectively.

[0048] For another example, for Figure 3 the graph structure of the neural network model in is sliced into 3 slices along the H dimension. After inserting auxiliary operators in each slice and linking each slice through the auxiliary operators, the resulting graph structure is as shown in Figure 7 As shown, a splitting operator is inserted between the first layer (layer 1) of each slice and the input layer of the neural network model respectively.

[0049] Specifically, in the embodiments of the present disclosure, during the process of inserting auxiliary operators in each slice and linking each slice through the auxiliary operators, a splicing operator needs to be inserted between the last layer of all slices and the output layer of the neural network model.

[0050] For example, for Figure 2 the graph structure of the neural network model in is evenly sliced into 4 slices along the H dimension. After inserting auxiliary operators in each slice and linking each slice through the auxiliary operators, the resulting graph structure is as shown in Figure 6 As shown, a splicing operator is inserted between the last layer (layer 4) of all slices and the output layer of the neural network model. The output feature maps of the 4th layer of each slice are all (1, 14, 56, 64). After being processed by the splicing operator, the output feature map is (1, 56, 56, 64).

[0051] For another example, for Figure 3 the graph structure of the neural network model in is sliced into 3 slices along the H dimension. After inserting auxiliary operators in each slice and linking each slice through the auxiliary operators, the resulting graph structure is as shown in Figure 7 As shown, a splicing operator is inserted between the last layer (layer 10) of all slices and the output layer of the neural network model. The output feature maps of the 10th layer of the 3 slices are (1, 18, 56, 256), (1, 19, 56, 256), and (1, 19, 56, 256) respectively. After being processed by the splicing operator, the output feature map is (1, 56, 56, 256).

[0052] The inference method of the neural network model provided by the present disclosure first inserts a splitting operator between the first layer of each slice and the input layer of the neural network model respectively, and inserts a splicing operator between the last layer of all slices and the output layer of the neural network model. Then, the processed graph structure is loaded into the arithmetic unit of the neural network processor for inference, avoiding the repeated calculation when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly sliced, improving the computing efficiency of the processor and enhancing the performance of the processor.

[0053] In some embodiments, auxiliary operators are inserted into each slice and each slice is linked through the auxiliary operators, including: When the outputs of all operators in the first graph structure of the neural network model are all an edge: For the first slice, traverse the operators of each layer from the last layer upwards in sequence, and determine the feature map input to the operator of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; For the nth slice, traverse the operators of each layer from the last layer upwards in sequence, and determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the layer above the (n - 1)th slice, and the feature map output by the operator of the layer above the current layer. When it is determined that there is an overlap with the (n - 1)th slice, a splitting operator is inserted after the output of the operator of the layer above the (n - 1)th slice, and a splicing operator is inserted before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the layer above the current layer, where n is an integer greater than 1.

[0054] Specifically, in the embodiments of the present disclosure, for the case where the outputs of all operators in the first graph structure of the neural network model are all an edge, first start from the first slice, traverse the operators of each layer from the last layer upwards in sequence, and determine the feature map input to the operator of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer.

[0055] It should be noted that: in the embodiments of the present disclosure, before inserting auxiliary operators into each slice and linking each slice through the auxiliary operators, there is no difference between each slice, and any slice can be used as the first slice to start traversing first. That is, no matter which slice starts traversing first, the method of the embodiments of the present disclosure is applicable.

[0056] For example, Figure 2 The specific specifications of the first layer Convolution operator in the first graph structure of the neural network model in are: input (in): [1, 224, 224, 3], output (out): [1, 112, 112, 64], kernel: 7x7, stride 2, dilation: 1.

[0057] The specific specifications of the second layer Pooling operator are: in: [1, 112, 112, 64], out: [1, 56, 56, 64], kernel: 5x5, stride: 2, dilation: 1.

[0058] The specific specifications of the 3rd layer Convolution operator are: in: [1, 56, 56, 64], out: [1, 56, 56, 64], kernel: 1x1, stride: 1, dilation: 1.

[0059] The specific specifications of the 4th layer Convolution operator are: in: [1, 56, 56, 64], out: [1, 56, 56, 64], kernel: 3x3, stride: 1, dilation: 1.

[0060] Take the example where the graph structure of the neural network model is sliced into n slices along the H dimension.

[0061] The starting point offset of the output of each slice is represented by the array out_offset, out_offset = [0, 1 H / n, 2 H / n,..., (n - 1) H / n].

[0062] The size (H dimension) of the output of each slice is represented by the array out_size, out_size = [H / n, H / n, H / n,..., H / n].

[0063] The starting point offset of the input of each slice is represented by the array in_offset.

[0064] The size of the input of each slice is represented by the array in_size.

[0065] The overlapping part of the input of each slice is represented by the array overlap_size.

[0066] The calculation method of in_offset is as follows: When i = 0, in_offset[0] = [0], where i is the index of the number of slices, starting from 0. That is, when the index of the number of slices is 0, it represents the 1st slice.

[0067] When i > 0, in_offset[i] = out_offset[i] kernel_h – pad_up The calculation method of in_size is as follows: in_size[i] = (size – 1) stride_h – pad + (kernel – 1) (dilation – 1) + kernel Among them, the uppermost pad = pad_up (left), the lowermost pad = pad_down (right), and in other cases, pad = 0.

[0068] As Figure 4 shown, Figure 4 is a schematic diagram of the first slice after the graph structure of the neural network model in Figure 2 is sliced. That is, when the feature map of the output of the Convolution operator in the 4th layer of the first slice is [1, 14, 56, 64], substituting into the above calculation method, the value of the H dimension of the feature map of the input of the Convolution operator in the 4th layer of the first slice is calculated as in_size[0] = 15, and then the input feature map is obtained as [1, 15, 56, 64]. Then, traversing the Convolution operator in the 3rd layer of the first slice, the Pooling operator in the 2nd layer of the first slice, the Convolution operator in the 1st layer of the first slice, and the inserted Crop operator in turn, the input feature maps of the Convolution operator in the 3rd layer of the first slice are calculated as [1, 15, 56, 64], the input feature map of the Pooling operator in the 2nd layer of the first slice is [1, 32, 112, 64], the input feature map of the Convolution operator in the 1st layer of the first slice is [1, 66, 224, 3], and the input feature map of the inserted Crop operator is [1, 224, 224, 3].

[0069] Starting from the second slice, for each slice, traverse the operators of each layer from the last layer upwards. According to the feature map of the output of the operator of the current layer and the calculation method of the operator of the current layer, determine the feature map of the input of the operator of the current layer, the overlapping feature map with the operator of the upper layer of the previous slice, and the feature map of the output of the operator of the upper layer of the current layer. And when it is determined that there is an overlap with the (n - 1)th slice, insert a splitting operator after the output of the operator of the upper layer of the previous slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the upper layer of the current layer.

[0070] The calculation method of overlap_size is as follows: overlap_size[0] = 0 (when i = 0) overlap_size[i] = out_offset[i - 1] + out_size[i - 1] - out_offset[i] (when i > 0).

[0071] As Figure 5 shown,Figure 5 As the schematic diagram of the second slice after the graph structure of the neural network model in Figure 2 is sliced, that is, when the feature map of the output of the Convolution operator in the fourth layer of the second slice is determined to be [1, 14, 56, 64], substituting into the above calculation method, the value of the H dimension of the feature map of the input of the Convolution operator in the fourth layer of the second slice is calculated as in_size[1]=16, and then the input feature map is obtained as [1, 16, 56, 64], and the value of the H dimension of the overlapping feature map with the operator of the previous layer of the first slice is obtained as overlap_size[1]=2, and then the overlapping feature map with the operator of the previous layer of the first slice is obtained as [1, 2, 56, 64]. A splitting operator is inserted after the output of the Convolution operator in the third layer of the first slice, and a splicing operator is inserted before the input of the Convolution operator in the fourth layer of the second slice. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the previous layer of the current layer, and the feature map of the output of the Convolution operator in the third layer of the second slice can be calculated as [1, 14, 56, 64].

[0072] After traversing all the operators of all the slices, the graph structure of the obtained neural network model is as Figure 6 shown. For the calculation processes of other operators in the embodiments of the present disclosure, refer to the calculation processes of the Convolution operator in the fourth layer of the first slice and the Convolution operator in the fourth layer of the second slice above, which will not be elaborated here.

[0073] The inference method of the neural network model provided by the present disclosure first inserts Crop operators and Concat operators in each slice and links each slice through the Crop operators and Concat operators, and then loads the processed graph structure into the arithmetic unit of the neural network processor for inference, avoiding the repeated calculation when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly sliced, improving the calculation efficiency of the processor and enhancing the performance of the processor.

[0074] In some embodiments, inserting auxiliary operators in each slice and linking each slice through the auxiliary operators includes: When there are multiple edges in the output of the target operator in the first graph structure of the neural network model: For the first slice, traverse each operator layer by layer from the last layer upwards for each path. According to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, determine the feature map input to the operator of the current layer; when it is determined that the feature maps output by the target operator are different according to different paths, insert a splitting operator after a partial output of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths and the offset values corresponding to different paths are different, insert a splitting operator after all outputs of the target operator. For the nth slice, traverse each operator layer by layer from the last layer upwards for each path. According to the feature map output by the operator of the current layer, the calculation method of the operator of the current layer, determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the layer above the (n - 1)th slice, and the feature map output by the operator of the layer above the current layer. And when it is determined that there is an overlap with the (n - 1)th slice, insert a splitting operator after the output of the operator of the layer above the (n - 1)th slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the layer above the current layer, where n is an integer greater than 1; when it is determined that the feature maps output by the target operator are different according to different paths, insert a splitting operator after a partial output of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths and the offset values corresponding to different paths are different, insert a splitting operator after all outputs of the target operator.

[0075] Specifically, in this embodiment of the present disclosure, for the case where the output of the target operator in the first graph structure of the neural network model has multiple edges. For example, Figure 3 in the first graph structure of the neural network model, the output of the 2nd layer Pooling operator has two edges, and the output of the 6th layer Eltwise operator has two edges. First, for the first slice, traverse each operator layer by layer from the last layer upwards for each path. According to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, determine the feature map input to the operator of the current layer.

[0076] It should be noted that: in this embodiment of the present disclosure, before inserting auxiliary operators in each slice and linking each slice through the auxiliary operators, there is no difference between each slice, and any slice can be used as the first slice to start traversing first. That is, no matter which slice starts traversing first, the method of this embodiment of the present disclosure is applicable. Figure 7 In, an example is given with the slice whose output feature map of the 10th layer is (1, 18, 56, 256) as the first slice.

[0077] For example, Figure 3The specific specifications of the 1st layer Convolution operator in the first graph structure of the neural network model are: in: [1, 224, 224, 3], out: [1, 112, 112, 64], kernel: 7x7, stride 2, dilation: 1.

[0078] The specific specifications of the 2nd layer Pooling operator are: in: [1, 112, 112, 64], out: [1, 56, 56, 64], kernel: 5x5, stride: 2, dilation: 1.

[0079] The specific specifications of the 3rd layer Convolution operator are: in: [1, 56, 56, 64], out: [1, 56, 56, 64], kernel: 1x1, stride: 1, dilation: 1.

[0080] The specific specifications of the left Convolution operator in the 4th layer are: in: [1, 56, 56, 64], out: [1, 56, 56, 256], kernel: 1x1, stride: 1, dilation: 1.

[0081] The specific specifications of the right Convolution operator in the 4th layer are: in: [1, 56, 56, 64], out: [1, 56, 56, 64], kernel: 3x3, stride: 1, dilation: 1.

[0082] The specific specifications of the 5th layer Convolution operator are: in: [1, 56, 56, 64], out: [1, 56, 56, 256], kernel: 1x1, stride: 1, dilation: 1.

[0083] The specific specifications of the 6th layer Eltwise operator are: in: [1, 56, 56, 256], [1, 56, 56, 256], out: [1, 56, 56, 256].

[0084] The specific specifications of the 7th layer Convolution operator are: in: [1, 56, 56, 256], out: [1, 56, 56, 64], kernel: 1x1, stride: 1, dilation: 1.

[0085] The specific specifications of the 8th layer Convolution operator are: in: [1, 56, 56, 64], out: [1, 56, 56, 64], kernel: 3x3, stride: 1, dilation: 1.

[0086] The specific specifications of the 9th layer Convolution operator are: in: [1, 56, 56, 64], out: [1, 56, 56, 256], kernel: 1x1, stride: 1, dilation: 1.

[0087] The specific specifications of the 10th layer Eltwise operator are: in: [1, 56, 56, 256], [1, 56, 56, 256], out: [1, 56, 56, 256].

[0088] As Figure 7 shown, Figure 7 For Figure 3 the schematic diagram after the graph structure of the neural network model in is sliced and completed, for the first slice, traverse each layer of operators on two paths in turn from the 10th layer Eltwise operator upwards, and determine the feature map input to the operator of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; when determining that the feature map output by the 6th layer Eltwise operator is [1, 18, 56, 256] according to the left path and the feature map output by the 6th layer Eltwise operator is [1, 19, 56, 256] according to the right path, and the two are different, at this time, a split operator needs to be inserted after one output of the 6th layer Eltwise operator, that is, a split operator is inserted in the left path in the graph. For the calculation processes of other operators in the embodiments of the present disclosure, refer to Figure 6 the calculation processes of the 4th layer Convolution operator of the first slice and the 4th layer Convolution operator of the second slice in, which will not be elaborated here.

[0089] Starting from the second slice, for each slice, traverse each layer of operators on each path in turn from the last layer upwards, and determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the previous layer of the previous slice, and the feature map output by the operator of the previous layer of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, and when it is determined that there is an overlap with the (n - 1)th slice, a split operator is inserted after the output of the operator of the previous layer of the previous slice, and a splicing operator is inserted before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted split operator and the feature map output by the operator of the previous layer of the current layer.

[0090] As Figure 7As shown, when the feature map of the output of the Convolution operator in the 8th layer of the 2nd slice is determined to be [1, 19, 56, 64], substituting into the above calculation method, the value of the H dimension of the feature map of the input of the Convolution operator in the 8th layer of the 2nd slice, in_size[1]=21, is calculated. Furthermore, the input feature map is obtained as [1, 21, 56, 64], and the value of the H dimension of the overlapping feature map with the operator of the previous layer of the 1st slice, overlap_size[1]=2, is obtained. Furthermore, the overlapping feature map with the operator of the previous layer of the 1st slice is obtained as [1, 2, 56, 64]. A split operator is inserted after the output of the Convolution operator in the 7th layer of the 1st slice, and a splicing operator is inserted before the input of the Convolution operator in the 8th layer of the 2nd slice. The inserted splicing operator is used to splice the feature map output by the inserted split operator and the feature map output by the operator of the previous layer of the current layer, and the feature map of the output of the Convolution operator in the 7th layer of the 2nd slice can be calculated as [1, 19, 56, 64].

[0091] In addition, when the feature map of the output of the Eltwise operator in the 6th layer is determined to be [1, 19, 56, 256] according to the left path and the feature map of the output of the Eltwise operator in the 6th layer is determined to be [1, 19, 56, 256] according to the right path, and the two are the same. However, since the offset values corresponding to different paths are different, split operators are inserted after the two outputs of the Eltwise operator in the 6th layer respectively.

[0092] After traversing all the operators of all paths of all slices, the graph structure of the obtained neural network model is as Figure 7 shown. For the calculation processes of other operators in the embodiments of the present disclosure, refer to Figure 6 the calculation processes of the Convolution operator in the 4th layer of the 1st slice and the Convolution operator in the 4th layer of the 2nd slice, which will not be elaborated here.

[0093] The inference method of the neural network model provided by the present disclosure first inserts Crop operators and Concat operators in each slice and links each slice through the Crop operators and Concat operators, and then loads the processed graph structure into the arithmetic unit of the neural network processor for inference, avoiding the repeated calculation when the operators in the slices are mapped to the arithmetic unit after the graph structure of the neural network model is directly segmented, improving the calculation efficiency of the processor and enhancing the performance of the processor.

[0094] Figure 8The following is a schematic structural diagram of an inference device for a neural network model provided by an embodiment of the present disclosure. An embodiment of the present disclosure also provides an inference device 80 for a neural network model, including: A slicing module 801 is configured to slice a first graph structure of a neural network model into multiple slices in a preset dimension; A processing module 802 is configured to insert auxiliary operators into each slice and link each slice through the auxiliary operators to obtain a second graph structure; An inference module 803 is configured to load the second graph structure into an arithmetic unit of a neural network processor for inference.

[0095] In some embodiments, the auxiliary operators include a splitting operator and a splicing operator.

[0096] In some embodiments, the inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: Inserting a splitting operator between the first layer of each slice and the input layer of the neural network model respectively; Inserting a splicing operator between the last layer of all slices and the output layer of the neural network model.

[0097] In some embodiments, the inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: When the output of all operators in the first graph structure of the neural network model is one edge: For the first slice, traverse the operators of each layer from the last layer upwards, and determine the feature map input to the operator of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; For the nth slice, traverse the operators of each layer from the last layer upwards, and determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the upper layer of the (n - 1)th slice, and the feature map output by the operator of the upper layer of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer. When it is determined that there is an overlap with the (n - 1)th slice, insert a splitting operator after the output of the operator of the upper layer of the (n - 1)th slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the upper layer of the current layer, where n is an integer greater than 1.

[0098] In some embodiments, the inserting auxiliary operators into each slice and linking each slice through the auxiliary operators includes: When there is an operator in the first graph structure of the neural network model whose output is multiple edges: For the first slice, traverse each operator layer by layer from the last layer upwards for each path. Based on the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, determine the feature map input to the operator of the current layer. When it is determined that the feature maps output by the target operator are different according to different paths, insert a splitting operator after a partial output of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths and the offset values corresponding to different paths are different, insert a splitting operator after all outputs of the target operator. For the nth slice, traverse each operator layer by layer from the last layer upwards for each path. Based on the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, determine the feature map input to the operator of the current layer, the overlapping feature map with the operator of the layer above the (n - 1)th slice, and the feature map output by the operator of the layer above the current layer. And when it is determined that there is an overlap with the (n - 1)th slice, insert a splitting operator after the output of the operator of the layer above the (n - 1)th slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted splitting operator and the feature map output by the operator of the layer above the current layer, where n is an integer greater than 1; when it is determined that the feature maps output by the target operator are different according to different paths, insert a splitting operator after a partial output of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths and the offset values corresponding to different paths are different, insert a splitting operator after all outputs of the target operator.

[0099] In some embodiments, the preset dimensions include at least one of the following: Batch size; Height; Number of channels; Width.

[0100] The device according to the embodiments of the present disclosure can execute the method provided by the embodiments of the present disclosure, and its implementation principle is similar and has corresponding technical effects. The actions performed by each module in the device according to the embodiments of the present disclosure correspond to the steps in the method according to the embodiments of the present disclosure. For the detailed function description of each module of the device, reference can specifically be made to the description in the corresponding method shown above, and details are not described herein again.

[0101] An electronic device is provided in an embodiment of the present disclosure, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method provided by any optional embodiment of the present disclosure.

[0102] In an optional embodiment, an electronic device is provided, as Figure 9 shown. Figure 9The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present disclosure.

[0103] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure content of the present disclosure. The processor 4001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0104] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0105] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0106] The memory 4003 is used to store the computer program for implementing the embodiments of the present disclosure and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0107] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0108] The embodiments of the present disclosure also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0109] It should be understood that although the flowcharts in the embodiments of the present disclosure indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.

[0110] The above are only alternative implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present disclosure, other similar implementation means based on the technical idea of the present disclosure also belong to the protection scope of the embodiments of the present disclosure.

Claims

1. A method for reasoning a neural network model, characterized in that: include: The first graph structure of the neural network model is divided into multiple slices in a preset dimension; Insert an auxiliary operator into each slice and link each slice through the auxiliary operator to obtain a second graph structure; The second graph structure is loaded into a computing unit of a neural network processor for reasoning.

2. The inference method of the neural network model according to claim 1, characterized in that: The auxiliary operators include a splitting operator and a splicing operator.

3. The inference method of the neural network model according to claim 2, characterized in that: The inserting of the auxiliary operator into each slice and linking each slice through the auxiliary operator includes: Insert a split operator between the first layer of each slice and the input layer of the neural network model; A concatenation operator is inserted between the last layer of all slices and the output layer of the neural network model.

4. The inference method of the neural network model according to claim 2, characterized in that: The inserting of the auxiliary operator into each slice and linking each slice through the auxiliary operator includes: When the output of all operators in the first graph structure of the neural network model is an edge: For the first slice, traverse the operators of each layer from the last layer upwards, and determine the feature map of the operator input to the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; For the nth slice, traverse the operators of each layer from the last layer upwards, and determine the feature map of the operator input to the current layer, the overlapping feature map with the operator of the previous layer of the n-1th slice, and the feature map output by the operator of the previous layer of the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer. If it is determined that there is an overlap with the n-1th slice, insert a split operator after the output of the operator of the previous layer of the n-1th slice, and insert a splicing operator before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted split operator and the feature map output by the operator of the previous layer of the current layer. n is an integer greater than 1.

5. The inference method of the neural network model according to claim 2, characterized in that: The inserting of the auxiliary operator into each slice and linking each slice through the auxiliary operator includes: In the case where the output of the target operator in the first graph structure of the neural network model is multiple edges: For the first slice, traverse the operators of each layer of each path from the last layer upwards, and determine the feature map of the operator input to the current layer according to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer; when the feature maps output by the target operator determined according to different paths are different, insert a split operator after part of the output of the target operator; when the feature maps output by the target operator determined according to different paths are the same and the offset values ​​corresponding to different paths are different, insert a split operator after all the outputs of the target operator; For the nth slice, the operators of each layer of each path are traversed from the last layer upwards in turn. According to the feature map output by the operator of the current layer and the calculation method of the operator of the current layer, the feature map of the operator input to the current layer, the overlapping feature map with the operator of the previous layer of the n-1th slice, and the feature map output by the operator of the previous layer of the current layer are determined. When it is determined that there is an overlap with the n-1th slice, a split operator is inserted after the output of the operator of the previous layer of the n-1th slice, and a splicing operator is inserted before the input of the operator of the current layer. The inserted splicing operator is used to splice the feature map output by the inserted split operator and the feature map output by the operator of the previous layer of the current layer, and n is an integer greater than 1; when it is determined that the feature maps output by the target operator are different according to different paths, a split operator is inserted after part of the output of the target operator; when it is determined that the feature maps output by the target operator are the same according to different paths, and the offset values ​​corresponding to different paths are different, a split operator is inserted after all outputs of the target operator.

6. The inference method of the neural network model according to any one of claims 1 to 5, characterized in that: The preset dimension includes at least one of the following: Batch quantity; high; Number of channels; width.

7. An inference device for a neural network model, characterized in that: include: A segmentation module, used for segmenting the first graph structure of the neural network model into multiple slices in a preset dimension; A processing module, used for inserting an auxiliary operator into each slice and linking each slice through the auxiliary operator to obtain a second graph structure; An inference module is used to load the second graph structure into a computing unit of a neural network processor for inference.

8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the inference method of the neural network model as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the inference method of the neural network model described in any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the inference method of the neural network model described in any one of claims 1 to 6 are implemented.