End-cloud collaborative reasoning method and device for neural network operator fusion

By converting neural networks into directed acyclic graphs and performing chain structure slicing and operator fusion, collaborative inference between edge devices and cloud servers is optimized, and the problems of large resource consumption and time extension are solved, and more efficient model inference is achieved.

CN115062784BActive Publication Date: 2025-08-19INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210666123.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-08-19
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

In the prior art, the neural network inference schemes of edge devices and cloud servers have problems such as large resource consumption and time delay, and the operator fusion technology is not fully utilized, resulting in the inadequate optimization of model division.

Method used

The neural network is converted into a directed acyclic graph, divided into a chain structure, and operator fusion is performed on each chain structure, predicting the inference time and output data size of each fusion block and the unfusion network layer, and calculating the intermediate data transmission time based on this, determining the best division point for end-cloud collaborative inference.

Benefits of technology

By optimizing model division and operator fusion, the total delay of neural network inference is reduced, and resource utilization and inference efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062784B_ABST
    Figure CN115062784B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for end-cloud collaborative reasoning for neural network operator fusion, the method comprising converting a neural network into a directed acyclic graph; dividing the directed acyclic graph into a plurality of chain structures; fusing the network layers in each chain structure, and replacing the fused network layers with the obtained fusion blocks; predicting the inference time and output data size of each fusion block and each unfused network layer based on the data to be inferred, and calculating the intermediate data transmission time based on the output data size and the network bandwidth between the end and the cloud; segmenting the neural network based on the inference time and the intermediate data transmission time, and performing end-cloud collaborative reasoning based on the segmentation results. The present invention solves the minimum latency problem of network models with fusionable operators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network operator fusion and end-cloud collaborative reasoning, and specifically to an end-cloud collaborative reasoning method and device for neural network operator fusion. Background Art

[0002] Machine learning is widely used, and deep learning models are an indispensable part of it, becoming a very important research topic in academia. Deep learning plays an important role in fields such as computer vision, pattern recognition, and natural language processing. In computer vision and object detection, network models based on deep learning are characterized by large number of parameters, long computation time, and limited bandwidth. In previous mainstream inference solutions, there are usually two methods for implementing deep learning model inference on edge / mobile devices: the first is to deploy the deep learning network model in the cloud. After the edge detects the data, it sends the data to the cloud for processing through network conditions. The cloud then feeds the data into the deep learning network model for inference, obtains the inference results, and then returns the inference results to the edge device. In such an inference scheme, since the data is usually images or videos with a large amount of data, it takes a lot of time and resources to transmit it from the edge to the cloud, resulting in certain resource consumption and time delay, and cannot well achieve the task requirements; the second is to deploy the deep learning network model on the edge, so that the edge no longer infers by transmitting data to the cloud, but directly performs neural network model inference on itself to obtain the return result. This inference scheme effectively reduces the transmission time of a large amount of data and saves resources, but the deployment of deep learning network models requires edge devices to have very strong performance. Usually, edge devices have poor computing power and small storage size, and cannot realize deep learning network inference. If new powerful devices are added to each edge to deploy deep learning network models, it will be a considerable expense. Moreover, each edge is only for its corresponding scenario and cannot play the same role as the cloud in other scenarios. Therefore, the feasibility is poor.

[0003] In recent years, the performance of edge devices has been improved to a certain extent. They can run some network tasks and divide the model to realize the reasoning acceleration of deep learning neural networks. This reasoning scheme takes into account the different output data sizes of different layers in the deep learning neural network model layer. The edge device first infers part of the model. After reasoning to a certain layer, the output data is used as the input of the next layer. The output data size of this layer is small, and it is split at this point. That is, the model reasoning layer after this layer is executed on the cloud. After the edge's own reasoning task is completed, the data is transmitted to the cloud. The cloud performs subsequent model reasoning. After the reasoning is completed, the reasoning result is returned to the edge device. Compared with the deep learning network model deployed entirely on the edge or cloud, it can improve resource utilization and model reasoning time, and reduce latency.

[0004] Operator fusion is an emerging technology for accelerating neural network inference. Edge devices typically have limited memory capacity and numerous neural network layers. Frequent transmission of intermediate computational data between layers consumes significant resources. Fusion of model operators eliminates the need for additional data exchange between the on-chip memory used for computation and external memory. Intermediate computational data between operators and the fused operators is stored directly in the on-chip memory of the AI chip, eliminating the need to store the data in external memory awaiting I / O access. This significantly saves memory bandwidth and improves model inference efficiency. However, many current partitioning algorithms fail to consider operator fusion. The optimal model partitioning point is likely to occur within the fused block, which is unfavorable and conflicts with fusion. While operator data transmission time within the actual fused block can be reduced, partitioning algorithms should also consider the fused block model of the fused operators. Currently, research on how to use partitioning algorithms to partition network models that may contain operator fusion is lacking.

[0005] In summary, in a general model inference system, due to the difference in computing power between edge devices and cloud servers, the data size output by each layer of the deep neural network is different. A model partitioning strategy can be formulated to enable the edge device to infer a part of the model, and then transmit the data after the inference model to the cloud server to complete the entire inference. The cloud server returns the inference result, such as Figure 1 As shown in the figure. In practice, the inference time of a network model includes the time it takes for the edge device to infer the model with the cloud server, as well as the time it takes for the edge device to transmit data to the cloud server. Data transmission time depends on the data size and network bandwidth. The inference time between the edge device and the cloud server includes not only the runtime of each layer of the inference model, but also the interaction time between the intermediate data between each layer of the inference model and the device memory. In common deep learning frameworks like Caffe and TensorFlow, the runtime between intermediate data and memory is not explicitly reflected in the general inference model, but it does exist and can be optimized. Summary of the Invention

[0006] To address the above problems, the present invention proposes an end-cloud collaborative reasoning method and device for neural network operator fusion, thereby solving the minimum latency problem of network models with fusionable operators.

[0007] The technical solution of the present invention includes:

[0008] A method for end-cloud collaborative reasoning for neural network operator fusion includes the following steps:

[0009] Converting the neural network into a directed acyclic graph, wherein the nodes in the directed acyclic graph represent network layers of the neural network, and the edges connecting the nodes in the directed acyclic graph represent data transmission between the network layers of the neural network;

[0010] Cutting the directed acyclic graph into a plurality of chain structures;

[0011] Performing a fusion operation on the network layers in each chain structure, and replacing the fused network layers with the obtained fusion blocks;

[0012] Based on the data to be inferred, the inference time and output data size of each fused block and each unfused network layer are predicted, and the intermediate data transmission time is calculated based on the output data size and the network bandwidth between the end and the cloud.

[0013] Based on the inference time and the intermediate data transmission time, the neural network is segmented, and end-cloud collaborative inference is performed based on the segmentation results.

[0014] Furthermore, the directed acyclic graph is divided into several chain structures, including:

[0015] Get the starting node P0 in the directed acyclic graph, and get the successor node of the starting node P0 according to the node connection edge, until a successor node P i Then there are multiple nodes connecting edges, thus forming a chain structure;

[0016] The successor node P i Each node P after i+1 They are respectively used as starting nodes to construct a chain structure until the directed acyclic graph is traversed.

[0017] Furthermore, performing a fusion operation on the network layers in each chain structure and replacing the fused network layers with the obtained fusion blocks includes:

[0018] Obtain fusion block rules, where the fusion block rules include:

[0019] The convolutional layer is fused with the batch normalization layer;

[0020] The convolution layer, batch normalization layer and activation function layer are integrated;

[0021] The convolution layer, batch normalization layer, activation function layer and pooling layer are integrated;

[0022] The batch normalization layer is fused with the activation function layer;

[0023] Batch normalization layer, activation function layer and pooling layer are integrated;

[0024] The activation function layer is integrated with the pooling layer;

[0025] fusing the network layers based on the fusion block rule;

[0026] The resulting fused block is used to replace the fused network layer.

[0027] Furthermore, the predicting of the inference time of each fusion block based on the data to be inferred includes:

[0028] When the convolutional layer is fused with the batch normalization layer, the inference time of the fusion block is Time CB =w1*x1+w2*x2+b1, where x1 represents the number of features of the input feature map corresponding to the data to be inferred, x2 represents the convolution kernel parameter, w1 represents the first weight coefficient, w2 represents the second weight coefficient, and b1 represents the first bias coefficient;

[0029] When the convolution layer, batch normalization layer and activation function layer are fused, the inference time of the fusion block is Time CBA =w3*x1+w4*x2+b2, where w3 represents the third weight coefficient, w4 represents the fourth weight coefficient, and b2 represents the second bias coefficient;

[0030] When the convolution layer, batch normalization layer, activation function layer and pooling layer are fused, the inference time of the fusion block is Time CBAP =w5*x1+w6*x2+b3, where w5 represents the fifth weight coefficient, w6 represents the sixth weight coefficient, and b3 represents the third bias coefficient;

[0031] When the batch normalization layer is fused with the activation function layer, the inference time of the fusion block is Time BA =w7*x3+b4, where x3 represents the size of the data to be inferred, w7 represents the seventh weight coefficient, and b4 represents the fourth bias coefficient;

[0032] When the batch normalization layer, activation function layer and pooling layer are fused, the inference time of the fusion block is Time BAP =w8*x4+w9*x5+b5, where x4 represents the size of the input feature map corresponding to the data to be inferred, x5 represents the size of the output feature map, w8 represents the eighth weight coefficient, w9 represents the ninth weight coefficient, and b5 represents the fifth bias coefficient;

[0033] When the activation function layer is fused with the pooling layer, the inference time of the fusion block is Time AP =w 10 *x4+w 11 *x5+b6, where w 10represents the tenth weight coefficient, w 11 represents the eleventh weight coefficient, and b6 represents the sixth bias coefficient.

[0034] Furthermore, the step of predicting the inference time of each unfused network layer based on the data to be inferred includes:

[0035] When the unfused network layer is a convolutional layer, the inference time of the unfused network layer is Time C =w 12 *x1+w 13 *x2+b7, where x1 represents the number of features of the input feature map corresponding to the data to be inferred, x2 represents the convolution kernel parameter, and w 12 represents the twelfth weight coefficient, w 13 represents the thirteenth weight coefficient, b7 represents the seventh bias coefficient;

[0036] When the unfused network layer is a batch normalization layer, the inference time of the unfused network layer is Time A =w 14 *x3+b8, where x3 represents the size of the data to be inferred, and w 14 represents the fourteenth weight coefficient, b8 represents the eighth bias coefficient;

[0037] When the unfused network layer is an activation layer, the inference time Time of the unfused network layer A =w 15 *x6+b9, where x6 represents the total number of input neurons, w 15 represents the fifteenth weight coefficient, b9 represents the ninth bias coefficient;

[0038] When the unfused network layer is a pooling layer, the inference time of the unfused network layer is Time P =w 16 *x4+w 17 *x5+b 10 , where x4 represents the size of the input feature map corresponding to the data to be inferred, x5 represents the size of the output feature map, and w 16 represents the sixteenth weight coefficient, w 17 represents the seventeenth weight coefficient, b 10 represents the tenth bias coefficient;

[0039] When the unfused network layer is a fully connected layer, the inference time of the unfused network layer is Time F =w 18 *x7+w 19 *x8+b 11, where x7 represents the number of input neurons, x8 represents the number of output neurons, and w 18 represents the eighteenth weight coefficient, w 19 represents the nineteenth weight coefficient, b 11 Indicates the eleventh bias coefficient.

[0040] Furthermore, based on the inference time and the intermediate data transmission time, the neural network is segmented, and end-cloud collaborative inference is performed based on the segmentation results, including:

[0041] By setting inference segmentation points for the neural network, a plurality of segmentation schemes are obtained;

[0042] For each segmentation scheme, calculate the total time T of the edge device inference segmentation point and the model before the segmentation point. e , the total time T required for the edge device to transmit data to the cloud server after completing the inference task t , the total time T of the model after the cloud server inference split point c and the total time T e , the total time T t With the total time T c Add them together to get the inference time of the segmentation scheme;

[0043] Select the best split point based on the inference time of each split scheme;

[0044] The optimal split point and the model before the optimal split point are deployed on the edge, and the model after the optimal split point is deployed in the cloud center to achieve end-cloud collaborative processing.

[0045] A device-cloud collaborative inference device for neural network operator fusion, comprising:

[0046] a conversion module, configured to convert the neural network into a directed acyclic graph, wherein the nodes in the directed acyclic graph represent network layers of the neural network, and the edges connecting the nodes in the directed acyclic graph represent data transmission between the network layers of the neural network;

[0047] A splitting module, used for splitting the directed acyclic graph into several chain structures;

[0048] a fusion module, configured to perform a fusion operation on the network layers in each chain structure, and use the obtained fusion block to replace the fused network layer;

[0049] The prediction module predicts the inference time and output data size of each fused block and each unfused network layer based on the data to be inferred, and calculates the intermediate data transmission time based on the output data size and the network bandwidth between the end and the cloud;

[0050] A deployment module is used to segment the neural network based on the inference time and the intermediate data transmission time, and perform end-cloud collaborative inference based on the segmentation results.

[0051] A computer-readable storage medium stores a computer program, which implements the above method when executed by a processor.

[0052] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor loads and executes the computer program to implement the above method.

[0053] A computer program product, when running on a computer device, enables the computer device to execute the above method.

[0054] Compared with the prior art, the present invention has the following advantages and effects:

[0055] Existing neural network end-cloud collaborative reasoning focuses on model partitioning selection for complex neural network models. The partitioning algorithm does not consider operator fusion, and the model partitioning points are likely to appear within the fusion block. Such a partitioning algorithm is not optimal. This invention considers a model partitioning algorithm based on operator fusion, which fuses the operators of the complex model and then partitions the simplified model after fusion. The partitioned models are deployed on the edge side and the cloud center respectively, further reducing the model's inference latency. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram of end-cloud collaborative inference.

[0057] Figure 2 A flow chart of a method according to an embodiment of the present invention.

[0058] Figure 3 Neural Network Directed Acyclic Graph.

[0059] Figure 4 Schematic diagram of neural network operator fusion. DETAILED DESCRIPTION

[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only specific embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0061] The end-cloud collaborative reasoning method of the present invention is as follows: Figure 2 As shown, the following steps are included:

[0062] Step 1: Convert the deep learning neural network graph into a directed acyclic graph.

[0063] The present invention converts a complex neural network into a directed acyclic graph, each network layer (operator) is converted into a node of the directed acyclic graph, and the data transmission between the neural network layers is converted into the node connection edge of the directed acyclic graph, such as Figure 3 shown.

[0064] Step 2: Due to the orderliness of the neural network inference model, the directed acyclic graph is divided into a multi-branched chain structure.

[0065] The principle of a multi-branch chain structure is to separate all branches. The specific steps to divide it into a multi-branch chain structure are:

[0066] (1) Traverse from the starting node of the directed acyclic graph (for example, starting node A) until there are multiple links after the node (for example, node C), as a branch chain structure (then A, B, and C are the first branch chain structure);

[0067] (2) In step (1), the multiple connecting lines of node C correspond to multiple branches. The first node of each branch is used as the starting node. Step (1) is executed for each branch to form a new chain structure.

[0068] (3) Repeat (1)-(2) until the end node of the directed acyclic graph is reached.

[0069] Step 3: Perform fusion operation on each chain structure and use the resulting fusion block as a new layer, such as Figure 4 As shown,

[0070] After being divided into multiple branches, only the network layers (operators) of each chain need to be fused. The fusion of network layers (operators) requires certain conditions. The present invention provides 6 fusion block rules. Not any two adjacent layers can be fused. When the sequence of network layers (operators) in the deep learning neural network meets the fusion rules, the network layers are fused to form a new fusion block.

[0071] The present invention provides six fusion block rules, including:

[0072] The convolutional layer is fused with the batch normalization layer;

[0073] The convolution layer, batch normalization layer and activation function layer are integrated;

[0074] The convolution layer, batch normalization layer, activation function layer and pooling layer are integrated;

[0075] The batch normalization layer is fused with the activation function layer;

[0076] Batch normalization layer, activation function layer and pooling layer are integrated;

[0077] The activation function layer is fused with the pooling layer.

[0078] Step 4: Establish a logistic regression function between the operator features and the inference time for the fusion block, and predict the inference time of each fusion block on the edge and cloud.

[0079] For the fusion block after the network structure is fused, a regression model is established. The independent variables vary depending on the type of fusion layer. Based on the six fusion block rules given above, for unified representation, the parameters of the convolution layer are: filter size represents the size of the convolution kernel, number of filters represents the number of filters, and stride represents the convolution step size. It is worth noting that the edge and cloud have different training coefficients for the logistic regression function of the fusion block, including w and b. For the same fusion block, the logistic regression function on the edge and cloud has the same form, as shown in the following formula:

[0080] (1) When the fusion block is (Convolution+BatchNorm), it means that the convolution layer and the batch normalization layer are fused, and the logistic regression function is

[0081] Time CB =w1*x1+w2*x2+b1

[0082] x1=The number of features in the input feature maps

[0083] x2=(filter size / stride) 2 *(number of filters)

[0084] Among them, Time CB is the prediction time of the fusion block (Convolution+BatchNorm), w1, w2, and b1 are the coefficients obtained by training with measured data in advance, x1 is the number of features of the input feature map, x2 is the calculated convolution kernel parameters, filter size is the size of the convolution kernel, number of filters is the number, and stride is the step size of the convolution.

[0085] (2) The fusion block is (Convolution+BatchNorm+ReLU / Sigmoid / Tanh), which means the fusion of convolution layer, batch normalization layer, and activation function layer. The logistic regression function is:

[0086] Time CBA=w3*x1+w4*x2+b2

[0087] x1=The number of features in the input feature maps

[0088] x2=(filter size / stride) 2 *(number of filters)

[0089] Among them, Time CBA is the prediction time of the fusion block (Convolution+BatchNorm+ReLU / Sigmoid / Tanh), w3, w4, and b2 are the coefficients obtained by training with measured data in advance, x1 is the number of features of the input feature map, x2 is the calculated convolution kernel parameters, filter size is the size of the convolution kernel, number of filters is the number, and stride is the step size of the convolution.

[0090] (3) The fusion block is (Convolution+BatchNorm+ReLU / Sigmoid / Tanh+Pooling), which means the fusion of convolution layer, batch normalization layer, activation function layer, and pooling layer. The logistic regression function is:

[0091] Time CBAP =w5*x1+w6*x2+b3

[0092] x1=The number of features in the input feature maps

[0093] x2=(filter size / stride) 2 *(number of filters)

[0094] Among them, Time cBAP is the prediction time of the fusion block (Convolution+BatchNorm+ReLU / Sigmoid / Tanh+Pooling), w5, w6, and b3 are the coefficients obtained by training with measured data in advance, x1 is the number of features of the input feature map, x2 is the calculated convolution kernel parameters, filter size is the size of the convolution kernel, number of filters is the number, and stride is the step size of the convolution.

[0095] (4) The fusion block is (BatchNorm+ReLU / Sigmoid / Tanh), which means the fusion of batch normalization layer and activation function layer. The logistic regression function is:

[0096] Time BA =w7*x3+b4

[0097] x3=Input data size

[0098] Among them, Time BA is the prediction time of the fusion block (BatchNorm+ReLU / Sigmoid / Tanh), w7 and b4 are the coefficients obtained by training with measured data in advance, and x3 is the size of the input data.

[0099] (5) The fusion block is (BatchNorm+ReLU / Sigmoid / Tanh+Pooling), which means the fusion of batch normalization layer, activation function layer, and pooling layer. The logistic regression function is:

[0100] Time BAP =w8*x4+w9*x5+b5

[0101] x4=The size of input feature maps

[0102] x5=The size of output feature maps

[0103] Among them, Time BAP is the prediction time of the fusion block (BatchNorm+ReLU / Sigmoid / Tanh+Pooling), w8, w9, and b5 are the coefficients obtained by training with measured data in advance, x4 is the size of the input feature map, and x5 is the size of the output feature map.

[0104] (6) The fusion block is (ReLU / Sigmoid / Tanh+Pooling), which means the fusion of activation function layer and pooling layer. The logistic regression function is:

[0105] Time AP =w 10 *x4+w 11 *x5+b6

[0106] x4=The size of input feature maps

[0107] x5=The size of output feature maps

[0108] Among them, Time AP is the prediction time of the fusion block (ReLU / Sigmoid / Tanh+Pooling), w 10 , w 11 , b6 represents the coefficient obtained by training the measured data in advance, x4 represents the size of the input feature map, and x5 represents the size of the output feature map.

[0109] The above logistic regression function is based on the fusion block of the fused network layer. Taking the Convolution layer and ReLU layer as an example, that is, after the fusion of Convolution and ReLU, the inference time of the fusion block (Convolution, ReLU) on the client side and the inference time on the cloud server are predicted.

[0110] Step 5: The unfused network layer predicts the inference time of each network layer on the edge and cloud server based on the logistic regression function.

[0111] In this example, the inference time of the neural network model layer on the edge and cloud servers is predicted by the logistic regression function.

[0112] For the prediction of network layer inference time in deep learning model inference, this paper proposes five network layer logistic regression models. It is worth noting that the edge and cloud have different training coefficients for the logistic regression models of the fusion block, including w and b. However, for the same fusion block, the logistic regression models on the edge and cloud have the same form, as follows:

[0113] (1) The model layer is the convolution layer, and the logistic regression model is a binary linear function:

[0114] Time C =w 12 *x1+w 13 *x2+b7

[0115] x1=The number of features in the input feature maps

[0116] x2=(filter size / stride) 2 *(number of filters)

[0117] Among them, Time C is the prediction time of the convolution layer, w 12 , w 13, b7 represents the coefficient obtained by training with measured data in advance, x1 represents the number of features of the input feature map, x2 represents the calculated convolution kernel parameters, filtersize represents the size of the convolution kernel, number of filters represents the number, and stride represents the step size of the convolution.

[0118] (2) The model layer is the batch normalization layer BatchNorm, and the logistic regression model is a one-dimensional linear function:

[0119] Time A =w 14 *x3+b8

[0120] x3=Input data size

[0121] Among them, Time B is the prediction time of the batch normalization layer BN, w 14 , b8 represents the coefficient obtained by training with measured data in advance, and x3 represents the size of the input data.

[0122] (3) The model layer is the activation layer, including ReLU / Sigmoid / Tanh, and the logistic regression model is a linear function:

[0123] Time A =w 15 *x6+b9

[0124] x6=The number of neurons

[0125] Among them, Time A is the prediction time of the activation layer Activation, w 15 , b9 represents the coefficient obtained by training with measured data in advance, and x6 represents the total number of input neurons.

[0126] (4) The model layer is the pooling layer, and the logistic regression model is a binary linear function:

[0127] Time P =w 16 *x4+w 17 *x5+b 10

[0128] x4=The size of input feature maps

[0129] x5=The size of output feature maps

[0130] Among them, Time P is the prediction time of the pooling layer, w 16 , w 17 , b 10 It represents the coefficient obtained by training the measured data in advance, x4 represents the size of the input feature map, and x5 represents the size of the output feature map.

[0131] (5) The model layer is a fully connected layer, and the logistic regression model is a binary linear function:

[0132] Time F =w 18 *x7+w 19 *x8+b 11

[0133] x7=The number of input neurons

[0134] x8=The number of output neurons

[0135] Among them, Time F is the prediction time of the full connection, w 18 , w 19 , b 11 represents the coefficient obtained by training with measured data in advance, x7 represents the number of input neurons, and x8 represents the number of output neurons.

[0136] Step 6: Replace the corresponding neural network layer in the original neural network model with the fusion block as a new network layer, obtain the output data size of each layer of the new model after the network layer fusion, obtain the current network bandwidth B, and thus obtain the transmission time of the intermediate data.

[0137] like Figure 4 As shown in the figure, we need to obtain the output data size of "Data", "Fusion Block 1 (including Conv, BN, ReLU, and Pooling layers)", "Fusion Block 2", "Fc", "ReLU", "Fc", and "ReLU" layers. The ratio of the output data size to the bandwidth is the transmission time of the intermediate data.

[0138] Step 7: Determine the optimal partition point of the model based on the inference time of all network layers after the above fusion and the intermediate data transmission time to achieve end-cloud fusion inference.

[0139] Assume that the number of layers in a neural network is X. Then, when layer X and the layers before it are inferred on the edge, the intermediate results output by layer X need to be transmitted to the cloud. The cloud uses the intermediate results as input to complete the network inference after layer X. In summary, the time for end-cloud collaborative inference consists of three parts: the end-side inference time T e , Intermediate data transmission time T t , cloud inference time T c . Traverse T under all the cut points of the model e ,T t ,T c The partition point corresponding to the minimum of the three is the optimal partition point. The optimal partition point and the previous model are deployed at the edge, and the model after the optimal partition point is deployed at the cloud center to realize the end-cloud collaborative reasoning of the neural network and obtain the shortest reasoning delay.

[0140] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Those skilled in the art may modify or make equivalent substitutions for the technical solutions of the present invention without departing from the spirit and scope of the present invention. The scope of protection of the present invention shall be based on the claims.

Claims

1. A method for end-cloud collaborative reasoning for neural network operator fusion, comprising the following steps: Converting the neural network into a directed acyclic graph, wherein the nodes in the directed acyclic graph represent network layers of the neural network, and the edges connecting the nodes in the directed acyclic graph represent data transmission between the network layers of the neural network; Cutting the directed acyclic graph into a plurality of chain structures; Performing a fusion operation on the network layers in each chain structure and replacing the fused network layers with the obtained fusion blocks; wherein performing a fusion operation on the network layers in each chain structure and replacing the fused network layers with the obtained fusion blocks include: Obtain fusion block rules, where the fusion block rules include: The convolutional layer is fused with the batch normalization layer; The convolution layer, batch normalization layer and activation function layer are integrated; The convolution layer, batch normalization layer, activation function layer and pooling layer are integrated; The batch normalization layer is fused with the activation function layer; Batch normalization layer, activation function layer and pooling layer are integrated; The activation function layer is integrated with the pooling layer; fusing the network layers based on the fusion block rule; Use the obtained fused block to replace the fused network layer; Based on the data to be inferred, the inference time and output data size of each fused block and each unfused network layer are predicted, and the intermediate data transmission time is calculated based on the output data size and the network bandwidth between the end and the cloud. Based on the inference time and the intermediate data transmission time, the neural network is segmented, and end-cloud collaborative inference is performed based on the segmentation results.

2. The method according to claim 1, wherein The directed acyclic graph is divided into several chain structures, including: Get the starting node P0 in the directed acyclic graph, and get the successor node of the starting node P0 according to the node connection edge, until a successor node P i Then there are multiple nodes connecting edges, thus forming a chain structure; The successor node P i Each node P after i+1 They are respectively used as starting nodes to construct a chain structure until the directed acyclic graph is traversed.

3. The method according to claim 1, wherein The method of predicting the inference time of each fusion block based on the data to be inferred includes: When the convolutional layer is fused with the batch normalization layer, the inference time of the fusion block is Time CB =w1*x1+w2*x2+b1, where x1 represents the number of features of the input feature map corresponding to the data to be inferred, x2 represents the convolution kernel parameter, w1 represents the first weight coefficient, w2 represents the second weight coefficient, and b1 represents the first bias coefficient; When the convolution layer, batch normalization layer and activation function layer are fused, the inference time of the fusion block is Time CBA =w3*x1+w4*x2+b2, where w3 represents the third weight coefficient, w4 represents the fourth weight coefficient, and b2 represents the second bias coefficient; When the convolution layer, batch normalization layer, activation function layer and pooling layer are fused, the inference time of the fusion block is Time CBAP =w5*x1+w6*x2+b3, where w5 represents the fifth weight coefficient, w6 represents the sixth weight coefficient, and b3 represents the third bias coefficient; When the batch normalization layer is fused with the activation function layer, the inference time of the fusion block is Time BA =w7*x3+b4, where x3 represents the size of the data to be inferred, w7 represents the seventh weight coefficient, and b4 represents the fourth bias coefficient; When the batch normalization layer, activation function layer and pooling layer are fused, the inference time of the fusion block is Time BAP =w8*x4+w9*x5+b5, where x4 represents the size of the input feature map corresponding to the data to be inferred, x5 represents the size of the output feature map, w8 represents the eighth weight coefficient, w9 represents the ninth weight coefficient, and b5 represents the fifth bias coefficient; When the activation function layer is fused with the pooling layer, the inference time of the fusion block is Time AP =w 10 *x4+w 11 *x5+b6, where w 10 represents the tenth weight coefficient, w 11 represents the eleventh weight coefficient, and b6 represents the sixth bias coefficient.

4. The method according to claim 1, wherein The method of predicting the inference time of each unfused network layer based on the data to be inferred includes: When the unfused network layer is a convolutional layer, the inference time of the unfused network layer is Time C =w 12 *x1+w 13 *x2+b7, where x1 represents the number of features of the input feature map corresponding to the data to be inferred, x2 represents the convolution kernel parameter, and w 12 represents the twelfth weight coefficient, w 13 represents the thirteenth weight coefficient, b7 represents the seventh bias coefficient; When the unfused network layer is a batch normalization layer, the inference time of the unfused network layer is Time A =w 14 *x3+b8, where x3 represents the size of the data to be inferred, and w 14 represents the fourteenth weight coefficient, b8 represents the eighth bias coefficient; When the unfused network layer is an activation layer, the inference time Time of the unfused network layer A =w 15 *x6+b9, where x6 represents the total number of input neurons, w 15 represents the fifteenth weight coefficient, b9 represents the ninth bias coefficient; When the unfused network layer is a pooling layer, the inference time of the unfused network layer is Time P =w 16 *x4+w 17 *x5+b 10 , where x4 represents the size of the input feature map corresponding to the data to be inferred, x5 represents the size of the output feature map, and w 16 represents the sixteenth weight coefficient, w 17 represents the seventeenth weight coefficient, b 10 represents the tenth bias coefficient; When the unfused network layer is a fully connected layer, the inference time of the unfused network layer is Time F =w 18 *x7+w 19 *x8+b 11 , where x7 represents the number of input neurons, x8 represents the number of output neurons, and w 18 represents the eighteenth weight coefficient, w 19 represents the nineteenth weight coefficient, b 11 Indicates the eleventh bias coefficient.

5. The method according to claim 1, wherein Segmenting the neural network based on the inference time and the intermediate data transmission time, and performing end-cloud collaborative inference based on the segmentation results, including: By setting inference segmentation points for the neural network, a plurality of segmentation schemes are obtained; For each segmentation scheme, calculate the total time T of the edge device inference segmentation point and the model before the segmentation point. e , the total time T required for the edge device to transmit data to the cloud server after completing the inference task t , the total time T of the model after the cloud server inference split point c and the total time T e , the total time T t With the total time T c Add them together to get the inference time of the segmentation scheme; Select the best split point based on the inference time of each split scheme; The optimal split point and the model before the optimal split point are deployed on the edge, and the model after the optimal split point is deployed in the cloud center to achieve end-cloud collaborative processing.

6. A device-cloud collaborative inference device for neural network operator fusion, comprising: a conversion module, configured to convert the neural network into a directed acyclic graph, wherein the nodes in the directed acyclic graph represent network layers of the neural network, and the edges connecting the nodes in the directed acyclic graph represent data transmission between the network layers of the neural network; A splitting module, used for splitting the directed acyclic graph into several chain structures; A fusion module is configured to perform a fusion operation on the network layers in each chain structure and replace the fused network layers with the obtained fusion blocks; wherein the performing of the fusion operation on the network layers in each chain structure and replacing the fused network layers with the obtained fusion blocks includes: Obtain fusion block rules, where the fusion block rules include: The convolutional layer is fused with the batch normalization layer; The convolution layer, batch normalization layer and activation function layer are integrated; The convolution layer, batch normalization layer, activation function layer and pooling layer are integrated; The batch normalization layer is fused with the activation function layer; Batch normalization layer, activation function layer and pooling layer are integrated; The activation function layer is integrated with the pooling layer; fusing the network layers based on the fusion block rule; Use the obtained fused block to replace the fused network layer; The prediction module predicts the inference time and output data size of each fused block and each unfused network layer based on the data to be inferred, and calculates the intermediate data transmission time based on the output data size and the network bandwidth between the end and the cloud; A deployment module is used to segment the neural network based on the inference time and the intermediate data transmission time, and perform end-cloud collaborative inference based on the segmentation results.

7. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements any one of the methods of claims 1 to 5 when executed by a processor.

8. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement any one of the methods of claims 1 to 5.

9. A computer program product, which, when run on a computer device, causes the computer device to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Robot task division-oriented end, side and cloud collaborative computing device

    CN112287609A