A Neural Network Edge-Cloud Collaborative Inference Method and Device for High Sampling Rate Video Stream Analysis
By modeling the deep learning model as a chain graph and adopting a periodic cyclic division strategy in the high-sampling rate video stream analysis scenario, the problem that the existing technology is difficult to improve system throughput without reducing the sampling rate is solved, and the system's ultimate throughput and real-time inference is achieved.
Patent Information
- Application Number
- CN202210369401.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-08
AI Technical Summary
When facing high sampling rate video streams, existing neural network-end cloud collaborative inference methods are difficult to maximize system throughput without reducing the sampling rate and meet the requirements of real-time inference.
By modeling the deep learning model as a chain graph, the data transmission edges between the model layers are divided, and the edge end and the cloud perform neural network layered delay and energy consumption evaluation based on the hardware environment and data sets. The periodic cyclic division strategy is adopted to switch the corresponding segmented edges when each frame of video stream arrives to achieve the optimal periodic cyclic division strategy.
In the high sampling rate video stream analysis scenario, the system's ultimate throughput is improved, the inference waiting time is reduced, the computing power at the edge and cloud is fully utilized, and the requirements of real-time inference are met.
Smart Images

Figure CN114723058B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep neural network edge-cloud collaborative inference acceleration optimization, and particularly relates to a neural network edge-cloud collaborative inference method for high-sampling-rate video stream analysis. Background Art
[0002] The latest progress in deep neural networks (DNNs) has greatly improved the accuracy and speed of computer vision and video analysis, creating new avenues for a new generation of intelligent applications. The maturity of cloud computing, equipped with powerful hardware such as TPUs and GPUs, has become a typical choice for such computationally intensive DNN tasks. For example, in the application of autonomous vehicles, cameras continuously monitor the surrounding scenes and transmit them to the server, which then performs video analysis and feeds back control signals to the pedals and steering wheel. In an augmented reality application, smart glasses continuously record their current views and transmit the information stream to the cloud server, which performs object recognition and sends back context-enhanced labels for seamless display on the actual scene.
[0003] The DNN network consists of multiple network layers, and the DNN that operates on each frame separately using the feed-forward algorithm is used to infer the video. The algorithm starts from the input layer and progresses layer by layer. Each layer receives the output of the previous layer as input, performs a series of calculations on the input data to obtain the output, and provides its output to the subsequent layer. Once the calculation of the output layer is completed, the process terminates. However, in the inference process of the DNN, data is generated at the edge, and data frames enter the DNN as raw inputs. The calculations of each layer in the DNN can be performed at the edge or in the cloud. The computing layer of the edge device does not need to transmit data to the cloud, but due to limited device resources, the amount of computation and computation time increase. The cloud computing layer can reduce the amount of computation, but there will be a transmission delay in the process of transmitting data from the edge device to the cloud. How to perform DNN inference at the edge and in the cloud to achieve the shortest latency or the maximum throughput is the main problem at present.
[0004] Since the data size of some intermediate DNN layers is significantly smaller than the original input data, most of the current neural network edge-cloud collaborative inference methods divide the deep neural network model, establish a mathematical model for the sum of the divided total time, and solve its optimal value. These optimal solutions are all based on the low load of the system, that is, the input rate of the data stream meets certain system requirements. If the input rate is too large, it is required to reduce the input rate. This limitation of (sampling rate) conditions obviously does not conform to the actual application scenarios. For example, at highway entrances and toll stations, it is necessary to identify whether the vehicle is legal and record the entry and exit times. The license plate recognition system has greatly improved the charging efficiency and communication costs. When the traffic flow is large, if the system is not adjusted, it will cause some license plates not to be recognized and pass through the station, resulting in system chaos; if the entry and exit rates are restricted, it will cause vehicle congestion at the entrance and exit, and it is easy to occur jams and traffic accidents. The neural network edge-cloud collaborative inference method for high sampling rate video stream analysis can maximize the system throughput without reducing the sampling rate, meeting the requirements for inference real-time performance and real-world scenario applications. Summary of the Invention
[0005] The purpose of the present invention is to provide a neural network edge-cloud collaborative inference method and device for high sampling rate video stream analysis in view of the deficiencies of the prior art.
[0006] A neural network edge-cloud collaborative inference method for high sampling rate video stream analysis of the present invention includes the following steps:
[0007] Step 1, in the scenario of a neural network architecture for high sampling rate video stream analysis, model the deep learning model as a chain diagram (line), where the vertices of the chain diagram represent the model layers of the neural network, and the arrows of the chain diagram represent the data transmission between the neural network model layers.
[0008] Step 2, perform segmentation between any two adjacent nodes in the chain diagram, and denote the segmentation edge as t. The node on the left end of t represents that the node performs inference at the edge, and the node on the right end of t represents that the node performs inference in the cloud. The segmentation edge t represents the time required for the data at the left end point of the segmentation edge to be transmitted to the right end point, and it is also the time required for the data to be transmitted from the edge to the cloud.
[0009] Step 3, for different segmentation edges t, the edge and the cloud evaluate the layer delay and energy consumption of the neural network according to their own hardware environments and data sets. According to the division of the corresponding neural network model (i.e., the segmentation in Step 2), record the delay and data volume of the edge and the cloud for processing tasks, and record the delay required for transmitting data between the edge and the cloud according to the network bandwidth and the size of the transmitted data.
[0010] Step 4: The edge device and the cloud end, based on the current sampling rate Q, bandwidth B, and the deep learning network model, use the model partitioning algorithm to give an optimal periodic loop partitioning strategy. That is, according to the periodic loop partitioning strategy, when each frame of video stream arrives, the edge device and the cloud end switch to the corresponding segmented edge in sequence. In the case of a high sampling rate, the system reaches the maximum throughput of the neural network model.
[0011] Step 5: The edge device receives the first video frame and, according to the first partitioning strategy in the periodic loop partitioning strategy, infers the tasks before the segmented edge for the input first video frame, obtains the first inference result, and sends the first inference result to the cloud end. After completion, the edge device switches to the second partitioning strategy and infers the second video frame input, and so on.
[0012] Step 6: The cloud end takes the first inference result sent by the edge device as input, completes the subsequent inference process according to the first partitioning strategy, and returns the inference result of the first frame. After completion, the cloud end switches to the second partitioning strategy and waits for the second inference result from the edge device, and so on.
[0013] Step 7: When the input of k frames of video stream ends, the system outputs all inference results and ends the inference.
[0014] Further, the segmented edge t in Step 2 traverses all the edges between the model layers. There are m layers in the deep learning model layer, and there are m - 1 segmented edges t, t ∈ {t1, …, t m-1}, t ≥ 1. Under each segmented edge t, the number of model layers that the edge device needs to infer is n, and the size of the data output by the edge device is s. Under a certain network bandwidth B, it is transmitted to the cloud end as transmission data, and the transmission time T t = s / B. The size of the input data received by the cloud end is s, and the number of model layers that the cloud end needs to infer is m - n.
[0015] Further, in Step 3, the processing times T e and T c of the deep learning model layer at the edge device and the cloud end are actually obtained through actual measurement or simulation model prediction; the transmission time T t of the data between the model layers of the deep learning at the edge device and the cloud end is obtained through actual measurement, or by detecting the network bandwidth between the edge device and the cloud end and calculating the ratio of the size of the data output by the edge device to the bandwidth.
[0016] Further, the total time of the subtasks inferred by the edge device is T e , The total time of the subtasks inferred by the cloud end is where T iis the inference time of the i-th layer of the deep neural network, and the time for data transmission between the edge and the cloud is T t = s / B.
[0017] Further, when the sampling rate satisfies 1 / Q > minmax{T e , T t , T c}, the partitioning strategy is determined according to the shortest total time T of a single frame, that is, according to min(T e + T t + T c ) to determine the partitioning strategy to maximize the system throughput; when in the high-load mode, that is, 1 / Q < minmax{T e , T t , T c}, the optimal periodic cycle partitioning strategy is adopted to partition the deep learning model at this time.
[0018] Further, in step four, the input of the model partitioning algorithm is a (m - 1) x 3-dimensional vector, the sampling rate Q, and the bandwidth B, and the output is the partitioning strategy of the neural network model. In the input (m - 1) x 3-dimensional vector, each row vector represents {T e , T t , T c} under a splitting edge t. For a deep learning network model with m layers, there are m - 1 splitting edges. Input = {T e1 , T r1 , T c1 ; T e2 , T t2 , T c2 ;...; T e(m-1) , T t(m-1) , T c(m-1)}. The sampling rate is Q, and Q can represent the maximum value of the theoretical throughput, that is, the number of valid pictures transmitted in 1 second. For the case of an n-frame video stream, in the parallel mode, when the system is in the low-load mode, 1 / Q > minmax{T e , T t , T c}, where minmax{T e , T t , T c} represents the time with the largest proportion among the three times T e , T t , T c under each partitioning strategy, making it the smallest. In this case, there will be no waiting time when running in parallel in the system. At this time, according to min(T e + T t + T c)Determine the partitioning strategy σ. That is, when the load is low, determine the partitioning strategy σ according to the shortest total time T of a single frame. T = T e +T t +T c . At this time, the maximum limit throughput that the system can achieve is greater than the maximum value of the theoretical throughput of the sampling rate Q. At this time, adopting the above partitioning strategy σ can satisfy maximizing the system throughput. When the system is in the high-load mode, 1 / Q < minmax{T e ,T t ,T c}. At this time, the maximum limit throughput of the entire system adopting the above single-frame partitioning strategy σ is approximately equal to n / (n*minmax{T e ,T t ,T c}+T1+T2), where T1 and T2 are the other two items in max{T e ,T t ,T c}. For example, if max{T e ,T t ,T c} is T e , then T1 and T2 are T t ,T c . When T varies greatly or n is large, the above formula is approximately equal to 1 / (min max{T e ,T t ,T c}). It can be seen that the maximum limit throughput according to the single-frame partitioning strategy is less than the theoretical maximum throughput achieved by the sampling rate Q. No matter what, the single-frame partitioning strategy will not exceed the theoretical maximum throughput of the sampling rate Q.
[0019] Furthermore, with the partitioning strategy of the shortest total time T of the above single frame, in the case of multiple frames, it is not necessarily the optimal partitioning strategy. Existing technologies all operate at the same partitioning point. Even if the best pipelining execution partitioning node is found, there will still be waiting time, so there is still room for optimization. The model partitioning algorithm of the present invention formulates a periodic cyclic model partitioning strategy, and adopts different model partitioning strategies σ i for each frame, which is obtained by the periodic cyclic partitioning strategy algorithm of the present invention.
[0020] Furthermore, the optimal periodic cyclic partitioning strategy σ = {σ1, σ2, …, σ q}, where σ1, σ2, …, σ qrespectively represent the first partitioning strategy, the first partitioning strategy, …, the q-th partitioning strategy. The edge side and the cloud side perform inference according to this cyclic partitioning strategy. For the input of a k-frame video stream, after the edge side finishes inferring the first frame of the video stream using the first partitioning strategy σ1, it transmits the obtained first inference result s1 to the cloud side and switches its own partitioning strategy to the second partitioning strategy σ2; the cloud side uses the first partitioning strategy σ1 with s1 as the input, after finishing inferring the remaining tasks, returns the inference result of the first frame of the video stream, and switches its own partitioning strategy to the second partitioning strategy σ2 and repeats the above process; after the edge side finishes inferring the q-th frame of the video stream using the q-th partitioning strategy σ q After finishing inferring the q-th frame of the video stream, it transmits the obtained q-th inference result s q To the cloud side and switches its own partitioning strategy to the first partitioning strategy σ1; the cloud side uses the q-th partitioning strategy σ q With the q-th inference result s q As the input, after finishing inferring the remaining tasks, returns the inference result of the q-th frame of the video stream, and switches its own partitioning strategy to the first partitioning strategy σ1, achieving a cyclic process. Therefore, for the input of a k-frame video stream, the total number of times the edge side and the cloud side switch the partitioning strategy is N = k % q - 1, where % represents the modulo operation.
[0021] Furthermore, for DNNs, the number of some intermediate results (outputs of intermediate layers) is significantly smaller than the number of the original input data. For example, the input data size of Tiny YOLOv2 is 0.95 MB, while the output data size of the intermediate layer max5 is 0.08 MB, a reduction of 93%. This provides us with an opportunity to utilize the powerful computing power of cloud computing and the proximity advantage of edge computing. Specifically, a part of the DNN can be calculated on the edge side, a small amount of intermediate results can be transmitted to the cloud, and then the remaining part can be calculated on the cloud side. The partitioning of the DNN constitutes a trade-off between computing and transmission. Partitioning at different layers will result in different computing times and transmission times. Therefore, an optimal partitioning is desirable.
[0022] Based on the same inventive concept, the present invention also provides a neural network edge-cloud collaborative inference device for high-sampling-rate video stream analysis, which is an electronic device, including a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the method of the present invention described above.
[0023] Compared with the prior art, the positive effects of the present invention are:
[0024] (1) The present invention is a neural network edge-cloud collaborative inference method for high-sampling-rate video stream analysis. In the actual high-sampling-rate video stream scenario, through the modeling and analysis of deep neural networks, while dividing the inference task into different stages of the edge / cloud, a periodic cyclic division strategy for the deep neural network model is completed, breaking through the current bottleneck when facing high-sampling-rate video streams.
[0025] (2) Through the periodic cyclic division strategy, the edge / cloud reduces the inference waiting time in the system under the specified deep neural network model, reaches the limit throughput of the system, and fully utilizes the computing capabilities of the edge and cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic diagram of the total transmission time of different split edges of AlexNet;
[0027] Figure 2 is a schematic diagram of the high-sampling-rate situation;
[0028] Figure 3 is a schematic diagram of the cyclic division strategy in the embodiment;
[0029] Figure 4 is a schematic diagram of a single division strategy in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in more detail below through the accompanying drawings and corresponding embodiments.
[0031] As Figure 1 shown, taking the AlexNet network as an example, the split edge t traverses all the edges between the model layers. There are 24 layers in the deep learning model layer, and there are 23 split edges t in total, t ∈ {1, 2, 3,..., 23}, t ≥ 1. In each split case, the number of model layers that need to be inferred at the edge is n, and the size of the data output by the edge is s n , under a certain network bandwidth B, s n is transmitted as transmission data to the cloud, and the transmission time T t = s n / B. The size of the input data received by the cloud is s n , and the number of model layers that the cloud needs to infer is m - n. The processing times T e , T c of the deep learning model layer at the edge and in the cloud are actually obtained through actual measurement or simulation model prediction; the data transmission time T tIt is to calculate the ratio of the data size output by the edge side to the network bandwidth by actually measuring or detecting the network bandwidth between the edge side and the cloud side.
[0032] Through system actual measurement, the total time of the sub-tasks of edge-side inference is T e , The total time of the sub-tasks of cloud-side inference is T c , The total time for the edge side and the cloud side to transmit data is T t = s t / B. For the partitioning strategy with the shortest total time T of a single frame, the total time of deep neural network inference is T one partition = T e + T t + T c .
[0033] Experiments were conducted in the built system, where the edge server processor model is: Intel(R) Core(TM) i7-10510U CPU@1.80GHz 2.30GHz, the cloud server processor model is: CPU Intel(R) Xeon(R) Silver 4116 CPU@2.10GHz, and the operating system is Ubuntu 18.04. Calculate the total time of deep neural network inference for all split edges t in the case of a single frame, and represent it in the form of Figure 1 . The abscissa represents after a certain network layer where the split edge t is located, and the ordinate represents the total time calculated with T one partition = T e + T t + T c . At this time, m - 1 = 23 partitioning strategies are obtained, which are part of the input of the model partitioning algorithm.
[0034] First, the definition of high sampling rate is given. As Figure 2 shown, where frame A, frame B, and frame C represent three video frames in the video stream. When referring to the case of high sampling rate, what is faced must be a continuous video stream. Figure 2 The t0, t1, and t2 in it represent the arrival of video frames, that is, at the sampling rate Q, a frame of picture is captured at this moment and input into the deep neural network for inference. Figure 2 The representation of is representative (assuming that the edge-side inference time is greater than the transmission time and the cloud-side inference time). Assuming that the initialized partitioning strategy is the shortest total time of a single-frame picture, and the edge side and the cloud side always adopt this partitioning strategy. When the next frame arrives at the edge side, if the edge side has already completed the sub-tasks of the previous frame at this time, that is, it can immediately infer the new frame that arrives. The sampling rate at this time is called the low sampling rate by the present invention. AsFigure 2 as shown below. When the next frame arrives at the edge side, if the edge side has not finished reasoning the subtasks of the previous frame at this time and cannot immediately reason the new frame that arrives, there will inevitably be a waiting time for the new frame. After waiting for the edge side to complete the reasoning task of the previous frame, the reasoning of the new frame will start. The sampling rate at this time is called the high sampling rate by the present invention. For example Figure 2 as shown above. In the high sampling rate state, the total waiting time of the edge side in n frames of video stream is T wait = n * (T e - 1 / Q). When the video stream is longer, the waiting time is longer. During the waiting time, both the transmission process and the cloud reasoning process are in an idle state, wasting some resources.
[0035] The steps of the periodic cycle division algorithm for high sampling rate video streams proposed by the present invention are summarized as follows:
[0036] 1) Calculate all the division strategies σ = {σ1, σ2,..., σ m-1} of m - 1 dividing edges. Each division strategy σ corresponds to three times, namely the edge - side processing time T e (that is, the edge - side processing time in Figure 2 ), the transmission time T t (that is, the data transmission time in Figure 2 ), and the cloud - side processing time T c (that is, the cloud - side processing time in Figure 2 ). The total reasoning time is: T = T e + T t + T c . The optimal strategy for a single frame in the division strategy is denoted as which represents the division strategy σ1 when the total single - frame reasoning time T is the shortest. Set this division strategy as the first division strategy of the periodic cycle division strategy.
[0037] 2) Classify all the division strategies into three categories, that is, three sets: ① T e < T t and T e < T c , ② T t < T e and T t < T c , ③ T c < T e and T c < T tAccording to which set $\sigma_1$ belongs to, algorithm reasoning is carried out within that set. If the set has only one element $\sigma_1$, the result of the periodic cycle partitioning strategy is $\sigma' = \{\sigma_1\}$, which is constant. $\sigma'$ represents the finally obtained periodic cycle partitioning strategy. The problems of the three cases that $\sigma_1$ belongs to can be transformed into each other. The following step 3) takes case ③ as an example to illustrate the subsequent algorithm.
[0038] 3) For $T$ c The smallest set is ③, and the partitioning strategy is denoted as $\sigma$ ③ $= \{\sigma_1, \sigma_2, \ldots, \sigma$ r}\}, and at this time the sampling rate $1 / Q < \min\{\max\{T$ e , $T$ t , $T$ t}\}, $T$ e , $T$ t , $T$ c $\in \sigma$. For the convenience of the formula, $T$ e $> T$ t , that is, $1 / Q < T$ e . The next frame $k + 1$ arrives when the current frame $k$ has not been fully reasoned. $k$ starts from 1. The difference between $T$ e and $T$ t is the waiting time $T$ p , $T$ p $= |T$ e - $T$ t |. The waiting time for the first frame $k = 1$ is zero, that is, no waiting is required. When $k > 1$, due to pipeline parallel processing, when the edge device finishes processing the first stage of the current frame $k$, it can start processing the first stage of the next frame $k + 1$. After the current frame $k$ finishes processing the first stage, it enters the data transmission stage. After the next frame $k + 1$ finishes processing the first stage, it also enters the data transmission stage. However, when the next frame $k + 1$ enters the data transmission stage, it needs to satisfy that the data transmission stage of the current frame $k$ has been completed, that is, the network is in an available state. If the data of the current frame $k$ has not been fully transmitted, the data transmission of the next frame $k + 1$ needs to wait. The waiting time between the next frame and the current frame $T$ p $(k + 1) = |T$ e $(k + 1) - T$ t $(k)|$.
[0039] 4) To obtain the optimal periodic cycle partitioning strategy in the current situation, search according to the following steps:
[0040] i. The known set of partitioning strategies is: $\sigma$ ③ $= \{\sigma$ i $| i = 1...r\}$, and the set of optimal periodic cycle partitioning strategies is: $\{\sigma$ j $| i = 1...q\}$, where $q \leq r$. The partitioning strategy obtained according to the shortest total inference time per frame is $\sigma$one partition , this partitioning strategy is in the set of total partitioning strategies. Take the partitioning strategy of σ one partition as the initial partitioning strategy, denoted as σ tmp , and add it to the optimal cyclic partitioning strategy;
[0041] ii. Perform inference on the first arriving picture according to the partitioning strategy in step i. At this time, the edge inference time under this partitioning strategy can be determined: Data transmission time:
[0042] iii. When the next picture arrives, find the partitioning strategy in the partitioning strategy set σ ③ that is closest to as the partitioning strategy for the current frame. If exit the loop to obtain the final periodic cyclic partitioning strategy; otherwise, add this partitioning strategy to the optimal cyclic partitioning strategy and execute step iv;
[0043] iv. Update For the next picture, repeat step iii.
[0044] 5) The above loop result gives the optimal periodic cyclic partitioning strategy σ = {σ1, σ2, …, σ q}. At this time, the total time from the 1st to the kth frame of the video stream is:
[0045]
[0046] For the case of the shortest partitioning for a single frame, the total time for k-frame inference is:
[0047] T one partition = k * T e + T t + T c
[0048] The final partitioning strategy is:
[0049]
[0050] Now, a specific example is given to illustrate. The sampling interval in this example is: 1 / Q = 10 ms. The running time of the partitioning result for this neural network is shown in Table 1:
[0051] Table 1 shows the time for three stages under 4 partitioning strategies
[0052] Partitioning strategy <![CDATA[T e (ms)]]> <![CDATA[T t (ms)]]> <![CDATA[T c (ms)]]> <![CDATA[σ1]]> 30 20 10 <![CDATA[σ2]]> 20 30 15 <![CDATA[σ3]]> 40 25 8 <![CDATA[σ4]]> 10 15 40
[0053] The following gives a detailed description of the algorithm based on the content of Table 1:
[0054] 1) Calculate the edge - side processing time \(T\) of 4 partitioning strategies according to algorithm step 1 e , the transmission time \(T\) for data to be transmitted from the edge - side to the cloud t , and the processing time \(T\) of the cloud c ;
[0055] 2) The partitioning strategies \(\sigma_1\), \(\sigma_2\), \(\sigma_3\) satisfy \(T\) c \(<T\) e and \(T\) c \(<T\) t , and it is the set with the smallest \(T\). Search for the optimal periodic - loop partitioning strategy within this set through the algorithm; c
[0056] 3) The sampling interval satisfies \(1 / Q = 10\text{ms}<20\text{ms}<30\text{ms}\), which ensures that at any moment when the current frame is processed, the next frame has arrived;
[0057] 4) Traverse to find the minimum value of \(T\) At this time, the corresponding partitioning strategy for the first frame is \(\sigma_1\);
[0058] \(T\) p \((2)=\min T\) p \((j)=\min\{|T\) e \((j)-T\) t \((1)|\}=|T\) e \((j)-T\) t \((1)|\) j=2 \( = 0\text{ms}\), and the partitioning strategy for the second frame is \(\sigma_2\);
[0059] \(T\) p \((3)=\min T\) p \((j)=\min\{|T\) e \((j)-T\) t \((2)|\}=|T\) e \((j)-T\) t \((2)|\) j=1 \( = 0\text{ms}\), and the partitioning strategy for the third frame is \(\sigma_3\);
[0060] At this time, there exist \(a = 1\), \(b = 3\) such that \(T\) e \((a)==T\) e \((b)\), and jump out of the loop.
[0061] 5) \(\sigma=\{\sigma_1,\sigma_2\}\), the optimal periodic - loop partitioning strategy contains \(q\) partitioning methods, that is, \(q = 2\). The number of times the partitioning strategy is switched in the system during the final calculation is \(N=k\%q - 1\).
[0062] The total time of the cyclic partitioning strategy, the schematic diagram is Figure 3 , and the total time is as shown in the following expression:
[0063]
[0064] The total division time with the shortest single division point, and the schematic diagram is Figure 4 , and the total time is as shown in the following expression:
[0065] T one partition = k * T e + T t + T c = 30k + 30
[0066] When k is 99, the total time for inference according to the periodic cycle division strategy is 2510 ms, and the total time for inference according to the division strategy of the shortest single frame time is 3000 ms; when k is 100, the total time for inference according to the periodic cycle division strategy is 2545 ms, and the total time for inference according to the division strategy of the shortest single frame time is 3030 ms. The periodic cycle division algorithm of the present invention is superior to the division algorithm of the shortest single frame time.
[0067] Therefore, it can be obtained that when k is odd, to satisfy T ≤ T one partition , the value range of k obtained is k ≥ 1; when k is even, to satisfy T ≤ T one partition , the value range of k obtained is k ≥ 3.
[0068] In the scenario of high sampling rate video stream analysis, the periodic cycle division strategy neural network segment cloud collaborative inference method proposed by the present invention improves the inference speed.
[0069] Based on the same inventive concept, another embodiment of the present invention provides a neural network edge-cloud collaborative inference device for high sampling rate video stream analysis, which is an electronic device (such as a computer, a server, a smart phone, etc.), and includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing each step in the method of the present invention.
[0070] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc), and the computer-readable storage medium stores a computer program. When the computer program is executed by a computer, each step of the method of the present invention is implemented.
[0071] It should be noted that the above-described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
Claims
1. A neural network edge-cloud collaborative inference method for high-sampling-rate video stream analysis, characterized in that, It includes the following steps: In the scenario of a neural network architecture for high-sampling-rate video stream analysis, the deep learning model is modeled as a chain graph. The vertices of the chain graph represent the model layers of the neural network, and the arrows of the chain graph represent the data transmission between the model layers of the neural network; Perform segmentation between any two adjacent nodes in the chain graph. The segmented edge is denoted as t. The node at the left end of t indicates that the node performs inference at the edge side, and the node at the right end of t indicates that the node performs inference in the cloud; For different segmented edges t, record the latency and data volume of the processing tasks at the edge side and in the cloud, and record the latency required for data transmission between the edge side and the cloud according to the network bandwidth and the size of the transmitted data; The edge side and the cloud, based on the current sampling rate, bandwidth, and deep learning model, give the optimal periodic cycle partitioning strategy through the model partitioning algorithm; The edge side receives the first video frame, and according to the first partitioning strategy in the periodic cycle partitioning strategy, infers the tasks before the segmented edge for the input first video frame, obtains the first inference result, and sends the first inference result to the cloud. Then the edge side switches to the second partitioning strategy and infers the second video frame input, and so on; The cloud uses the first inference result sent by the edge side as input, completes the subsequent inference process according to the first partitioning strategy, and returns the inference result of the first frame; The cloud switches to the second partitioning strategy and waits for the second inference result from the edge side, and so on, until the input video stream ends, and outputs all inference results; Suppose there are m layers in the deep learning model layer and m - 1 splitting edges t. The input of the model partitioning algorithm is a (m - 1) x 3 - dimensional vector, the sampling rate Q, and the bandwidth B, and the output is the partitioning strategy of the neural network model; in the input (m - 1) x 3 - dimensional vector, each row vector represents {T e , T t , T c} under a splitting edge t, where T e is the processing time of the deep learning model layer at the edge, T c is the processing time of the deep learning model layer in the cloud, and T t is the transmission time of the data between the deep learning model layers at the edge and in the cloud; When the sampling rate satisfies 1 / Q > minmax{T e ,T t ,T c}, the partitioning strategy is determined according to the shortest total time T of a single frame, that is, according to min(T e +T t +T c ), to maximize the system throughput; when in the high-load mode, that is, 1 / Q < minmax{T e ,T t ,T c}, at this time, the optimal periodic loop partitioning strategy is adopted to partition the deep learning model; The model partitioning algorithm includes the following steps: Calculate all the partitioning strategies σ = {σ1, σ2, …, σ -1} of m - 1 splitting edges. Each partitioning strategy corresponds to three times, namely T e , T t , T c . The total inference time is: T = T e + T t + T c . The optimal strategy for a single frame among all the partitioning strategies is denoted as represents the partitioning strategy σ1 when the total inference time T of a single frame is the shortest. Set this partitioning strategy as the first partitioning strategy of the periodic cyclic partitioning strategy; Classify all partitioning strategies into three categories, namely three sets, ① T e <T t and T e <T c , ② T t <T e and T t <T c , ③ T c <T e and T c <T t ; According to which set σ1 belongs to, perform algorithmic reasoning within that set. If the set has only one element, σ1, then the result of the periodic cyclic partitioning strategy is σ′ = {σ1}, which is constant. σ′ represents the finally obtained periodic cyclic partitioning strategy; For set ③, the following steps are used to obtain the optimal periodic cycle partitioning strategy: The set of known partitioning strategies is: σ ③ = {σ i | i = 1...r}, and the set of optimal periodic loop partitioning strategies is: {σ j | i = 1...q}, where q ≤ r; the partitioning strategy obtained based on the shortest total inference time per frame is σ onepartition , and the partitioning strategy of σ onepartition is used as the initial partitioning strategy, denoted as σ tmp , and added to the optimal periodic loop partitioning strategies; Perform inference on the picture arriving in the first frame according to the step division strategy σ tmp to determine the edge inference time under this division strategy and the data transmission time When the next frame of image arrives, search for the set of partitioning strategies σ ③ in the closest partitioning strategy as the partitioning strategy for the current frame, If exit the loop to obtain the final periodic loop partitioning strategy; Otherwise, add this partitioning strategy to the optimal cycle partitioning strategy and execute step iv; Update For the next frame of the picture, repeat the above steps to finally obtain the optimal periodic cycle division strategy σ = {σ1, σ2, …, σ q}.
2. The method according to claim 1, characterized in that, For a high-sampling-rate video stream, the first arriving picture is inferred according to the first partitioning strategy of the optimal periodic cycle partitioning strategy, the second arriving picture is inferred according to the second partitioning strategy of the optimal periodic cycle partitioning strategy, and the frame stream is processed according to this rule. When traversing to the last partitioning strategy of the optimal periodic cycle partitioning strategy, return to the first partitioning strategy to start cycling, and so on.
3. A neural network edge-cloud collaborative inference device for high-sampling-rate video stream analysis, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method according to claim 1 or 2.
4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to claim 1 or 2 is implemented.