A Deep Learning Model Partitioning Method and Device for Pipeline Distributed End-Cloud Collaborative Inference
By adopting the deep learning model division method of pipeline distributed end cloud collaborative inference in video stream processing, the blocking problem in video stream processing at high sampling rate is solved, and higher system throughput and efficiency is achieved.
Patent Information
- Application Number
- CN202210368166.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-08
AI Technical Summary
The prior art has blocking problems in video stream processing at high sampling rates, resulting in a significant reduction in the throughput and efficiency of the inference system.
The deep learning model division method for pipeline distributed end cloud collaborative inference is adopted, and the deep neural network model is divided into three parts: edge device inference, intermediate result transmission, and cloud server inference, and it is processed as three parallel processes to achieve real-time inference.
Through this method, the blockage of the inference system can be avoided at high sampling rates, the system throughput and efficiency can be improved, and the real-time and stability of video stream processing can be ensured.
Smart Images

Figure CN115130649B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of edge computing and deep neural network model inference, and particularly relates to a pipeline-oriented end-cloud collaborative distributed inference acceleration method and device. Background Art
[0002] Video analysis is the core for implementing a series of applications ranging from surveillance and autonomous vehicles to personal digital assistants and autonomous drone control. The current state-of-the-art method is to use deep neural networks, where video frames are processed by a well-trained convolutional neural network or recurrent neural network. Video analysis uses a deep neural network to extract features from the input frames of a video and classify the objects in the frames into one of the predefined classes. With the continuous development of cloud computing and the continuous improvement of the performance of cloud servers, the current popular approach is to upload complex deep neural networks to cloud servers for inference calculation. For example, in the application of an intelligent camera, the camera continuously monitors the indoor scene and transmits it to the server. The neural network model preset in the server analyzes the video frames and feeds back abnormal situations to the user or an alarm device. This technology is widely used in scenarios such as homes, shopping malls, and factories. In the application of autonomous driving, multiple cameras installed at the front, side, and rear of the vehicle collect environmental data in real time and upload it to the server. The DNN detects physical objects or lane markings by analyzing the images and sends signals to devices such as the steering wheel or pedals. Due to the huge amount of data generated by the video stream, on the one hand, data explosion may occur in the cloud service center, and on the other hand, the change in bandwidth causes a high latency in the transmission of raw data.
[0003] To avoid the above impacts, edge computing emerged, that is, placing computing resources near the data source. However, edge computing itself is limited by computing power and storage capacity and cannot completely replace cloud computing. A better method is needed to combine the two. Currently, a neural network model splitting method for single-frame data has been proposed. The model is split into two parts. The model of the inference part on the edge side obtains an intermediate result and uploads the intermediate result to the cloud server. The cloud server infers the remaining part of the model to achieve collaborative inference between the edge and the cloud server.
[0004] In actual industrial applications, the application of video stream processing is more extensive, such as the above-mentioned autonomous driving and factory intelligent monitoring scenarios. The single-frame-based model partitioning method does not consider the blocking problem in the frame stream environment and still has certain room for optimization. Since data-intensive tasks require real-time continuous inference and feedback of results, when using the single-frame-based model partitioning method for inference, blocking will occur when the sampling rate is relatively high. At this time, the throughput and efficiency of the entire inference system will be greatly reduced. Summary of the Invention
[0005] The object of the present invention is to provide a deep learning model partitioning method for pipeline distributed edge-cloud collaborative inference aiming at the deficiencies of the existing technology, which can solve the blocking problem under high sampling rates to a certain extent, avoid the inference system from being blocked when the frame stream arrives, and well make up for the deficiencies of the existing technology. Edge-cloud collaborative inference can be divided into three parts: edge device inference, intermediate result transmission, and cloud server inference. The three parts can be used as three processes and processed in parallel during the inference process. That is, while the cloud is inferring the current video frame, the edge device can infer the first half of the next frame. Based on the above pipeline inference steps, the present invention obtains the optimal partitioning point of the model to achieve real-time inference of model edge-cloud collaboration.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A deep learning model partitioning method for pipeline distributed edge-cloud collaborative inference, comprising the following steps:
[0008] 1) The edge device obtains the hardware environments of itself and the cloud server, and evaluates the time delay of each layer in the deep neural network model;
[0009] 2) The edge device obtains the output data size of each layer in the deep neural network model, and calculates the transmission time of the output data of each layer according to the current network conditions;
[0010] 3) The edge device obtains the partitioning point of the current deep neural network model, and calculates the times of three stages under this partitioning point, namely: edge-side inference time, data transmission time, and cloud-side inference time;
[0011] 4) Traverse the above total time, and obtain the partitioning strategy by constraining the minimum total;
[0012] 5) Obtain the maximum ratio of the three stages under the current partitioning strategy, and compare it with the sampling interval. If the maximum value is less than or equal to the sampling interval, the current partitioning strategy is the optimal partitioning strategy. If this condition is not met, perform step 6);
[0013] 6) Traverse the maximum values of the three stages under all partitioning points, constrain the minimum of the maximum values to obtain the partitioning strategy, obtain the maximum ratio of the three stages under the current partitioning strategy, and compare it with the sampling interval. If the maximum value is less than or equal to the sampling interval, the current partitioning strategy is the optimal partitioning strategy. If not, adjust the sampling rate;
[0014] 7) The edge device calculates all the inference subtasks before the segmentation point of the deep neural network according to the optimal partitioning strategy and sends the intermediate results of the inference to the cloud; the cloud takes the intermediate results of the inference sent from the edge device as input and calculates the inference subtasks between the segmentation point and the last layer according to the neural network layering; while the video frame stream arrives, the inference tasks for each frame are repeated until the entire inference task is completed.
[0015] Further, the end-cloud collaborative inference process is divided into three parts according to the model partitioning point: edge-side inference, intermediate data transmission, and cloud inference. These three parts are carried out in parallel as three processes. That is, when the result of the first part of the inference of the neural network of the video frame is transmitted to the cloud, the second part of the neural network of the previous frame is processed on the edge device at the same time.
[0016] Further, the implementation of the segmentation scheme of the deep neural network model is as follows:
[0017] Let the number of layers of a deep neural network be n, the sampling rate be Q, then the sampling interval is 1 / Q, and the total delay of the entire inference is denoted as T. Assume that the segmentation starts from the j-th layer of the deep neural network, that is, the first j layers perform inference on the edge side, and the (j + 1)-th layer and subsequent layers perform inference on the cloud. Then the inference time on the edge side is expressed as T e =[t e1 ,t e2 ,…t ej , the total transmission time is T t =D j / B, the total time for cloud processing is T c =t c(j+1) +t c(j+2) +…+t cn , the total time for the neural network to execute a single frame is T = T e +T c +T t , and the total time for n frames to execute is expressed as
[0018] Then obtain the optimal partitioning point according to the following steps:
[0019] First, minimize the total delay for processing each frame, that is, minimize T = T e +T t +T c , to obtain the partitioning point of the entire network model; if each stage can be completed before the next frame arrives, that is Then obtain the shortest time delay according to the above partitioning point; in summary, when Minimize T = T e +T t +T c To obtain the optimal partitioning point;
[0020] When happens, minimize max{T e , T t , T c}, that is, min max{T e , T t , T c}. The maximum sampling rate is
[0021] If the sampling rate is greater than then force the sender or user to reduce the sampling rate.
[0022] Furthermore, the inference process of video frames is carried out in a pipeline form. Use frame(i) to represent the i-th video frame arriving at the edge side. The inference steps include:
[0023] 1) When frame(i + 2) arrives, the edge side calculates the inference task of the first half of the network of frame(i + 2). At the same time, frame(i + 1) completes the edge-side calculation task, and the intermediate result is transmitted from the edge side to the cloud;
[0024] 2) frame(i + 1) arrives at the cloud, frame(i) completes the inference task at the cloud, and at the same time the edge device sends the intermediate result of frame(i + 2) to the cloud;
[0025] 3) The cloud imports the intermediate result of frame(i + 1) from the edge side into the second half of the neural network layer, calculates the calculation task after the segmentation point, and completes the calculation task of the second half of the network of the current frame;
[0026] 4) Repeat the whole process while the video frame arrives until there are no incoming video frames at the edge side.
[0027] Furthermore, the used deep neural network model is pre-saved on the edge device and the cloud server.
[0028] Furthermore, the edge device and the cloud respectively load and infer the corresponding neural network models. The inference work is based on the dataset and uses the same sample data. According to the size of the output feature map of each layer under the current network condition, obtain the inference latency of each layer of the model at the edge side and the cloud, and transmit it through the TCP protocol to save the obtained data information to the edge device and the cloud server.
[0029] Based on the same inventive concept, the present invention also provides a deep learning model partitioning device for pipeline distributed end-cloud collaborative inference, which is an electronic device including a memory and a processor. The memory stores a computer program configured to be executed by the processor, and the computer program includes instructions for executing the method of the present invention described above.
[0030] Compared with the prior art, the positive effects of the present invention are as follows:
[0031] (1) Traditional model partitioning methods are based on single-frame pictures, while the present invention is based on video stream analysis, which is more in line with the actual situation and is flexible and efficient.
[0032] (2) Traditional model partitioning methods do not consider the pipeline parallel inference mode. The present invention regards the three stages of edge-side inference, data transmission, and cloud inference as three independent processes, and on this basis, considers the selection of the optimal partitioning point.
[0033] (3) Traditional model partitioning methods do not consider the system blocking problem under high sampling rates. The present invention adjusts the partitioning strategy according to the system sampling rate, and avoids system congestion on the premise of ensuring inference latency and system throughput. Description of the Drawings
[0034] Figure 1 is the schematic diagram of pipeline parallel inference of the present invention;
[0035] Figure 2 is the schematic diagram of the total time of pipeline parallel inference;
[0036] Figure 3 is the pipeline comparison diagram under different sampling rates, where (a) is the case, and (b) is the case;
[0037] Figure 4 is the schematic diagram of the deep neural network model in the embodiment. Detailed Embodiments
[0038] The technical method of the present invention will be further described below with reference to the drawings and examples, but the scope of the present invention is not limited in any way.
[0039] The present invention relates to a method for partitioning a deep learning model for pipeline distributed edge-cloud collaborative inference. In the segmentation point evaluation stage of the model, the resource conditions and network status under the edge and cloud systems are regularly evaluated and analyzed. By evaluating the model and the segmentation algorithm, the partitioning method of the current deep neural network model is determined. According to the segmentation point (or partitioning point), the neural network is divided into two parts. The model calculation tasks of the first part are executed on the edge device, and the model calculation tasks of the second part are executed in the cloud. The whole process is in the form of a pipeline, divided into three stages: edge-side execution, data transmission, and cloud execution. The execution stage of the latter frame on the edge side can be carried out simultaneously with the data transmission stage of the previous frame, so as to improve the throughput rate of the whole system.
[0040] As Figure 1 shown, a key idea of this patent is to allow the video frame stream to be processed in a pipeline manner. For example, when the inference result of the first part of the neural network of a video frame is transmitted to the cloud, the second part of the neural network of the previous frame can be processed on the edge device at the same time.
[0041] 1) Let the number of layers of a deep neural network be n, and denote the computational hierarchical delay on the edge side as T e =[t e1 ,t e2 ,…t en , and the data volume of each layer is D = [D1, D2, …, D n ; denote the computational hierarchical delay on the cloud side as T c =[t c1 ,t c2 ,…t cn ; denote the transmission bandwidth from the edge side to the cloud side as B; the sampling rate as Q; and denote the inference delay of the whole system as T.
[0042] Assume that the segmentation starts from the j-th layer of the deep neural network, that is, the first j layers are inferred on the edge side, and the (j + 1)-th layer and subsequent layers are inferred in the cloud. Then the total inference time on the edge side can be expressed as T e =[t e1 ,t e2 ,…t ej , the total transmission time is T t =D t / B, and the total time for the cloud server to process the second part of the model is T c =t c(j+1) +t c(j+2) +…+t cn .
[0043] The total time for the neural network to execute a single frame is:
[0044] T=T e +T c +Tt (1)
[0045] As Figure 2 shown, the completion time for executing n-frame blocks on the neural network is:
[0046] T n = n * T e + T t + T c (2)
[0047] For the case where the number of frames n is very large, the overall total processing time can be expressed as:
[0048]
[0049] In this case, the stage with the longest inference time has the greatest impact on the total delay of the system.
[0050] 2) To achieve real-time inference for end-cloud collaborative inference and obtain the minimum inference delay, the present invention first minimizes the total delay for processing each frame, that is, minimizes T = T e + T t + T c , to obtain the partitioning point of the entire network model. If each stage can be completed before the arrival of the next frame, as Figure 3 shown in (a) below, that is the system obtains the shortest delay according to the above partitioning point. In summary, when To obtain the optimal partitioning point, this algorithm minimizes T = T e + T t + Tx.
[0051] When As Figure 3 shown in (b) below, the next frame arrives before the current frame is processed on the edge device. At this time, the system needs to wait for the current video frame inference to complete and then continue the inference. If the partitioning strategy is not adjusted, system blocking will occur. Therefore, in this case, it is necessary to maximize the throughput of the system, that is, to make the system process as many frames as possible per unit time to avoid blocking. At this time, the goal is to minimize max{T e , T t , T c}, that is, min max{T e , T t , T c}, and the maximum sampling rate of the system is
[0052] In addition, it is also necessary to consider that if the sampling rate is greater than the system will eventually become congested. The system must force the sender or user to reduce the sampling rate.
[0053] The specific steps are as follows:
[0054] 1) The edge device evaluates the latency of each layer of the deep neural network model based on its own and the cloud server's hardware environment, obtaining the latency and output size of each layer of the model on different servers.
[0055] 2) Obtain the bandwidth of the current network environment, and based on the current bandwidth, obtain the transmission time of the output data of each layer of the model.
[0056] 3) Traverse the sum of the times of the three stages of the edge-side inference time, intermediate data transmission time, and cloud-side inference time under all division points of the current model, and obtain the division point with the minimum sum of times as the division strategy.
[0057] 4) Obtain the maximum value of the ratios of the three stages under the current division strategy, and compare it with the system sampling interval. If the maximum value is less than or equal to the sampling interval, the current division strategy is the optimal division. If this condition is not met, proceed to step 6).
[0058] 5) Traverse the maximum values of the three stages under all division points, constrain its maximum value to be the minimum, obtain the division strategy, obtain the maximum value of the ratios of the three stages under the current division strategy, and compare it with the system sampling interval. If the maximum value is less than or equal to the sampling interval, the current division strategy is the optimal division. If not, adjust the sampling rate.
[0059] 6) Synchronize the deep neural network model for inference to the edge device and the cloud through a download or distribution mechanism.
[0060] 7) The edge device calculates all inference subtasks before the division point of the deep neural network according to the division strategy, and sends the intermediate results of the inference to the cloud.
[0061] 8) The cloud uses the intermediate results of the inference sent from the edge device as input to infer the subtasks of the second half (after the division point).
[0062] 9) When the video frame stream arrives, repeat the inference tasks for each frame (repeat steps 7)-8)) until the entire inference task is completed.
[0063] In an embodiment of the present invention, the following test scenario is set: one edge server, an edge server, and one cloud server, a cloud server; the inference model is AlexNet, and the inference framework is Caffe.
[0064] The implementation steps are as follows:
[0065] 1) Model synchronization. Train a deep neural network model and synchronize the deep neural network model to be used to the terminal side, edge side, and cloud side through data synchronization technology. The neural network model can also be a pre-trained model stored in the cloud model repository, and the edge device and the corresponding cloud can download the same version of the model from the same location. Figure 4 It is a schematic diagram of the deep neural network model in this embodiment.
[0066] 2) Performance prediction. The edge device and the cloud first load and infer the corresponding neural network models respectively. The inference work is based on the dataset, and the same sample data is used to predict the inference latency of each layer of the model on the edge side and the cloud side. In practical applications, Caffe time can be used to obtain the inference latency data of each layer. Assume that the neural network model to be inferred has n layers, and the latency of each neural network layer is calculated through multiple inferences. The latency of each layer on the edge side determined by the method of taking the average value through multiple inferences is [t e1 , t e2 , …, t en , the inference latency on the cloud side is [t c1 , t c2 , … t cn , and the amount of intermediate data that needs to be transmitted over the network is [D1, D2, …, D n .
[0067] 3) Obtain the current bandwidth and calculate the output data transmission time T of each layer according to the bandwidth t = D / B.
[0068] 4) Combining the current example, the inference time at the lower end side, data transmission time, and cloud side inference time of the AlexNet model at different partitioning points can be obtained, as shown in Table 1.
[0069] 5) If inference is performed in the form of a single frame, selecting the partitioning point at Pool1 can achieve the optimal efficiency in this case. The total time for inferring 1 frame is 291.5 ms. Among them, the inference time of the edge device is 54 ms, the data transmission time is 35 ms, and the processing time of the cloud side is 201 ms. The largest among them is the cloud side processing time of 201 ms. The maximum theoretical throughput (sampling rate) is 5 FPS.
[0070] 6) In the case of a video stream, using the method of this patent, the optimal partitioning point can be obtained as Pool2. The processing time of one frame on the edge device is 162 ms, the transmission time is 22 ms, and the processing time on the cloud side is 167 ms. The largest among them is the cloud side processing time of 167 ms. At this time, the theoretical throughput can reach a maximum of 6 FPS. It can be seen that it is superior to the partitioning method in the single frame case.
[0071] 7) Real-time inference. After obtaining the division point, the edge device infers up to the division point, and transmits the intermediate result and the division information via TCP. After the cloud obtains this information, the inference model processes the part after the division point. Use frame(i) to represent the i-th video frame arriving at the edge side. The steps are as follows:
[0072] i. When frame(i+2) arrives, the edge side calculates the inference task of the first half of the network for frame(i+2). At the same time, frame(i+1) has completed the edge-side calculation task, and the intermediate result is transmitted from the edge side to the cloud;
[0073] ii. frame(i+1) arrives at the cloud, frame(i) completes the inference task at the cloud, and at the same time the edge device sends the intermediate result of frame(i+2) to the cloud;
[0074] iii. The cloud imports the intermediate result of frame(i+1) from the edge side into the second half of the neural network layer, calculates the calculation task after the division point, and completes the calculation task of the second half of the network;
[0075] iv. Repeat the whole process while the video frames are arriving until there are no more incoming video frames at the edge side.
[0076] Table 1 shows the time descriptions under different division points
[0077]
[0078] Based on the same inventive concept, another embodiment of the present invention provides a deep learning model partitioning device for pipelined distributed end-cloud collaborative inference, which is an electronic device (computer, server, smart phone, etc.), and includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the method of the present invention.
[0079] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disc), and the computer-readable storage medium stores a computer program. When the computer program is executed by a computer, each step of the method of the present invention is implemented.
[0080] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is defined by the scope of the claims.
Claims
1. A deep learning model partitioning method for pipeline distributed edge-cloud collaborative inference, characterized in that It includes the following steps: 1) The edge device obtains the hardware environments of itself and the cloud server, and evaluates the latency of each layer in the deep neural network model; 2) The edge device obtains the output data size of each layer in the deep neural network model, and calculates the transmission time of the output data of each layer according to the current network conditions; 3) The edge device obtains the partitioning point of the current deep neural network model, and calculates the time of three stages under this partitioning point, namely: the end-side inference time, the data transmission time, and the cloud-side inference time; 4) Traverse the total time of the end-side inference time, the data transmission time, and the cloud-side inference time, and obtain the partitioning strategy by constraining the minimum total; 5) Obtain the maximum ratio of the three stages under the current partitioning strategy, and compare it with the sampling interval. If the maximum value is less than or equal to the sampling interval, the current partitioning strategy is the optimal partitioning strategy. If this condition is not met, go to step 6); 6) Traverse the maximum values of the three stages under all partitioning points, constrain the minimum of the maximum values to obtain the partitioning strategy, obtain the maximum ratio of the three stages under the current partitioning strategy, and compare it with the sampling interval. If the maximum value is less than or equal to the sampling interval, the current partitioning strategy is the optimal partitioning strategy. If not, adjust the sampling rate; 7) The edge device calculates all inference subtasks before the deep neural network segmentation point according to the optimal partitioning strategy, and sends the intermediate results of the inference to the cloud; the cloud uses the intermediate results of the inference sent from the edge device as input, and calculates the inference subtasks from the segmentation point to the last layer according to the neural network layering; While the video frame stream arrives, repeat the inference task for each frame until the entire inference task is completed; The end-cloud collaborative inference process is divided into three parts according to the model partitioning point: end-side inference, intermediate data transmission, and cloud-side inference. These three parts are carried out in parallel as three processes. That is, when the result of the first part of the neural network inference of the video frame is transmitted to the cloud, the second part of the neural network of the previous frame is processed on the edge device at the same time.
2. The deep learning model partitioning method for pipeline distributed end-cloud collaborative inference according to claim 1, wherein The implementation of the segmentation scheme of the deep neural network model is as follows: Let the number of layers of a deep neural network be \(n\), the sampling rate be \(Q\), then the sampling interval is \(1 / Q\), and the total inference latency is denoted as \(T\). Assume that the splitting starts from the \(j\)-th layer of the deep neural network, that is, the first \(j\) layers perform inference on the edge side, and the \((j + 1)\)-th layer and subsequent layers perform inference in the cloud. Then the inference time on the edge side is expressed as \(T\) e =\([t e1 ,t e2 ,…t ej \), the total transmission time is \(T\) t =D j / B, the total time for cloud processing is \(T\) c =t c(j+1) +t c(j+2) +…+t cn , the total time for the neural network to execute a single frame is \(T = T\) e +T c +T t , and the total time for executing \(n\) frames is expressed as Then obtain the best partitioning point according to the following steps: First, minimize the total delay for each frame, that is, minimize T = T e + T t + T c , to obtain the partitioning point of the entire network model; if each stage can be completed before the arrival of the next frame, that is then obtain the shortest delay according to the above partitioning point; in summary, when minimize T = T e + T t + T c to obtain the optimal partitioning point; When minimize max{T e , T t , T c}, that is, min max{T e , T t , T c}, and the maximum sampling rate is If the sampling rate is greater than then force the sender or user to reduce the sampling rate.
3. The method for dividing a deep learning model for pipeline distributed end-cloud collaborative inference according to claim 1, wherein The inference process of the video frame is carried out in a pipeline form. Use frame(i) to represent the i-th video frame arriving at the edge side. The inference steps include: 1) When frame(i + 2) arrives, the edge side calculates the inference task of the first half of the network of frame(i + 2). At the same time, frame(i + 1) completes the edge-side calculation task, and the intermediate result is transmitted from the edge side to the cloud; 2) frame(i + 1) arrives at the cloud, frame(i) completes the inference task in the cloud, and at the same time the edge device sends the intermediate result of frame(i + 2) to the cloud; 3) The cloud imports the intermediate result of frame(i + 1) from the edge side into the neural network layer of the second half, calculates the calculation task after the segmentation point, and completes the calculation task of the second half of the network of the current frame; 4) Repeat the whole process while the video frame arrives until there are no incoming video frames on the edge side.
4. The method for partitioning a deep learning model for pipeline distributed edge-cloud collaborative inference according to claim 1, wherein The pre - used deep neural network model is pre - saved on the edge device and the cloud server.
5. The method for dividing a deep learning model for pipeline distributed edge-cloud collaborative inference according to claim 1, wherein The edge device and the cloud respectively load and infer the corresponding neural network models. The inference work is based on the dataset and uses the same sample data. According to the size of the output feature map of each layer under the current network condition, the inference latency of each layer of the model on the edge side and the cloud is obtained and transmitted through the TCP protocol, so that the obtained data information is saved on the edge device and the cloud server.
6. A deep learning model partitioning device for pipeline distributed edge-cloud collaborative inference, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The computer - readable storage medium stores a computer program, and when the computer program is executed by the computer, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Deep learning model training acceleration method based on end-to-edge cloud cooperation
CN111242282A
Robot task division-oriented end, side and cloud collaborative computing device
CN112287609A