A method and system for optimizing the allocation of terminal edge joint resources
Through the method of optimized configuration of terminal edge joint resource, the problems of high computational complexity, delay and energy consumption of deep neural network inference tasks on mobile devices are solved, and the effect of improving model inference accuracy while reducing delay and energy consumption is achieved.
Patent Information
- Application Number
- CN202210455191.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-27
AI Technical Summary
When performing inference tasks for deep neural networks (DNNs) on mobile devices, there are problems such as high computational complexity, long inference delay and high energy consumption, resulting in less use of deep learning in existing AR applications.
Through the terminal edge joint resource optimization configuration method, combined with edge controller, edge DNN inference module, video sampling management module and local controller, dynamically adjust the number of video frames, decide whether the inference task is completed locally or at the edge, and allocate communication and computing resources to optimize inference delay, energy consumption and accuracy.
It effectively reduces the average inference delay of the system and the average energy consumption of mobile devices, while improving the accuracy of model inference, realizing multi-dimensional performance optimization between inference delay, energy consumption and accuracy.
Smart Images

Figure CN114756371B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of computer technology, and in particular to a method and system for optimizing the configuration of terminal edge joint resources. Background Art
[0002] The development of network, cloud computing, edge computing, artificial intelligence and other technologies has triggered people's infinite imagination of the metaverse. In order to enable users to interact between the real world and the virtual world, augmented reality (AR) technology plays a vital role. At the same time, artificial intelligence plays an important role in automatic speech recognition, natural language processing, computer vision and other fields due to its learning and reasoning capabilities. With the assistance of AI technology, AR can achieve deeper scene understanding and more immersive interaction.
[0003] However, the computational complexity of artificial intelligence algorithms, especially deep neural networks (DNNs), is usually very high. On mobile devices with limited computing and energy capacity, it is difficult to complete the inference of neural networks in a timely and reliable manner. Experiments show that even with the acceleration of mobile GPUs, a typical single-frame image processing AI inference task takes about 600 milliseconds. In addition, continuous execution of the above inference tasks can only last up to 2.5 hours on commodity devices. The above problems have led to only a few AR applications using deep learning at present. In order to reduce the inference time of DNNs, one way is to prune the neural network. However, if too many channels are pruned, the model may be damaged, and it may not be possible to restore satisfactory accuracy through fine-tuning.
[0004] Mobile edge computing-assisted AI is another approach to address these issues. The integration of mobile edge computing and AI technologies has recently become a promising paradigm for supporting computationally intensive tasks. Edge computing transfers the inference and training process of AI models to the edge of the network close to the data source. Therefore, it will alleviate network traffic load, latency, and privacy issues. However, there are still a lot of challenges for edge computing-assisted AI applications, which are as follows:
[0005] (1) Although edge computing resources are much stronger than those of end users, they are also limited. Simply relying on the computing power of edge devices cannot effectively solve the problem of insufficient terminal computing power;
[0006] (2) Users offload AI reasoning tasks to edge computing, which can reduce the impact of insufficient computing power to a certain extent, but will also introduce communication delays;
[0007] (3) Latency, energy consumption, and accuracy constrain each other. Improving the performance of one of them will inevitably lead to a decline in the performance of the other two. Summary of the invention
[0008] The purpose of the present invention is to overcome the shortcomings of the prior art. The present invention provides a method and system for optimizing the configuration of terminal edge joint resources. Based on the optimization method, the average inference delay of the system and the average energy consumption of mobile devices are reduced at the same time, and the accuracy of model inference is improved.
[0009] The present invention provides a method for optimizing the configuration of terminal edge joint resources, the method comprising:
[0010] When the edge controller records different frames of video and inputs them into the neural network, the number of multiplications and additions required by the neural network and the accuracy of neural network recognition are calculated; when receiving video recognition requests from various mobile devices, the control task is completed according to the video recognition requests of the mobile devices, and the control task includes:
[0011] Determine the frame number of video sampling of each mobile device, and send sampling frame number control information to the video sampling management module corresponding to the mobile device based on the frame number of video sampling of each mobile device;
[0012] Determine the user's uninstallation decision, and determine whether the user's reasoning task is completed in the local DNN reasoning module or in the edge DNN reasoning module based on the user's uninstallation decision;
[0013] Based on the allocation strategy of communication resources of each mobile video device, the time proportion of uplink video transmission of each mobile device is determined;
[0014] Determine the resource allocation strategy of the edge DNN inference module to determine the CPU computing frequency of each offload user.
[0015] The method further comprises:
[0016] The edge DNN inference module obtains the uploaded video from the edge device that needs to be offloaded, then completes the inference according to the computing resources allocated by the edge controller, and sends the inference results to each mobile device.
[0017] The method further comprises:
[0018] The video sampling management module obtains sampling frame number control information of edge computing, controls the sampling frame number of the mobile device based on the sampling frame control information, and determines the frame number of the input video used for neural network reasoning.
[0019] The method further comprises:
[0020] The local controller obtains the video from the video sampling management module, and decides whether to transmit the video to the edge server according to the user unloading decision obtained from the edge controller.
[0021] The local controller obtains the video from the video sampling management module, and determines whether to transmit the video to the edge server according to the user unloading decision obtained from the edge controller, including:
[0022] If the user's offloading decision is 1, the video is transmitted to the edge server for inference, and the communication resources for transmission are configured by the base station;
[0023] If the user's uninstall decision is 0, the video is allowed to complete inference locally, and local CPU computing resources are allocated according to local device information.
[0024] When the mobile device needs to complete DNN reasoning locally, the local DNN reasoning module obtains the video and uses the allocated local computing resources to complete the DNN reasoning.
[0025] Accordingly, the present invention also provides an edge-end collaborative video AI reasoning system, the system comprising:
[0026] The edge controller is used to record when different frames of video are input into the neural network, calculate the number of multiplications and additions required by the neural network and the accuracy of neural network recognition; when receiving video recognition requests from various mobile devices, it completes the control task according to the video recognition requests of the mobile devices;
[0027] The edge DNN inference module is used to obtain the uploaded video from the edge device that needs to be offloaded, then complete the inference based on the computing resources allocated by the edge controller, and send the inference results to each mobile device;
[0028] A video sampling management module is used to obtain sampling frame control information of edge computing, control the sampling frame number of the mobile device based on the sampling frame control information, and determine the frame number of the input video for neural network reasoning;
[0029] A local controller, used for obtaining videos from the video sampling management module and determining whether to transmit the videos to the edge server according to the user offloading decision obtained from the edge controller;
[0030] The local DNN inference module is used to obtain video and use the allocated local computing resources to complete DNN inference when the mobile device needs to complete DNN inference locally.
[0031] The completing the control task according to the video recognition request of the mobile device includes:
[0032] Determine the frame number of video sampling of each mobile device, and send sampling frame number control information to the video sampling management module corresponding to the mobile device based on the frame number of video sampling of each mobile device;
[0033] Determine the user's uninstallation decision, and determine whether the user's reasoning task is completed in the local DNN reasoning module or in the edge DNN reasoning module based on the user's uninstallation decision;
[0034] Based on the allocation strategy of communication resources of each mobile video device, the time proportion of uplink video transmission of each mobile device is determined;
[0035] Determine the resource allocation strategy of the edge DNN inference module to determine the CPU computing frequency of each offload user.
[0036] If the user uninstall decision is 1, the local controller transmits the video to the edge server for reasoning, and the communication resources for transmission are configured by the base station.
[0037] If the user uninstall decision is 0, the local controller allows the video to complete reasoning locally and allocates local CPU computing resources according to local device information.
[0038] The embodiments of the present invention have the following beneficial effects:
[0039] (1) The present invention provides an architecture of an edge-cooperative video AI inference system based on an edge-cooperative AI inference algorithm offloading architecture. The edge server can determine the number of video frames used for detection by the user according to the number of requesting users, and provide the user's offloading strategy and the user's communication computing resource allocation plan.
[0040] (2) Multi-dimensional performance optimization. Under the given system architecture, an effective algorithm is proposed by jointly considering inference delay, terminal energy consumption and recognition accuracy. It can improve the recognition accuracy of the neural network while reducing delay and energy consumption.
[0041] (3) Performance trade-off analysis. The invention provides the trade-off relationship between inference delay, terminal energy consumption and recognition accuracy. By using this relationship, targeted system optimization can be performed to determine the system performance of AI inference under different applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 It is a schematic diagram of the structure of the edge-end collaborative video AI reasoning system in an embodiment of the present invention;
[0044] Figure 2is a schematic diagram of performance comparison of different unloading strategies in an embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of the trade-off relationship between delay, energy consumption, and recognition accuracy in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] The main technical problem solved by the present invention is to propose an edge-to-edge collaborative video AI reasoning system architecture for deep learning-based video recognition tasks, and based on the architecture, a method for simultaneously optimizing the latency, energy consumption, and accuracy of deep learning reasoning tasks is proposed. At the same time, the present invention provides a trade-off relationship between latency, energy consumption, and accuracy, which can be used as a reference for system design.
[0048] The present invention proposes an architecture of an edge-end collaborative video AI inference system, whose function is to reduce the inference delay and energy consumption of the neural network while improving the inference accuracy of the neural network by reasonably adjusting the number of video frames to be recognized and jointly configuring wireless and computing resources in a scenario where an edge server serves multiple mobile devices through wireless access.
[0049] Specifically, Figure 1 A schematic diagram of the structure of the edge-end collaborative video AI inference system in an embodiment of the present invention is shown, and the system includes:
[0050] The edge controller is used to record when different frames of video are input into the neural network, calculate the number of multiplications and additions required by the neural network and the accuracy of neural network recognition; when receiving video recognition requests from various mobile devices, it completes the control task according to the video recognition requests of the mobile devices;
[0051] The edge DNN inference module is used to obtain the uploaded video from the edge device that needs to be offloaded, then complete the inference based on the computing resources allocated by the edge controller, and send the inference results to each mobile device;
[0052] A video sampling management module is used to obtain sampling frame control information of edge computing, control the sampling frame number of the mobile device based on the sampling frame control information, and determine the frame number of the input video for neural network reasoning;
[0053] A local controller, used for obtaining videos from the video sampling management module and determining whether to transmit the videos to the edge server according to the user offloading decision obtained from the edge controller;
[0054] The local DNN inference module is used to obtain video and use the allocated local computing resources to complete DNN inference when the mobile device needs to complete DNN inference locally.
[0055] The method of completing the control task according to the video recognition request of the mobile device includes: determining the frame number of video sampling of each mobile device, and sending sampling frame number control information to the video sampling management module corresponding to the mobile device based on the frame number of video sampling of each mobile device; determining the user unloading decision, and determining whether the user's reasoning task is completed in the local DNN reasoning module or in the edge DNN reasoning module based on the user unloading decision; determining the time proportion of uplink video transmission of each mobile device based on the allocation strategy of communication resources of each mobile video device; and determining the resource allocation strategy of the edge DNN reasoning module to determine the CPU calculation frequency of each unloaded user.
[0056] If the user's unloading decision is 1, the local controller transmits the video to the edge server for reasoning, and the communication resources for transmission are configured by the base station; if the user's unloading decision is 0, the video is allowed to complete reasoning locally, and local CPU computing resources are allocated according to local device information.
[0057] based on Figure 1 The system structure shown in the figure is a method for optimizing the configuration of terminal edge joint resources in an embodiment of the present invention, the method comprising: based on the edge controller recording different frames of video input to the neural network, calculating the number of multiplications and additions required by the neural network and the accuracy of neural network recognition; when receiving the video recognition request of each mobile device, completing the control task according to the video recognition request of the mobile device, the control task comprising: determining the number of frames of video sampling of each mobile device, and sending sampling frame control information to the video sampling management module corresponding to the mobile device based on the number of frames of video sampling of each mobile device; determining the user unloading decision, and determining whether the user's reasoning task is completed in the local DNN reasoning module or in the edge DNN reasoning module based on the user unloading decision; determining the time proportion of uplink transmission of video by each mobile device based on the allocation strategy of communication resources of each mobile video device; determining the resource allocation strategy of the edge DNN reasoning module to determine the CPU calculation frequency of each unloaded user.
[0058] Furthermore, the edge DNN reasoning module obtains the uploaded video from the edge device that needs to be unloaded, then completes the reasoning according to the computing resources allocated by the edge controller, and sends the reasoning results to each mobile device.
[0059] Furthermore, the video sampling management module obtains sampling frame control information of edge computing, controls the sampling frame number of the mobile device based on the sampling frame control information, and determines the frame number of the input video used for neural network reasoning.
[0060] Furthermore, the local controller obtains the video from the video sampling management module, and decides whether to transmit the video to the edge server according to the user unloading decision obtained from the edge controller.
[0061] Furthermore, the local controller obtains the video from the video sampling management module, and decides whether to transmit the video to the edge server according to the user unloading decision obtained from the edge controller, including: if the user unloading decision is 1, the video is transmitted to the edge server for reasoning, and the communication resources for transmission are configured by the base station; if the user unloading decision is 0, the video is allowed to complete reasoning locally, and local CPU computing resources are allocated according to local device information.
[0062] Furthermore, when the mobile device needs to complete DNN reasoning locally, the local DNN reasoning module obtains the video and uses the allocated local computing resources to complete the DNN reasoning.
[0063] It should be noted that, based on the above system architecture, the present invention proposes a corresponding optimization scheme, which reduces the average inference delay of the system and the average energy consumption of mobile devices, and improves the accuracy of model inference. The specific algorithm is as follows:
[0064] First, the multiplication and addition number model of the neural network is used to obtain the multiplication and addition number C (M n ), considering that for each layer of the neural network, the number of multiplications and additions required to complete the inference is proportional to the input size, so after derivation, the total number of multiplications and additions can be roughly represented by a linear function of the number of input video frames, recorded as:
[0065] C(M n )=m c,0 M n +m c,1
[0066] Where m c,0 and m c,1 is the parameter obtained by fitting, which is determined by the architecture of the model. n is the sampling frame number control information, and then the calculation delay D of the neural network is given according to the multiplication and addition number n and energy consumption E n The expression is:
[0067]
[0068]
[0069] Where ρ represents the number of CPU revolutions required for each multiplication and addition operation, d represents the size of each frame of video, and R n represents the communication rate of the nth mobile device, t n represents the communication time ratio of the nth mobile device (i.e., the communication resource allocation ratio), κ represents the energy consumption coefficient, and p n represents the transmit power of the mobile device, x n Indicates the user's uninstall decision, Indicates the CPU computing resources allocated to the local computer. represents the computing resources allocated by the edge controller, x n Indicates whether to offload to edge computing (if x n =1 means uninstallation, otherwise it is calculated locally).
[0070] As for accuracy, the more frames of the input video, the higher the accuracy of the model prediction. As the number of video input frames increases, the gain of the model prediction accuracy will gradually decrease. Therefore, the accuracy Φ(M n ) can be expressed as the following function:
[0071]
[0072] Where m a,0 , m a,1 and m a,2 It is a parameter obtained through fitting, and its value is determined by the model and task of the neural network.
[0073] In summary, the optimization objective function is:
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081] Where: β 1 , β 2 , β 3 are weight coefficients respectively, and the optimization goal is to reduce the total delay D n and total energy consumption E n, and improve the user recognition accuracy Φ(M n ), the constraints are: ① the number of user frames is within the given range, ② the sum of the communication time ratios of the devices involved in the offloading is less than 1, ③ the sum of the allocated computing frequencies of the edge devices is less than its upper limit, ④ the allocated communication time and edge computing frequency need to be greater than 0, ⑤ the computing frequency of the mobile device is greater than 0 and less than The value of ⑥x n To uninstall or not to uninstall.
[0082] In order to solve the optimization problem, set the unloading strategy x n is given, and then the problem is decomposed into two sub-problems, namely the resource optimization problem of mobile devices that complete reasoning locally and the resource optimization problem of mobile devices that complete reasoning at the edge. Suppose the set of users who complete reasoning locally is N 0 , the set of users who complete reasoning at the edge is N 1 , so the optimization problem is transformed into two sub-optimization problems.
[0083] For N 0 The resource optimization problem is stated as follows:
[0084]
[0085]
[0086]
[0087] in The cost function for local calculation is checked. The constraints that need to be met are: ① The number of user frames is within the given range, and ② The calculation frequency of the mobile device is greater than 0 and less than Value
[0088] By deduction, we can get the closed-form expression for solving this subproblem:
[0089]
[0090]
[0091] For N 1 The resource optimization problem is stated as follows:
[0092]
[0093]
[0094]
[0095]
[0096]
[0097] in The cost function for offloading to edge computing. The constraints that need to be met are: ① the number of user frames is within the given range, ② the sum of the communication time ratios of the devices participating in the offloading is less than 1, ③ the sum of the allocated computing frequencies of the edge devices is less than their upper limit, and ④ the allocated communication time and edge computing frequency need to be greater than 0.
[0098] By derivation, we can get t n , With M n The relational expression is:
[0099]
[0100]
[0101] Then the problem can be solved by convex optimization method.
[0102] For uninstall policy x n The solution problem of this invention is implemented based on a greedy iterative algorithm. It can be observed that when performing inference locally, the cost function and optimization variable M n , It only depends on the parameters of the device itself and is not affected by other device parameters. Cost Functions and Sets The number of devices in the system and the parameters are related. The following is an introduction to the principle of the algorithm. First, calculate the set of tasks for each device when they are executed locally. Cost function Secondly, all devices are offloaded to the edge server for reasoning and In each iteration, we get The cost function corresponding to each device Comparing cost functions and the cost function exist The devices in the collection can be and and select the device with the largest difference as device y. Add the device y in the collection And calculate the cost of the new set. If the total cost of the new set is lower, continue to the next iteration. Otherwise, put device y back into the set The algorithm ends.
[0103] The downlink bandwidth is set to 5Mhz, and the path loss is modeled as PL=128.1+37.6log 10 (D), where D is the distance between the device and the wireless access point in kilometers. The devices are randomly distributed within the range of [500m 500m]. The computing resources of the MEC server and the device are set to 1.8GHz and 22GHz respectively. The recognition accuracy requirement and the maximum number of input video frames are set to Coefficient κ = 10 28 , determined by the corresponding device. The size of the input video is 112*112*M n In addition, the computational complexity coefficient is set to ρ = 12 which is obtained through multiple experiments. 1 , β 2 , β 3 Set them to 0.2, 0.2, and 0.6 respectively.
[0104] The proposed offloading scheme is compared with the local inference scheme (Local), the edge inference scheme (Edge), and the random offloading scheme (Random). The experimental results are as follows: Figure 2 As shown in the figure. When the number of devices is less than 10, the cost of the solution that only performs tasks on the edge is almost equal to the cost of the proposed offloading solution. This is because when the number of devices is small, all devices can benefit from performing inference on the edge server. If the inference task is only performed locally, the average cost of the device will not change because the local resources between devices do not affect each other. The curve of the Edge solution is linear because the AI model of all users in the experiment is the same.
[0105] This experiment uses different weights β 1 , β 2 , β 3 to analyze the trade-off between average latency, energy consumption, and accuracy. The performance of the trade-off surface is obtained by the proposed offloading and allocation scheme, with the constraint β 1 +β 2 +β 3 =1. Figure 3 As shown in the figure, latency, energy consumption, and accuracy are mutually restrictive and are traded off against each other. When latency is constant, higher recognition accuracy requires higher energy consumption. From another perspective, in order to improve accuracy, latency and energy consumption performance need to be sacrificed. In addition, with the same accuracy, higher energy consumption will make the device more inclined to perform reasoning tasks locally, thereby reducing latency.
[0106] The embodiments of the present invention have the following beneficial effects:
[0107] (1) The present invention provides an architecture of an edge-cooperative video AI inference system based on an edge-cooperative AI inference algorithm offloading architecture. The edge server can determine the number of video frames used for detection by the user according to the number of requesting users, and provide the user's offloading strategy and the user's communication computing resource allocation plan.
[0108] (2) Multi-dimensional performance optimization. Under the given system architecture, an effective algorithm is proposed by jointly considering inference delay, terminal energy consumption and recognition accuracy. It can improve the recognition accuracy of the neural network while reducing delay and energy consumption.
[0109] (3) Performance trade-off analysis. The invention provides a trade-off relationship between inference latency, terminal energy consumption and recognition accuracy. By using this relationship, targeted system optimization can be performed to improve the system performance of AI reasoning in different applications.
[0110] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0111] In addition, the embodiments of the present invention are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for optimizing the allocation of joint resources at the terminal edge, characterized in that, the method includes: Based on the edge controller recording the multiplication and addition numbers required by the neural network and the recognition accuracy of the neural network when different frames of video are input into the neural network; when receiving video recognition requests from each mobile device, completing control tasks according to the video recognition requests of the mobile devices, and the control tasks include: Determining the number of frames for video sampling of each mobile device, and sending sampling frame number control information to the corresponding video sampling management module of the mobile device based on the number of frames for video sampling of each mobile device; Determining the user offloading decision, and determining whether the user's inference task is completed in the local DNN inference module or in the edge DNN inference module based on the user offloading decision; Determining the time ratio of the uplink transmission of video for each mobile device based on the allocation strategy of the communication resources of each mobile video device; Determining the resource allocation strategy of the edge DNN inference module to decide the CPU computing frequency of each offloading user.
2. The method for optimizing the allocation of joint resources at the terminal edge according to claim 1, characterized in that, the method further includes: The edge DNN inference module obtains the uploaded video from the edge devices that need to be offloaded, then completes the inference according to the computing resources allocated by the edge controller, and sends the inference results to each mobile device.
3. The method for optimizing the allocation of joint resources at the terminal edge according to claim 1, characterized in that, the method further includes: The video sampling management module obtains the sampling frame number control information of edge computing, controls the sampling frame number of the mobile device where it is located based on the sampling frame control information, and determines the number of frames of the input video for neural network inference.
4. The method for optimizing the allocation of joint resources at the terminal edge according to claim 1, characterized in that, the method further includes: The local controller obtains the video from the video sampling management module, and decides whether to transmit the video to the edge server according to the user offloading decision obtained from the edge controller.
5. The method for optimizing the allocation of joint resources at the terminal edge according to claim 4, characterized in that, The local controller obtains the video from the video sampling management module, and decides whether to transmit the video to the edge server according to the user offloading decision obtained from the edge controller, including: If the user offloading decision is 1, then transmit the video to the edge server for inference, and the communication resources for transmission are configured by the base station; If the user offloading decision is 0, then allow the video to complete the inference locally, and allocate the local CPU computing resources according to the local device information.
6. The method for optimizing the allocation of joint resources at the terminal edge according to claim 1, characterized in that, When the mobile device needs to complete DNN inference locally, the local DNN inference module obtains the video and uses the allocated local computing resources to complete DNN inference.
7. A video AI inference system for edge-end collaboration, characterized in that, the system includes: An edge controller, which is used to record the number of multiply-accumulate operations required by the neural network and the recognition accuracy of the neural network when videos with different numbers of frames are input into the neural network; when receiving video recognition requests from various mobile devices, it completes control tasks according to the video recognition requests of the mobile devices; An edge DNN inference module, which is used to obtain the uploaded videos from the edge devices to be offloaded, then complete inference according to the computing resources allocated by the edge controller, and send the inference results to each mobile device; A video sampling management module, which is used to obtain the sampling frame number control information for edge computing, control the sampling frame number of the mobile device where it is located based on the sampling frame control information, and determine the number of frames of the input video for neural network inference; A local controller, which is used to obtain videos from the video sampling management module and decide whether to transmit the videos to the edge server according to the user offloading decision obtained from the edge controller; A local DNN inference module, which is used to obtain videos when the mobile device needs to complete DNN inference locally, and complete DNN inference using the allocated local computing resources.
8. The edge-end collaborative video AI inference system according to claim 7, wherein, the completion of control tasks according to the video recognition requests of the mobile devices includes: determining the number of frames of video sampling for each mobile device, and sending sampling frame number control information to the corresponding video sampling management module of the mobile device based on the number of frames of video sampling for each mobile device; determining the user offloading decision, and determining whether the user's inference task is completed in the local DNN inference module or in the edge DNN inference module based on the user offloading decision; determining the time ratio of uplink video transmission for each mobile device based on the allocation strategy of communication resources for each mobile video device; determining the resource allocation strategy of the edge DNN inference module to decide the CPU computing frequency of each offloading user.
9. The edge-end collaborative video AI inference system according to claim 8, wherein, when the user offloading decision is 1, the local controller transmits the video to the edge server for inference, and the communication resources for transmission are configured by the base station.
10. The edge-end collaborative video AI inference system according to claim 8, wherein, when the user offloading decision is 0, the local controller allows the video to be inferred locally and allocates local CPU computing resources according to the local device information.
Citation Information
Patent Citations
Method and system for hybrid deployment of depth learning neural networks on terminals and clouds
CN109543829A
Deep learning model reasoning acceleration method based on cooperation of edge server and mobile terminal equipment
CN110309914A