Task Offloading Method for Multi-View Inference Applications in Edge Computing Environments
By building an execution framework for multi-perspective inference tasks in an edge computing environment, dynamically dividing the model location and resource allocation ratio, the problem of extended execution time of multi-perspective inference tasks in the existing technology is solved, and more efficient resource utilization and low-latency task completion is achieved.
Patent Information
- Application Number
- CN202310136701.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-02-20
AI Technical Summary
The existing edge computing execution framework cannot dynamically divide the location and resource allocation ratio for models with better decision-making in each perspective based on the execution mode and environmental characteristics of multi-perspective inference tasks, resulting in strong terminals waiting for weak terminals and increasing task execution time.
Build an execution framework for multi-perspective inference tasks in an edge computing environment. Through system information collection, offload modeling analysis and task offload decisions, dynamically divide the location and resource allocation ratio of multi-perspective deep neural network models, and optimize the execution time of multi-perspective inference tasks.
Through dynamic resource allocation and model division, the resource utilization rate of terminal devices can be improved, the completion time of multi-view reasoning tasks can be reduced, and the low-latency requirements can be met.
Smart Images

Figure CN116166336B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of edge computing and deep learning, and specifically to a task offloading method for multi-view inference applications in an edge computing environment. Background Art
[0002] Since AlexNet emerged in the 2012 ImageNet competition, image analysis technologies based on artificial neural networks have gradually attracted people's attention and received extensive research. Researchers have found that as the number of layers of artificial neural networks increases within a certain range, the image analysis accuracy shows an upward trend. As a result, deep neural networks (DNNs) represented by VGG and GoogleNet have emerged. However, as the number of neural network layers continues to increase, the classification progress of DNNs will gradually decrease, that is, network degradation occurs, and the Residual Network (ResNet) is proposed. Since the above neural network models all accept only one image as input, these models can be collectively referred to as single-view models. Single-view models have achieved very good or even better results than human eye recognition in conventional image recognition scenarios. However, in some specific scenarios, the recognition effect is not good. For example, in the field of public safety, computer vision-based crowd density detection and social distance monitoring have received extensive attention and research. However, due to crowd congestion and easy occlusion, it is easy to miss people whose bodies are occluded during the detection process, resulting in poor crowd density detection and social distance monitoring effects and potential safety hazards. Moreover, in the field of industrial production, taking the steel billet flaw detection of Nanjing Iron and Steel as an example, in the steel billet flaw detection task of Nanjing Iron and Steel, the detection system often misdetects watermarks and oil prints as lamination. The reason is that the detection system takes images from directly above the steel billet, and there is not much difference between watermarks and oil prints and lamination when observed from the front. Therefore, the steel billet detection of Nanjing Iron and Steel has a high false detection rate, resulting in an increase in the cost of manual re-inspection and a reduction in the production efficiency of the enterprise.
[0003] Just because single-view (Single-View) cannot achieve good recognition effects in the above but not limited to the above scenarios, researchers have proposed whether multi-view image acquisition can be performed on the same recognition target and the advantages of multi-view image complementarity can be used to comprehensively recognize the target, that is, multi-view image recognition (Multi-View). Such as Figure 2As shown, it is a multi-view deep learning model structure. For the multi-view model, its inference computation amount is positively correlated with the number of views, and its computational intensity is much higher than that of single-view inference, resulting in low inference efficiency. In order to perform multi-view inference in real time and efficiently, it is necessary to accelerate multi-view inference. Currently, there are two relatively common paradigms for executing intelligent tasks, namely: 1) The local computing paradigm. The multi-view images are collected by the terminal and sent to the local server with multi-view inference capabilities, and the local server performs multi-view inference to obtain the results. However, the insufficient resources and limited computing power of the local server will lead to poor response ability of the local server. In addition, with the continuous development of deep learning and the rapid popularization of intelligent terminals with certain inference capabilities such as intelligent cameras and intelligent sensors, using the local computing paradigm cannot effectively integrate the computing power of intelligent terminals; 2) The cloud computing paradigm. The images collected by the terminal are directly sent to the cloud computing, and the cloud computing center completes multi-view inference. Although the execution process of the multi-view inference task can enjoy the support of the powerful computing power and rich resources of the cloud computing center, due to the long physical distance between the cloud computing center and the terminal and the need for the data transmission to pass through the backbone network, when the multi-view inference tasks are submitted frequently, it will cause congestion in the backbone network, and then a relatively high latency will be generated, making it difficult to meet the real-time requirement.
[0004] In order to solve the shortcomings of the traditional computing paradigm in intelligent tasks such as multi-view inference, the industry and academia have proposed edge computing. Edge computing proposes that edge servers with computing capabilities can be deployed near the data source to provide the nearest service. The edge server is physically close to the terminal and is generally in the same local area network. The use of edge computing allows computing tasks to obtain execution results directly without going through the backbone network, greatly alleviating the congestion of the backbone network, reducing the burden on the cloud computing center, and reducing the task execution delay. At present, some work has focused on using edge computing architecture to support intelligent applications of smart terminals, but existing research is all focused on single-view inference tasks. Since the computation offloading method based on single-view inference tasks cannot directly accelerate multi-view inference tasks, it will inevitably lead to a waste of energy and computing resources and increase the response time of the application. The limitations of the existing edge computing execution framework and offloading mechanism are mainly reflected in the following two points: (1) According to the execution mode of multi-view inference, multi-view inference has the characteristic of shallow inference parallelism. Directly applying the task offloading method of single-view inference to multi-view inference cannot consider the characteristics of the multi-view inference execution model. (2) In a multi-view reasoning execution system, there may be heterogeneity of terminal devices and dynamic network connections between terminal devices and edge servers. Existing offloading methods cannot divide the location and resource allocation ratio of the terminal decision model based on the strength of the terminal in a multi-view execution environment, which will cause strong terminals to wait for weak terminals for a long time, thereby extending the execution time of the task. Therefore, the existing edge computing execution framework and offloading mechanism cannot take into account the characteristics of the multi-view reasoning execution mode and the characteristics of the execution environment, and have great limitations, and cannot meet the low latency requirements of multi-view reasoning applications. Summary of the invention
[0005] Technical problem: In order to overcome the problem pointed out in the background technology that the unloading mode in the existing edge computing execution framework does not take into account the multi-perspective reasoning execution mode and environmental characteristics, and cannot dynamically decide the location and resource allocation ratio of the model with better performance for each perspective according to the strength of the terminal devices involved in the multi-perspective reasoning and the network connection status, thereby causing the strong terminal to wait for the weak terminal for a long time, resulting in the problem of prolonged execution time of the multi-perspective reasoning task. The present invention proposes a task offloading method for multi-perspective reasoning applications in an edge computing environment, which implements the multi-perspective reasoning execution model and environmental characteristics, dynamically decides the location and resource allocation ratio of the model with better performance for each perspective according to the strength of the terminal device and the network connection status, improves the resource utilization of the terminal device, and minimizes the completion time of the multi-perspective reasoning task, thereby meeting the low latency requirements of the application.
[0006] Technical solution: The present invention is a task offloading method for multi-view inference applications in an edge computing environment, which includes four parts, namely, constructing an edge computing execution framework for multi-view inference tasks, system information collection, offloading modeling analysis, and task offloading decision-making. The specific implementation method is as follows:
[0007] In the part of constructing the edge computing execution framework for multi-view inference tasks, the present invention first constructs an edge computing execution framework for multi-view inference applications in the edge computing environment. Combining the multi-view inference execution mode, the multi-view neural network model is divided into N + 1 blocks, where N represents the number of views. Block i is deployed on the terminal device i, and all blocks are deployed on the edge server. The main execution steps of this execution framework are as follows:
[0008] Step 1, after receiving the task, the multi-view terminal device collects the current network parameters, sends the network parameters and computing power parameters to the edge server, and enters Step 2.
[0009] Step 2, the task offloading decision maker decides the model partitioning location and resource allocation ratio according to the network parameters and computing power parameters obtained in Step 1, allocates resources for the corresponding shallow inference modules of each terminal on the edge server according to the resource allocation ratio, and sends the model partitioning location to each terminal, and enters Step 3.
[0010] Step 3, the terminal device i executes the inference from the 1st to the jth layer of block i according to the model partitioning location j obtained in Step 2, and transmits the inference output of the jth layer and the model partitioning location j to the edge server, and enters Step 4.
[0011] Step 4, the shallow inference module of device i in the edge server receives the intermediate result and the model partitioning location transmitted in Step 3 and executes the remaining shallow inference, and transmits the shallow inference result to the deep inference module in the edge server, and enters Step 5.
[0012] Step 5, the deep inference module of the edge server collects the shallow inference results of each view obtained according to Step 4, fuses the shallow inference features, obtains the fused features, and enters Step 6.
[0013] Step 6, the deep inference module of the edge server performs deep inference on the fused features and obtains the inference result.
[0014] In the part of system information collection, based on the constructed edge computing execution framework for multi-view inference tasks, the computing resource data features of the edge server and the computing resource data features of the terminal device in the edge computing environment are collected; the number of layers of the multi-view model and the computing volume data features of each layer are collected; the network transmission ability features between the terminal device and the edge server are collected. The specific steps are as follows:
[0015] Step 1: Reasonably layer the shallow model of the multi-view model. Count the number of layers of the shallow model of the multi-view model as A, and count the inference computation amount l of the j-th layer of the shallow model. j Count the output size of the j-th layer as s. j ;
[0016] Step 2: Obtain the CPU performance of each layer in the edge computing execution architecture, including the CPU performance F of the edge server and the CPU performance f of the terminal device i. i ;
[0017] Step 3: Monitor the current network performance, and obtain the network bandwidth b from the terminal device i to the edge server i and the network latency d. i ;
[0018] In the offloading modeling analysis part, use the collected feature data as input parameters to construct an optimization model that minimizes the execution time of the multi-view inference task. The specific steps are as follows:
[0019] Step 1: Use the split position p of the shallow model of the terminal device i i as a parameter to divide the shallow model on the terminal device i into two parts. The inference computation of the 1st to p i layers is executed on the terminal device, and the inference computation from the p i th layer to the A-th layer needs to be offloaded to the edge server for execution. Then, the execution time of the shallow inference and the transmission time of the intermediate data on each terminal device are respectively and
[0020]
[0021]
[0022] Step 2: Use the resource ratio r allocated by the edge server for shallow inference execution for view i i as a parameter. Then, the remaining shallow inference time for the edge server to execute for the terminal device i is where R is the allocation granularity feature of the edge server's computing power;
[0023] Step 3: According to Step 1 and Step 2, the shallow inference time in each view can be obtained as
[0024] Step 4: Denote the deep inference execution time and the multi-view inference task execution time as and T(p, r)
[0025]
[0026] Step 5. With the objective of minimizing the completion time of the multi-view inference task, the optimization model for minimizing the execution time in the task can be expressed as:
[0027] min(T(p, r))
[0028]
[0029] In the task offloading decision part, a heuristic algorithm based on the Stackelberg game theory is used to solve the constructed optimization model for minimizing the completion time of the multi-view inference task, and the model price position and resource allocation ratio are obtained. The specific steps are as follows:
[0030] Step 1. Transform the optimization model into the following problem
[0031]
[0032] where C i (p, r) represents the loss of view i under the offloading strategy (p, r), and the optimization problem is transformed into a problem of minimizing the maximum loss;
[0033] Step 2. Initialize the model partitioning position solution p′;
[0034] Step 3. Solve the problem Obtain the optimal resource allocation solution r′ when the model partitioning position solution is p′, and enter Step 4;
[0035] Step 4. Solve the problem Obtain the optimal model partitioning position p″ when the resource allocation ratio is r′. Judge whether p″ is equal to p′. If they are equal, the obtained model partitioning and resource allocation strategy (p′, r′) is the target solution; otherwise, update p′ with p″ and enter Step 3.
[0036] Beneficial effects:
[0037] The effectiveness of the present invention lies in:
[0038] Through a task offloading method for multi-view inference applications in an edge computing environment, considering the execution characteristics of multi-view inference tasks, the computing and storage resources of the terminal and the edge server are effectively utilized, and the overall completion time of the multi-view inference task is significantly reduced.
[0039] Compared with the prior art, the present invention has the following advantages:
[0040] 1. By deploying the multi-view deep neural network model in an edge computing environment, the present invention realizes the divisibility of multi-view deep learning computing, providing a basis for a fine-grained multi-view task offloading method;
[0041] 2. The multi - perspective task offloading method of the present invention takes into account the characteristics of the multi - perspective reasoning task execution mode and environment, and considers the existing terminal heterogeneity and network dynamics. Therefore, it can effectively utilize the computing resources of the terminal and reduce the task execution time.
[0042] 3. The multi - perspective task offloading algorithm of the present invention is simple, effective, highly practical, and has low complexity. Therefore, it can be applied to large - scale task environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic diagram of the overall process framework for multi - perspective reasoning applications of the present invention;
[0044] Figure 2 is a schematic diagram of multi - perspective model segmentation of the present invention;
[0045] Figure 3 is an execution architecture diagram for multi - perspective reasoning applications of the present invention;
[0046] Figure 4 is a schematic diagram of problem transformation of the present invention;
[0047] Figure 5 is a flowchart of the execution of the task offloading algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS:
[0048] To deepen the understanding of the present invention, the following detailed description is given in conjunction with the accompanying drawings for this embodiment.
[0049] Embodiment 1: As Figure 1 shown, the present invention constructs an edge - computing execution framework for multi - perspective reasoning applications in combination with the multi - perspective deep neural network reasoning mode, including four logical steps: model training, shallow reasoning of terminal devices, task offloading, and edge - side reasoning. When training the multi - perspective model, the deep neural network is split into N + 1 model blocks and distributedly deployed at different positions in the "terminal - edge" of the edge - computing architecture, where N represents the number of perspectives. Preferably, in the present invention, block i is deployed on the terminal device corresponding to perspective i to collect data for shallow reasoning for perspective i; all blocks from block 1 to block N + 1 are deployed on the edge server, so that the edge server has both the ability to assist the terminal device in completing the complete shallow reasoning and the ability to complete feature fusion and deep reasoning.
[0050] Embodiment 2: As Figure 2 shown, the present invention combines the execution characteristics of multi - perspective reasoning tasks and splits the multi - perspective deep - learning model into N + 1 model blocks, where N represents the number of perspectives. Among them, block i (1 ≤ i ≤ N) represents the shallow - reasoning model of perspective i, and block N + 1 represents the multi - perspective feature - fusion and deep - reasoning model.
[0051] Embodiment 3: As Figure 3 shown, a task offloading method for multi-view inference applications in an edge computing environment disclosed by the present invention is used in the task offloading step of the execution framework. When a multi-view inference task is issued, the terminal devices participating in the multi-view inference will send the computing power of the acquisition device and the network communication status to the edge server. The task offloading decision maker of the edge server will make a decision, send the model partitioning location to the terminal devices, and allocate resources for the edge server to assist the terminal in performing the shallow inference model according to the resource allocation strategy. After receiving the decision result, the terminal device will complete the corresponding shallow inference, and send the intermediate result of the shallow inference and the model partitioning location to the corresponding shallow inference module of the edge server. This module will complete the remaining shallow inference and send the inference result to the deep inference module of the edge server. After receiving all the shallow inference results, the deep inference module of the edge server will perform multi-view feature fusion on the shallow inference eigenvalue and perform deep inference, and finally obtain the multi-view inference result. For the task offloading method described in the present invention, through the modeling analysis of the system, according to the computing power and network communication status of the terminal device, it is decided how much proportion of tasks of each view are directly executed on the terminal device, improving the resource utilization rate of the terminal resources, reducing the waiting time of the strong terminal for the weak terminal, so as to reduce the task completion time.
[0052] After the edge computing execution framework of the present invention is constructed, in the system information collection step, it is necessary to collect the edge server computing resource data characteristics and terminal device computing resource data characteristics in the edge computing environment based on the constructed edge computing execution framework for multi-view inference tasks; collect the number of layers of the multi-view model and the computing volume data characteristics of each layer; collect the network transmission performance characteristics between the terminal device and the edge server.
[0053] Embodiment 4: In the offloading modeling analysis part, the data obtained in the system information collection part is used as input parameters to establish an optimization model for minimizing the multi-view inference task completion time. The specific modeling process is as follows:
[0054] Let the model partitioning location obtained by view i be p i , and the resource allocation ratio be r i , which means that the terminal corresponding to view i needs to complete the inference calculation of the first layer to the p i th layer of the shallow model of view i. Its corresponding shallow inference module of the edge server will assist it to complete the shallow inference from the (p i + 1)th layer to the A-th layer (A represents the shallowness of the shallow model), so the shallow inference time of the device side of view i can be obtained The intermediate data transmission time The shallow inference time of the edge server are respectively:
[0055]
[0056]
[0057]
[0058] where f i represents the computing resource data feature of device i, and l j represents the inference computation volume of the j-th layer of the shallow model, and s k represents the output size of the k-th layer of the model. F represents the computing resource data feature of the edge server, R represents the allocation granularity feature of the edge server computing power, and b i and d i represent the bandwidth and latency of the network connection between device i and the edge server, respectively.
[0059] The deep inference completion time is T es_hl , and we get where B represents the inference computation volume of the deep model. Combining with the execution mode of multi-view inference, the execution time of the multi-view inference task can be obtained as where p = (p i ) 1≤i≤N , r = (r i ) 1≤i≤N , so the optimization model for minimizing the execution time of the multi-view task is
[0060] min(T(p, r))
[0061]
[0062] Furthermore, denote Transform the original optimization model into
[0063]
[0064] where
[0065]
[0066] As Figure 4 shown, transform this optimization problem into a two-stage Stackelberg game problem. In the first stage, solve the optimal resource allocation strategy by fixing the model partitioning position, and in the second stage, solve the model partitioning position by fixing the resource allocation strategy. Through two-stage iterative loops, finally obtain the equilibrium solution of the game problem, and prove the existence of the equilibrium solution of this problem based on game theory.
[0067] In the first stage, solve the optimal resource allocation strategy by fixing the model partitioning position as p′, that is The preferred embodiment of the present invention uses a greedy algorithm to solve the problem, and the algorithm is as follows:
[0068]
[0069] That is, initialize by allocating one resource to each perspective, and then continuously allocate one resource to the perspective with the largest C i value until all resources are allocated.
[0070] In the second stage, the optimal model partitioning position is solved through the resource allocation ratio r′, that is the algorithm is as follows:
[0071]
[0072] That is, an optimal model partitioning position is selected for each perspective by means of traversal.
[0073] Finally, as Figure 5 shown, the above two stages are iteratively looped to finally obtain the equilibrium solution of the game problem.
[0074] The specific steps are as follows:
[0075] Step 1, initialize the model partitioning position solution p′;
[0076] Step 2, solve the problem through Algorithm 1 to obtain the optimal resource allocation solution r′ when the model partitioning position solution is p′, and enter Step 3;
[0077] Step 3, through Algorithm 2, solve the problem to obtain the optimal model partitioning position p″ when the resource allocation ratio is r′. Judge whether p″ is equal to p′. If they are equal, the obtained model partitioning and resource allocation strategy (p′, r′) is the target solution, and the solution is ended; otherwise, update p′ with p″ and enter Step 2.
[0078] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, and these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A task offloading method for multi-view inference applications in an edge computing environment, characterized in that: This method utilizes the characteristics of the multi-view inference execution mode, the different computing capabilities of each terminal in the edge computing environment, the different inference computing amounts of each layer of the deep learning model, and the different output sizes of each layer. On the premise of considering the heterogeneity of multi-view multi-terminal devices and network dynamics, it constructs an edge computing execution framework for multi-view inference tasks, and through the collection and analysis of relevant system parameters, it completes the online offloading decision of dynamically arriving multi-view inference tasks. This method includes the following steps: Step 1) In the edge computing environment, the multi-view deep learning model is reasonably divided into N + 1 blocks according to the execution mode of multi-view inference, where N represents the number of views, and is distributed at different positions in the "terminal-edge". Specifically: Block i, where 1 ≤ i ≤ N, is deployed on terminal device i, and blocks 1 to N + 1 are deployed on the edge server; Based on this, an edge computing execution framework for multi-view inference tasks is constructed: The real-time inference task can make an online decision on the model division position and resource allocation ratio according to the terminal computing power and network status to achieve the acceleration of multi-view inference tasks; Step 2) Based on the edge computing execution framework for multi-view inference tasks constructed in Step 1), relevant data in the system is collected and corresponding features are analyzed, specifically including: the computing resource features of the edge server in the edge computing environment, the computing resource features in the terminal device computing, the number of layers of the multi-view model and the computing amount of each layer, and the network transmission capacity features in the edge computing environment; Step 3) Taking the feature data obtained in Step 2) as input parameters, according to the execution mode of multi-view inference, mathematical models are established for the shallow inference time, intermediate data transmission time, and multi-view deep inference time of each view in multi-view inference; Further analysis shows that the execution time of the multi-view inference task is the maximum value of the completion time of the shallow inference of each view plus the deep inference time; Taking the minimum value of the above multi-view inference task execution time as the objective function and the feature data obtained in Step 1) as the constraint conditions, an optimization model for minimizing the multi-view inference task completion time is constructed; Step 4) Based on the Stackelberg game theory, a heuristic algorithm is used to solve the task completion time minimization model obtained in Step 3) to obtain a better model division position and resource allocation scheme.
2. The task offloading method for multi-view inference applications in an edge computing environment according to claim 1, characterized in that: In the said Step 1), for the construction of the edge computing execution framework of the multi-view inference task, the task offloading decision-maker is located at the edge server side. After receiving the task, the terminal device will obtain the task offloading strategy, that is, the model division position, through the task offloader. After the terminal executes the shallow inference before the model division position, it pushes the intermediate result and the model division position to the edge server; The remaining shallow inference is completed by the corresponding shallow inference module in the edge server. After receiving all the shallow inference features, the deep inference module in the edge server will fuse the features and perform deep inference; Step 1) includes the following steps: Step 101) After receiving the task, the multi-view terminal device collects the current network parameters, sends the network parameters and computing power parameters to the edge server, and enters Step 102); Step 102) The task offloading decision-maker decides the model partitioning location and resource allocation ratio according to the network parameters and computing power parameters obtained in Step 101, allocates resources for the corresponding shallow inference modules of each terminal on the edge server according to the resource allocation ratio, and sends the model partitioning location to each terminal, and enters Step 103); Step 103) The terminal device i executes the inference from the 1st to the jth layer of the i-th block according to the model partitioning location j obtained in Step 102, and transmits the intermediate inference result of the jth layer and the model partitioning location j to the edge server, and enters Step 104); Step 104) The shallow inference module corresponding to device i in the edge server receives the intermediate result and model splitting location transmitted in Step 103, executes the remaining shallow inference, and transmits the shallow inference result to the deep inference module in the edge server, and enters Step 105); Step 105) The deep inference module in the edge server collects the shallow inference results of each view obtained according to Step 104, fuses the shallow inference features to obtain fused features, and enters Step 106); Step 106) The deep inference module in the edge server further infers the fused features and obtains the multi-view inference results.
3. The task offloading method for multi-view inference applications in an edge computing environment according to claim 1, characterized in that: Step 2) is used to collect the parameters required for modeling in Step 3), and Step 2) includes the following steps: Step 201) Reasonably layer the shallow model of the multi-view model. Statistically, the number of layers of the shallow model of the multi-view model is A, and the inference calculation amount l of the j-th layer of the shallow model is statistically calculated j , and the output size s of the j-th layer is statistically calculated j ; Step 202) Obtain the CPU performance of each layer in the edge computing execution architecture, including the CPU performance F of the edge server and the CPU performance f of the terminal device i i ; Step 203) Monitor the current network performance to obtain the network bandwidth b from the terminal device i to the edge server i and the network latency d i .
4. The task offloading method for multi-view inference applications in an edge computing environment according to claim 1, characterized in that: The method for constructing an optimization model that minimizes the completion time of the multi-view inference task is: Step 301) Take the shallow model segmentation position p of the terminal device i i As a parameter, divide the shallow model corresponding to the perspective i into two parts. The inference calculation from the first layer to the p i th layer is executed on the terminal device, and the inference calculation from the p i th layer to the A-th layer needs to be offloaded to the edge server for execution. Then, the shallow inference time and the time for transmitting intermediate data of each terminal device are respectively and Step 302) Use the resource ratio r allocated by the edge server for performing shallow inference from perspective i i as a parameter, then the edge server performs the remaining shallow inference time for the terminal device i where R is the allocation granularity feature of the computing power of the edge server; Step 303) According to Step 301) and Step 302), the shallow inference time for each perspective in Step 3) described in Claim 1 can be obtained as Step 304) The deep inference execution time in Step 3) described in Claim 1 is The multi-perspective inference task execution time is T(p, r) Step 305) In Step 3) described in claim 1, with minimizing the execution time of the multi-view inference task as the objective function, the optimization model for minimizing the task execution time can be expressed as: min(T(p, r)) 5. The task offloading method for multi-view inference applications in an edge computing environment according to claim 1, characterized in that: In Step 4), a heuristic algorithm based on the Stackelberg game theory is used to obtain the model splitting location and resource allocation ratio, and Step 4) includes: Step 401) Transform the optimization problem obtained in Step 305) into the following Among them C i (p, r) represents the loss of view i when unloading policy (p, r), and the optimization problem is transformed into minimizing the maximum loss problem; Step 402) Initialize the model partitioning location solution p′; Step 403) Solve the problem Obtain the optimal resource allocation solution r′ when the model partition position solution is p′, and proceed to Step 404); Step 404) Solve the problem Obtain the optimal model division position p″ when the resource allocation ratio is r′; determine whether p″ is equal to p′. If they are equal, the obtained model division and resource allocation strategy (p′, r′) are the target solutions. Otherwise, update p′ with p″ and enter Step 403).
Citation Information
Patent Citations
Task unloading method for deep learning application in edge computing environment
CN110347500A
Unloading scheduling and resource allocation method based on deep reinforcement learning
CN113452625A
Cited By
Multi-terminal-oriented reasoning task cooperative scheduling system and method
CN121560533A