Mobile edge computing dynamic offloading method and device medium for CNN network
By building a synesthesia computing integrated system model in a mobile edge computing environment and using multi-agent deep reinforcement learning algorithms, the computing offloading strategy of the CNN network is dynamically adjusted, and the resource optimization and quality assurance problems of CNN computing tasks in a dynamic environment are solved, achieving efficient computing task execution and system performance improvement.
Patent Information
- Application Number
- CN202510204142.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In dynamic and resource-constrained mobile IoT environments, CNN computing tasks face challenges in offload decision optimization and computing quality assurance, especially when device computing resources are heterogeneous and resource competition between mobile devices.
A dynamic unloading method for mobile edge computing for CNN networks is proposed. By building a synesthesia computing integrated system model, using multi-agent deep reinforcement learning algorithm and Markov decision model, the computational unloading strategy is dynamically adjusted to optimize resource allocation and task segmentation.
In an uncertain network environment and under the conditions of shortage of resources, the smooth execution of computing tasks is achieved, the overall performance and reliability of the system is improved, the delay and energy consumption are reduced, and the equipment is extended.
Smart Images

Figure CN119697701B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computing and communication technology, and in particular to a mobile edge computing dynamic unloading method and device medium for a CNN network. Background Art
[0002] With the popularity of smart terminals in various application scenarios in the Internet of Things, the demand for computing power for computationally intensive and delay-sensitive services based on convolutional neural network (CNN) technology (such as ocean observation and monitoring, forest fire prevention and warning, marine resource exploration and development, etc.) continues to increase. However, due to the limitations of physical space, cost and energy consumption, the local computing resources of terminal devices are difficult to meet the high computing demand marine applications. Mobile edge computing (MEC) uses the computing resources of edge servers (MESs) to reduce the computing load of mobile devices and becomes a feasible solution. Due to the sparse deployment of devices caused by cost, the computing and cache resources of MES are limited and unevenly distributed. With the increase in the number and computing requirements of mobile devices (MMDs), MES faces a high risk of overload and QoS cannot be guaranteed. Therefore, in actual application scenarios, considering the heterogeneity of CNN tasks on smart mobile devices, the seriality of CNN network structure, dynamic scheduling and resource allocation between multiple MESs, and providing the best partial offloading decision to cope with the dynamic MMEC environment are particularly necessary.
[0003] The network structure dependency of CNN computing tasks makes its offloading decision face more complex practical application challenges than traditional tasks. Multi-access edge computing strategies based on layering, such as hybrid edge computing solutions that integrate ground and air resources and fog-based three-layer edge computing architecture solutions, can help improve the quality of data transmission and computing task execution in dynamic MMECs. However, this type of strategy based on highly abstract computing tasks is difficult to meet specific computing deployment requirements. For example, the study of fine-grained partition offloading of CNN shows that detailed network partitioning can provide more offloading options, help find the optimal solution of the system and accelerate the model inference process. The partition point determination method of CNN network includes constructing an adaptive neural network model, establishing a cost matrix or using a deep learning inference delay model, and enumerating appropriate partition points through a neural network. In particular, for a large CNN image object detection model, a genetic algorithm-based offloading strategy is proposed, which regards each layer of the model as an offloading candidate point and finds the optimal offloading configuration through a genetic search mechanism.
[0004] Although some progress has been made in the research of CNN offloading strategies, environmental factors, heterogeneous computing resources of devices, and resource competition among mobile devices in the dynamic and resource-constrained mobile Internet of Things environment bring challenges to the optimization of offloading decisions and computing quality assurance for CNN computing tasks. Since MEC faces data loss due to unstable communication, the horizontal partitioning strategy of CNN network is difficult to apply due to the frequent data exchange between devices. In addition, with the diversification of the network structure and types of offshore equipment, the complexity of optimizing offloading strategies increases accordingly. Summary of the invention
[0005] In view of this, the purpose of the present invention is to propose a dynamic offloading method for mobile edge computing for CNN networks. For edge computing task scenarios with scarce computing and communication resources and dynamic network environments, a dynamic computing offloading strategy is formulated based on the high computing resource consumption tasks of the CNN network architecture, so as to achieve the smooth execution of computing tasks in an uncertain network environment and under the condition of scarce network resources.
[0006] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0007] The present invention provides a mobile edge computing dynamic unloading method for a CNN network, comprising the following steps:
[0008] Step 1: According to the target detection requirements, image task data, device status data and network status data are obtained in the integrated system of synaesthesia and computing to form a data set;
[0009] Step 2: Based on the synaesthesia-computing integrated system and data set, a synaesthesia-computing integrated system model is constructed. Aiming at the computing task, the synaesthesia-computing integrated system model is transformed into the computing task segmentation and resource allocation problem;
[0010] Step 3: Based on the integrated system model of synaesthesia and computing and the problem of computing task segmentation and resource allocation, a Markov decision model is constructed;
[0011] Step 4: Based on the Markov decision model and the multi-agent deep reinforcement learning algorithm, a joint optimization strategy is obtained, and the neural network parameter model that meets the termination conditions is obtained through training through the data set;
[0012] Step 5: Deploy the trained neural network parameter model to the synaesthesia computing integrated system;
[0013] Step 6: Generate computing task segmentation decisions and computing resource allocation decisions based on real-time image task data, device status data, and network status data, and transmit the decision result information to each node of the integrated synergy computing system, ultimately realizing the offloading of computing tasks.
[0014] Furthermore, the step 1 specifically includes:
[0015] Step 11: According to the task requirements submitted by the user, the integrated system of synergy and computing is composed of multiple mobile devices and sensing devices; the sensing devices in the integrated system of synergy and computing perform data sensing operations, collect and store image task data; the integrated system of synergy and computing simultaneously records the device status data and network status data of the mobile device;
[0016] Step 12: Divide the acquired image task data, device status data and network status data based on time slots to form a time series data set within a set time period; the time slot refers to a time period with a set time length, and the data acquired within the time period are all marked as data of the time slot, and the time series data refers to data with time slot attributes and numbered in the order of time slots;
[0017] Step 13: Divide the time series data set into a training set, a validation set, and a test set.
[0018] Furthermore, the step 2 specifically includes:
[0019] Step 21, the nodes with sensing, computing, communication and storage functions in the integrated system of synergy and computing are connected to each other through wireless and wired communication or wireless communication only to form an integrated system model of synergy and computing; the integrated system model of synergy and computing works according to the time slots divided in step 12;
[0020] Step 22: In the synaesthesia-computing integrated system model based on time slot division, the user initiates a task request in a certain time slot; specifically:
[0021] In the synaesthesia-computing integrated system model based on time slot division, the collection of mobile device MMDs , where device m represents a mobile device MMD selected from multiple mobile devices MMDs to implement task perception, and all of them are loaded with a computing module based on a CNN network. The CNN network contains multiple CNN models, and different CNN models correspond to different computing tasks; the collection of edge servers MESs , wherein device s represents selecting an edge server MES from multiple edge servers MESs to provide computing resources for surrounding mobile devices MMDs; the devices form a network of inter-sensing and computing integration through wired and wireless or only wireless means, the user terminal accesses the network of inter-sensing and computing integration through wired or wireless means, and initiates a task request to the network of inter-sensing and computing integration in a certain time slot; the request includes the type of data to be sensed and the required data processing requirements;
[0022] Step 23: After receiving the task request, the integrated system model of synaesthesia and calculation implements the perception and calculation process; the calculation process consists of data transmission and data processing, and the perception and calculation process is an integrated process of synaesthesia and calculation; specifically:
[0023] The mobile device MMDs are responsible for computing the implementation of the offload decision. First, the mobile device MMDs selects one of the devices Establish communication and offload part of the workload to MES; the entire integrated system model of synergy and computing runs in a time slot manner. As the entire running time of the synaesthesia computing integrated system, the entire running time is divided into time slots of equal length, the time slot duration is expressed as , Indicates rounding down, MMDs in time slot At the beginning, a target detection task is performed; the specific steps are as follows: in each time slot Initially, the device Different target detection tasks are generated according to the requirements submitted by users. The CNN models required for different types of tasks are heterogeneous. Need to choose the right equipment Establish connections and identify task offloading parts, each Receive multiple requests and allocate computing resources to them.
[0024] Furthermore, the step 23 specifically includes:
[0025] Step 231: partition and unload tasks according to the CNN network:
[0026] Assuming that the synaesthesia computing integrated system has a total of For the CNN model CNN models, among which, ,use Indicates The total number of layers in the CNN model, Each layer of the CNN model The properties are ,in, , Respectively Enter the height, width and number of channels, Represent the channel value, kernel size and stride size respectively; the CNN network is mainly composed of convolutional layers and pooling layer The convolutional layer and the pooling layer are unified with express, The total computational effort of the layer is recorded as , the calculation formula is:
[0027]
[0028] In the reasoning process of the CNN network, the output of the previous layer is used as the input of the next layer. The output of the layer is equal to the output of the next layer The amount of input, so the amount of output data is recorded as , the calculation formula is:
[0029]
[0030] in, is the memory usage of unit data;
[0031] Consider that there is only one optimal segmentation point for any CNN network on each MMD, assuming that the segmentation point is , Indicates partition point, the first layer to the second layer are executed locally. layer CNN network, and then execute The intermediate data after the layer is unloaded to the MES, and the MES executes layer until the last layer, and finally transmit the inference result back to MMD; When , all tasks are performed locally; when When the task is directly offloaded to MES;
[0032] Step 232: Obtain the computation delay and transmission delay of the CNN network:
[0033] (1) Execution time on MMD: Indicates the device The computing resources in the time slot The computing resources allocated to local execution applications are , ; Local execution time for:
[0034]
[0035] in, Indicates time slot Medium Equipment Above, first layer and partition layer The sum of the computational effort of all layers between Indicates CNN model The number of partition point layers; Indicates time slot Medium Equipment Previous The amount of computation of the layer;
[0036] (2) Transmission time of intermediate output data:
[0037] equipment To device The channel gain of the communication link between Modeled as:
[0038]
[0039] in, Represents the Euclidean distance between corresponding devices, Indicates the channel gain when the reference distance is 1m; the corresponding transmission rate Defined as:
[0040]
[0041] in, is the transmission bandwidth of the edge computing network, is the transmission power of the UAV, is the waveguide channel noise; let Indicates that during time slot t, device m is in the partition layer The size of the data transferred, get the device Offload to device Transmission delay for:
[0042]
[0043] Considering the computational delay of the target detection task in MES and the queuing delay between multiple tasks; let Indicated in Devices in time slot Assign to device of computing resources, of which, , 'For equipment A collection of MMD devices that provide computing services at the same time. For equipment The total computing resources of Time slot, the execution delay of device m on device s for:
[0044]
[0045] in, Indicates transmission to the device The CNN network partition on The local computation amount of the part after the layer; Indicates time slot Medium Equipment Previous The amount of computation of the layer;
[0046] Therefore, after dividing an object detection task, the device Total execution time It is equal to the sum of the local execution time on MMD, the intermediate data transmission time and the execution time on MES:
[0047]
[0048] Assume that each Each task must be completed within a single time slot, that is ;
[0049] Step 233, calculate energy consumption:
[0050] (1) Execution energy consumption on MMD: Assume Indicates the device The computing power of the device Local execution energy consumption for:
[0051]
[0052] (2) Equipment Energy consumption of intermediate output data transmission for:
[0053]
[0054] in, express The transmission power of
[0055] (3) Equipment The energy consumed in executing the task is:
[0056]
[0057] in, Indicates the device Average power when performing tasks;
[0058] (4) When the device Total energy consumption of executing some CNN network tasks It is the sum of local execution energy consumption, intermediate output data transmission energy consumption and task execution energy consumption:
[0059]
[0060] Step 234, calculate the quality of service:
[0061] (1) Use and They represent the improvement levels of energy consumption and latency of MMDs respectively; Describes the energy saved by calculating the ratio of the local energy consumption to the energy consumption saved by offloading MMDs. It represents the reduction ratio of latency compared with local latency through MESs computing offloading; therefore, the energy consumption and improvement level of MMD are expressed as:
[0062]
[0063] and
[0064]
[0065] in, and Respectively represent the time slot Period equipment Latency and energy consumption when performing tasks entirely locally;
[0066] (2) Introducing weight factors To measure the trade-off between energy saving and latency reduction in MMD, consider the remaining battery power of MMD and compare the remaining power to Further weighting factors are incorporated, where:
[0067]
[0068] in, and Respectively represent the current maximum battery capacity and current remaining power of the device;
[0069] (3) The weighted sum of energy consumption improvement and delay improvement is expressed as:
[0070]
[0071] (4) Using exponential functions to characterize the trade-off model between energy consumption improvement and delay improvement and MMDs service quality The mapping relationship between them is defined as:
[0072]
[0073] in, Represents variable parameters, is a positive integer, then express Index of
[0074] (5) Based on Derived , so it is deduced that The upper and lower bounds of the value are and ;
[0075] Step 235: Define the target problem:
[0076] The scheduling strategy mainly includes real-time dynamic adjustment of the weight ratio of delay and energy consumption according to the remaining battery power of the MMD, and minimizing system delay and energy consumption; each MMD makes a decision , and To implement computing tasks and optimize all MMDs in the system ,in, The decision of device m to select device s is as follows:
[0077]
[0078]
[0079]
[0080]
[0081]
[0082] Combining the equations, it can be inferred that the target problem is a mixed integer nonlinear programming problem, which is also a non-deterministic polynomial-hard problem;
[0083] constraint Indicates the device Assign to device The computing resources must not exceed the maximum available local computing resources; Indicates that the partition point selection of the target detection model must be within the range of the number of layers of the target detection model; Indicates that the selection of MESs by MMD must be from the set Select Indicates that the remaining battery power of the MMD does not exceed its maximum capacity.
[0084] Furthermore, the step 3 specifically includes:
[0085] Step 31: Each device It is regarded as an intelligent agent, and the problem of computing task partitioning and resource allocation is further expressed as a multi-agent Markov decision process, and a Markov decision model is constructed;
[0086] Step 32: Adaptively determine the offloading decision of each device according to the environmental state to maximize the user service quality; the Markov decision model is expressed as ,in, is the state space, is the action space, is the state transition probability matrix, is the reward; the state, action, and reward of the Markov decision model are as follows:
[0087] State: The state space consists of the states of MMD and MES, time slots The state space at time Defined as :
[0088]
[0089] in, For time slot Time equipment The CNN model category carried on, Indicates time slot Assigned to device The computing resources for local execution of applications, express Time slot equipment The remaining power, Indicated in Devices in time slot Assign to device The computing resources are sent to each MMD in the form of broadcast; Indicates the device With equipment Uplink capacity between
[0090] Action: All devices The action set is a mixed integer and non-integer offloading decision vector, where the device In time slot The action is expressed as:
[0091]
[0092] in, Indicates the offloading decision of the MMD offloading task, Represents the partition of the CNN network carried on MMD;
[0093] Reward: In the Markov decision process, the reward is the Status The system benefits obtained by taking actions under Complete the device When the task is transferred, the device Generating Quality of Service Feedback, MESs collaborate to achieve service quality for all devices Maximization is the goal; therefore, the reward function is expressed as:
[0094] .
[0095] Furthermore, the step 4 specifically includes:
[0096] Step 41: Determine the network segmentation nodes based on the network node retention algorithm:
[0097] (1) Select a CNN model j as the preprocessing target and calculate different bandwidths The execution time corresponding to all partition points , ≤ ;definition To preserve the set of partition layers; define For bandwidth The corresponding partition layer set; definition is the candidate point set corresponding to the current CNN model j;
[0098] Calculate the same bandwidth The average execution time of all partition points , as shown below:
[0099]
[0100] (2) Order is a weight factor in the range [0,1], and each node is associated with the average execution time and weight factor The product of is compared; if Less than , then keep the node until ; Otherwise, eliminate the node; finally get the bandwidth Corresponding ;
[0101] (3) Through the formula:
[0102]
[0103] renew , when taking different When the value is reached, repeat (1) to (3) in step 41 until all bandwidths are traversed. , and obtain the final ;
[0104] (4) Candidate point sets for all types of CNN models Integrate into middle;
[0105] Step 42: Joint optimization strategy based on multi-agent deep reinforcement learning algorithm:
[0106] (1) The MADDPG framework is integrated in a multi-agent collaborative environment, where each agent consists of an actor network and a network of critics Both are parameterized by deep neural networks; the actor network takes state observations as input and Output action; at the same time, the critic network uses the state-action function Evaluate the actions chosen by the actor network based on the current state;
[0107] (2) In the MADDPG framework, each target network takes the global states of multiple agents as input to evaluate the quality of actions under these states; the target network promotes stable learning by generating temporarily delayed target policies and value estimates; in the MADDPG framework, the actor network observes the state of its local agent based on the target network’s action quality. Determine local actions , the policy parameters of each agent are defined as ,in, Indicates The strategy parameters of each agent are Parameterized Continuous Strategy , is the actor network parameter of agent m; in the deterministic strategy framework, each agent adopts a continuous strategy ,in is the random noise used to enhance the system’s exploration capabilities;
[0108] (3) The critic network transforms the state of the agent and actions As input, it is represented as ,in, is the network parameter of the critic network, which is used to evaluate the collective performance A global evaluation is performed, whose input includes the actions and states of all agents; the target actor network of the agents is defined as , the corresponding network parameters are defined as , the target network is expressed as , the corresponding network parameters are ; The update formulas for these parameters are defined as:
[0109]
[0110] Among them, the parameters represents a coefficient in the range of 0 to 1;
[0111] (4) The MADDPG algorithm consists of two stages: centralized offline training and decentralized execution;
[0112] During the centralized offline training process, the system collects the state information of the agent ,in Indicates the status is Execute action when The next state entered after that, this state information is stored in the experience playback buffer; is the reward value obtained by the current action; after collecting data into the experience replay buffer, random sampling is used for the policy evaluation process; the MADDPG algorithm follows the offline policy method and uses the sampled data to update and Parameters; the objective function update is expressed as:
[0113]
[0114] in, represents the discount factor, It is through the formula: Get, update the critic network by minimizing the loss function:
[0115]
[0116] in, [ ] means to seek expectation;
[0117] After updating the critic network, the actor network is parameterized by the following formula Update:
[0118]
[0119] in, Representation parameters The gradient of
[0120] The critic network then updates its parameters according to the gradient rule of the neural network, obtaining:
[0121]
[0122] in is the learning rate of the critic network, and then the actor network is updated using the policy gradient method to obtain:
[0123]
[0124] The actor network parameters are updated as:
[0125]
[0126] in, is the learning rate of the critic network; represents the actor network; Represents the actor network Ask about parameters The gradient of
[0127] Step 43. Standardized and adaptive learning rate for multi-agent deep reinforcement learning algorithm:
[0128] (1) During the initialization of the multi-agent deep reinforcement learning algorithm, the environment state is standardized by storing the maximum and minimum values during training in temporary memory and calculating a scaling factor. The difference between the maximum and minimum values of each environment variable is used as the scaling factor. The standardized state information is obtained by subtracting the minimum value from the current state and dividing it by the obtained scaling factor.
[0129] (2) During the reinforcement learning training process, the decay learning rate strategy is used to adjust the learning rate according to the loss rate at different training stages:
[0130]
[0131] in, is the learning rate decay coefficient.
[0132] Furthermore, the step 5 specifically includes:
[0133] Step 51: synchronize the trained neural network parameters to each node in the CNN network at a set time interval according to the environmental dynamic information; the time interval depends on whether the new neural network parameter training process converges to a new state;
[0134] Step 52: Each node in the CNN network updates the unloading strategy and resource allocation strategy based on the obtained neural network parameters;
[0135] Step 53: Each node in the CNN network feeds back the updated strategy status to the integrated synaesthesia and computing system.
[0136] Furthermore, the step 6 specifically includes:
[0137] Step 61: The sensory-computing integrated system obtains real-time network status data and device status data, and the sensing device obtains real-time image task data according to user needs;
[0138] Step 62: Based on the input network status data, device status data and image task data, according to the multi-agent Markov decision model determined by the neural network, generate a decision on the division of computing tasks and the allocation of computing resources;
[0139] Step 63, synchronizing the decision result information to each node in the integrated system of synergy and computing;
[0140] Step 64: Each node implements and executes a specific computing task.
[0141] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the dynamic unloading method for mobile edge computing for a CNN network as described above is implemented.
[0142] The present invention also provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the dynamic unloading method of mobile edge computing for CNN network as described above is implemented.
[0143] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art:
[0144] (1) This paper proposes a weighted sum model of delay improvement and energy consumption improvement of marine intelligent mobile devices under resource constraints, and optimizes the unloading decision by dynamically adjusting the system weight. This method not only considers the trade-off between delay and energy consumption, but also pays special attention to the remaining power of the device, aiming to extend the use time of the device, thereby improving the overall performance and reliability of the MMEC system.
[0145] (2) The MINLP (mixed integer nonlinear programming) problem is modeled as MAMDP. Considering the complexity of the target problem, this paper proposes an innovative partial offloading strategy that reduces the data transmission demand by setting a single split point in the CNN network, which is particularly suitable for situations where data transmission is limited in remote environments. Combined with the network node reservation algorithm (NPPR), the decision-making process is effectively simplified and the implementation efficiency of the offloading strategy is improved.
[0146] (3) The present invention proposes a joint optimization strategy based on multi-agent deep reinforcement learning (MADRL) for the collaborative operation of drones and sea surface buoys in the marine IoT environment. By treating each marine mobile device as an independent agent, the CNN is vertically divided and the resource allocation of the edge server (MESs) is optimized. In addition, the invention also designs a SA-MADDPG algorithm that combines state normalization and adaptive decay learning rate to improve the algorithm's adaptability to environmental changes, and verifies its effectiveness in reducing latency energy consumption and improving endurance through simulation experiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0147] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0148] Figure 1 It is an execution flow chart of a mobile edge computing dynamic unloading method for a CNN network provided by an embodiment of the present invention.
[0149] Figure 2 It is a framework diagram of the SA-MADDPG algorithm provided by an embodiment of the present invention.
[0150] Figure 3 It is a schematic diagram of an electronic device provided by an embodiment of the present invention.
[0151] Figure 4 It is a schematic diagram of a computer-readable storage medium provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0152] The present invention will be further described in detail below in conjunction with the accompanying drawings and examples. It is particularly noted that the following examples are only used to illustrate the present invention, but are not intended to limit the scope of the present invention. Similarly, the following examples are only partial embodiments of the present invention rather than all embodiments, and all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present invention.
[0153] This paper proposes an improved partial offloading strategy, (1) by setting a single split point to reduce the data transmission of the CNN network, and combining the network node reservation algorithm (NPPR) to intelligently select the task division point to simplify the decision-making process. (2) In the research on CNN offloading of MEC, the focus is usually on the trade-off between energy consumption and latency or detection accuracy. However, considering the dynamic nature of the environment, the lack of infrastructure, the heterogeneity of device computing resources and resource competition among mobile devices, and the limited battery life of the device, this paper proposes a multi-agent reinforcement learning algorithm (MADRL) and incorporates the battery life of the device into the optimization target. That is, under the condition of limited resources, this paper not only pursues the optimal trade-off between latency and energy consumption between multiple devices, but also considers the remaining power of the device, optimizes the offloading decision by dynamically adjusting the system weight, and prolongs the use time of the device. This strategy aims to improve system performance and reliability, and provides a new optimization method for the application of remote MEC.
[0154] Specifically, in the IoT environment, the present invention targets remote area perception and computing scenarios consisting of multiple mobile device MMDs and multiple edge servers MESs, and perception devices. A joint optimization strategy based on MADRL is proposed to solve the problem of task offloading and resource allocation based on CNN. The strategy regards each mobile device MMD as an independent agent, vertically divides CNN, and optimizes the resource allocation of MESs. In order to enhance the environmental adaptability of the device, the agent's remaining power is introduced as the weight factor of the optimization target, and a non-deterministic polynomial-hard (NP-hard) mixed integer nonlinear programming (MINLP) problem of discrete variables and continuous variables is constructed, and the NPPR algorithm is used to simplify the solution space of the problem. Furthermore, the problem is transformed into a multi-agent Markov decision process (MAMDP), and a combined state normalization and adaptive decay learning rate (SA-MADDPG) algorithm is designed to improve the adaptability of the algorithm to environmental changes.
[0155] See also Figure 1 and Figure 2 , a mobile edge computing dynamic unloading method for a CNN network of the present invention comprises the following steps:
[0156] Step 1: Data preparation: According to the target detection requirements, image task data, device status data and network status data are obtained in the integrated system of synaesthesia and computing to form a data set;
[0157] In this embodiment, the step 1 specifically includes:
[0158] Step 11. According to the task requirements submitted by the user, the integrated synaesthesia and computing system is composed of multiple mobile devices and perception devices; the perception devices in the integrated synaesthesia and computing system perform data perception operations, collect and store image task data; the integrated synaesthesia and computing system simultaneously records the device status data and network status data of the mobile device; the mobile device has the functions of perception, communication, storage and calculation, but the mobile device is not required to have some or all of the above-mentioned perception, communication, storage and calculation functions. The integrated synaesthesia and computing system is composed of mobile devices and perception devices connected by a wired and wireless or only wireless network. The integrated synaesthesia and computing system refers to a network that has communication, perception and calculation capabilities at the same time.
[0159] Step 12: Divide the acquired image task data, device status data and network status data based on time slots to form a time series data set within a set time period; the time slot refers to a time period with a set time length, and the data acquired within the time period are all marked as data of the time slot, and the time series data refers to data with time slot attributes and numbered in the order of time slots;
[0160] Step 13: Divide the time series data set into a training set, a validation set, and a test set.
[0161] Step 2: Model establishment: Based on the synaesthesia-computing integrated system and data set, a synaesthesia-computing integrated system model is constructed. Aiming at the computing task, the synaesthesia-computing integrated system model is transformed into the computing task segmentation and resource allocation problem;
[0162] In this embodiment, step 2 specifically includes:
[0163] Step 21, the nodes with sensing, computing, communication and storage functions in the integrated system of synergy and computing are connected to each other through wireless and wired communication or wireless communication only to form an integrated system model of synergy and computing; the integrated system model of synergy and computing works according to the time slots divided in step 12;
[0164] Step 22: In the synaesthesia-computing integrated system model based on time slot division, the user initiates a task request in a certain time slot; specifically:
[0165] In the synaesthesia-computing integrated system model based on time slot division, the collection of mobile device MMDs , where device m represents a mobile device MMD selected from multiple mobile devices MMDs to implement task perception, and all are loaded with a computing module based on a CNN network. The CNN network contains multiple CNN models, and different CNN models correspond to different computing tasks; a collection of edge servers MESs (as fixed devices) , wherein device s represents selecting an edge server MES from multiple edge servers MESs to provide computing resources for surrounding mobile devices MMDs; the devices form a telepathic computing network through wired and wireless or only wireless means, and the user terminal accesses the telepathic computing network through wired or wireless means, and initiates a task request to the telepathic computing network in a certain time slot; the request includes the type of data to be sensed and the required data processing requirements; the requirements include but are not limited to the sensing range, sensing accuracy and sensing timeliness, etc.
[0166] Step 23: After receiving the task request, the integrated system model of synaesthesia and calculation implements the perception and calculation process; the calculation process consists of data transmission and data processing, and the perception and calculation process is an integrated process of synaesthesia and calculation; specifically:
[0167] The mobile device MMDs are responsible for the implementation of the computation offloading decision. First, the mobile device MMDs selects a suitable CNN model to handle the detection task. Due to the limited computing power and power of the MMDs themselves, the mobile device MMDs select one of the devices Establish communication and offload part of the workload to MES; the entire integrated system model of synergy and computing runs in a time slot manner. As the entire running time of the synaesthesia computing integrated system, the entire running time is divided into time slots of equal length, the time slot duration is expressed as , Indicates rounding down, MMDs in time slot At the beginning, a target detection task is performed; the specific steps are as follows: in each time slot Initially, the device Different target detection tasks are generated according to the requirements submitted by users. The CNN models required for different types of tasks are heterogeneous. Need to choose the right equipment Establish connections and identify task offloading parts, each Receive multiple requests and allocate computing resources to them.
[0168] It should be noted that for CNN-based computing networks, since each layer of the network structure is in series, that is, the execution of each layer of the structure needs to be activated by the execution result of the previous layer. MMDs determines a split point for each network, and divides the CNN network into two interdependent sub-parts through the split point. The sub-parts are executed sequentially on MMDs and MES.
[0169] In this embodiment, step 23 specifically includes:
[0170] Step 231: partition and unload tasks according to the CNN network:
[0171] Assuming that the synaesthesia computing integrated system has a total of For the CNN model CNN models, among which, ,use Indicates The total number of layers in the CNN model, Each layer of the CNN model The properties are ,in, , Respectively Enter the height, width and number of channels, Represent the channel value, kernel size and stride size respectively; the CNN network is mainly composed of convolutional layers and pooling layer In the reasoning process of the network, the output of the previous layer is used as the input of the next layer, resulting in dependencies between the layers, and each layer will not interfere with each other, providing a theoretical basis for subsequent network structure partitioning and unloading. Before calculating the execution delay and energy consumption of the CNN network, it is necessary to calculate the amount of calculation and output parameters of each layer of the network. The calculation of each layer is regarded as Floating Point Operations (FLOPs), which is used to measure the amount of calculation of the number of YOLO network layers.
[0172] Since the network's computation is mainly concentrated on and In order to facilitate representation, the convolutional layer and the pooling layer are unified as express, The total computational effort of the layer is recorded as , the calculation formula is:
[0173]
[0174] In the reasoning process of the CNN network, the output of the previous layer is used as the input of the next layer. The output of the layer is equal to the output of the next layer The amount of input, so the amount of output data is recorded as , the calculation formula is:
[0175]
[0176] in, is the memory usage of unit data;
[0177] For the partition offloading method, there are multiple feasible segmentation points for a CNN network; considering that there is only one optimal segmentation point for any CNN network on each MMD, assuming that the segmentation point is , Indicates partition point, the first layer to the second layer are executed locally. layer CNN network, and then execute The intermediate data after the layer is unloaded to the MES, and the MES executes layer until the last layer, and finally transmit the inference result back to MMD; When , all tasks are performed locally; when When the task is directly offloaded to MES;
[0178] In order to better optimize the partitioning and unloading of the network, it is necessary to determine the location of the segmentation points in each network and the computing resources allocated to the application by the MMD, and select the most suitable MES for the computing tasks on the MMD. By changing the partitioning points and unloading objects of the CNN network, the overall system's adaptability to the environment can be improved.
[0179] Step 232: Obtain the computation delay and transmission delay of the CNN network:
[0180] CNN inference latency is mainly composed of computational latency and transmission latency. The computational latency is caused by the execution of CNN on MMDs and MESs, which is related to the model size and device computing resources. The transmission latency is affected by the network bandwidth and the amount of data transmitted.
[0181] (1) Execution time on MMD: Indicates the device The computing resources in the time slot Assigned to device The computing resources for local execution of applications are , ; Local execution time for:
[0182]
[0183] in, Indicates time slot Medium Equipment Above, first layer and partition layer The sum of the computational effort of all layers between Indicates CNN model The number of partition point layers; Indicates time slot Medium Equipment Previous The amount of computation of the layer;
[0184] (2) Transmission time of intermediate output data:
[0185] equipment To device The channel gain of the communication link between Modeled as:
[0186]
[0187] in, Represents the Euclidean distance between corresponding devices, Indicates the channel gain when the reference distance is 1m; the corresponding transmission rate Defined as:
[0188]
[0189] in, is the transmission bandwidth of the edge computing network, is the transmission power of the UAV, is the waveguide channel noise; let Indicated in Device during time slot At the partition level The size of the data transferred is obtained by the device Offload to device Transmission delay for:
[0190]
[0191] Since the size of the result data is much smaller than the size of the input data, the delay in transmitting the detected results back to the local device is ignored.
[0192] Since MES needs to offload calculations from multiple MMDs, and the CNN model in MES is more complex than the CNN model in MMD, it is necessary to consider the calculation delay of the target detection task in MES and the queuing delay between multiple tasks; Indicates time slot Internal equipment Assign to device of computing resources, of which, , 'For equipment A collection of MMD devices that provide computing services at the same time. For equipment The total computing resources of Time slot, the execution delay of device m on device s for:
[0193]
[0194] in, Indicates transmission to the device The CNN network partition on The local computation amount of the part after the layer; Indicates time slot Medium Equipment Previous The amount of computation of the layer;
[0195] Therefore, after dividing an object detection task, the device Total execution time It is equal to the sum of the local execution time on MMD, the intermediate data transmission time and the execution time on MES:
[0196]
[0197] Assume that each Each task must be completed within a single time slot, that is
[0198] Step 233, calculate energy consumption:
[0199] (1) Execution energy consumption on MMD: Assume Indicates the device The computing power of the device Local execution energy consumption for:
[0200]
[0201] (2) Equipment Energy consumption of intermediate output data transmission for:
[0202]
[0203] in, express The transmission power of
[0204] (3) Equipment The energy consumed in executing the task is:
[0205]
[0206] in, Indicates the device Average power when performing tasks;
[0207] (4) When the device Total energy consumption of executing some CNN network tasks It is the sum of local execution energy consumption, intermediate output data transmission energy consumption and task execution energy consumption:
[0208]
[0209] Step 234, calculate the quality of service:
[0210] (1) By device and equipment Quality of service of the computing process Affected by both task delay and energy consumption; and They represent the improvement levels of energy consumption and latency of MMDs respectively; Describes the energy saved by calculating the ratio of the local energy consumption to the energy consumption saved by offloading MMDs. It represents the reduction ratio of latency compared with local latency through MESs computing offloading; therefore, the energy consumption and improvement level of MMD are expressed as:
[0211]
[0212] and
[0213]
[0214] in, and Respectively represent the time slot Period equipment Latency and energy consumption when performing tasks entirely locally;
[0215] (2) For drones with different remaining battery capacities, the offloading strategy has different preferences in terms of energy saving and latency reduction. Therefore, the weight factor To measure the trade-off between energy saving and latency reduction in MMD, consider the remaining battery power of MMD and compare the remaining power to Further weighting factors are incorporated, where:
[0216]
[0217] in, and Respectively represent the current maximum battery capacity and current remaining power of the device;
[0218] (3) The weighted sum of energy consumption improvement and delay improvement is expressed as:
[0219]
[0220] By adjusting the weight factors, MMDs can save more energy or reduce latency. Low latency requires faster computation time and transmission speed, which will aggravate the energy consumption of MMDs. On the contrary, lower residual power limits the level of latency improvement, even though MMDs have a strong preference for latency improvement.
[0221] (4) Using exponential functions to characterize the trade-off model between energy consumption improvement and delay improvement and MMDs service quality The mapping relationship between them is defined as:
[0222]
[0223] in, Represents variable parameters, is a positive integer, then express Index of
[0224] (5) Based on Derived , so it is deduced that The upper and lower bounds of the value are and ;
[0225] Step 235: Define the target problem:
[0226] The scheduling strategy mainly includes real-time dynamic adjustment of the weight ratio of delay and energy consumption according to the remaining battery power of the MMD, and minimizing system delay and energy consumption; each MMD has a corresponding Indicator, each MMD through decision , and To implement computing tasks and optimize all MMDs in the system ,in, The decision of device m to select device s is as follows:
[0227]
[0228]
[0229]
[0230]
[0231]
[0232] Combining the equations, it is inferred that the target problem is a mixed integer nonlinear programming (MINLP) problem, which is also a non-deterministic polynomial-hard (NP-Hard) problem;
[0233] constraint Indicates the device Assign to device The computing resources must not exceed the maximum available local computing resources; Indicates that the partition point selection of the target detection model must be within the range of the number of layers of the target detection model; Indicates that the selection of MESs by MMD must be from the set Select Indicates that the remaining battery power of the MMD does not exceed its maximum capacity.
[0234] Step 3: Model transformation: Based on the integrated system model of synaesthesia and computing and the problems of computing task segmentation and resource allocation, a Markov decision model is constructed;
[0235] In this embodiment, step 3 specifically includes:
[0236] Step 31: Each device It is regarded as an intelligent agent, and the problem of computing task partitioning and resource allocation is further expressed as a multi-agent Markov decision process (MAMDP), and a Markov decision model is constructed;
[0237] Step 32: Adaptively determine the offloading decision of each device according to the environmental state to maximize the user service quality; the Markov decision model is expressed as ,in, is the state space, is the action space, is the state transition probability matrix, is the reward; the state, action, and reward of the Markov decision model are as follows:
[0238] State: The state space consists of the states of MMD and MES, which evolve at different times; time slots The state space at time Defined as :
[0239]
[0240] in, For time slot Time equipment The CNN model category carried on, Indicates time slot Assigned to device The computing resources for local execution of applications, express Time slot equipment The remaining power, Indicated in Devices in time slot Assign to device The computing resources are sent to each MMD in the form of broadcast; Indicates the device With equipment Uplink capacity between
[0241] Action: A device affects the environment through its actions; all devices The action set is a mixed integer and non-integer offloading decision vector, where the device In time slot The action is expressed as:
[0242]
[0243] in, Indicates the offloading decision of the MMD offloading task, Represents the partition of the CNN network carried on MMD;
[0244] Reward: In a Markov decision process (MDP), the reward is the Status The system gain obtained by taking actions under the given conditions; since reinforcement learning models learn by maximizing cumulative rewards, an appropriate reward function is crucial to optimizing reinforcement learning performance. Complete the device When the task is transferred, the device Generating Quality of Service Feedback, MESs collaborate to achieve service quality for all devices Maximization is the goal; therefore, the reward function is expressed as:
[0245] .
[0246] Step 4: Model training: Based on the Markov decision model and the multi-agent deep reinforcement learning algorithm (MADRL), a joint optimization strategy is obtained, and the neural network parameter model that meets the termination conditions is obtained through training through the data set;
[0247] In this embodiment, step 4 specifically includes:
[0248] Step 41: Determine the network segmentation nodes based on the network node reservation algorithm (NPPR):
[0249] (1) Due to the complexity of the CNN network, there may be multiple potential partition points, which leads to a complex solution space in the multi-agent Markov decision process (MAMDP). Before determining the approximate optimal offloading and resource allocation strategy, a network node reservation algorithm (NPPR) is introduced. By retaining an appropriate number of network nodes, the complexity of the subsequent algorithm is reduced.
[0250] Since some network layers have large output data volumes, which will lead to high communication costs under low bandwidth conditions, they are not suitable as potential partitioning points. First, a CNN model j is selected as the target of preprocessing and the different bandwidths are calculated. The execution time corresponding to all partition points , ≤ ;definition To preserve the set of partition layers; define For bandwidth The corresponding partition layer set; definition is the candidate point set corresponding to the current CNN model j;
[0251] Calculate the same bandwidth The average execution time of all partition points , as shown below:
[0252]
[0253] (2) Order is a weight factor in the range [0,1], and each node is associated with the average execution time and weight factor The product of is compared; if Less than , then keep the node until ; Otherwise, eliminate the node; finally get the bandwidth Corresponding ;
[0254] (3) Through the formula:
[0255]
[0256] renew , when taking different When the value is reached, repeat (1) to (3) in step 41 until all bandwidths are traversed. , and obtain the final ;
[0257] (4) Candidate point sets for all types of CNN models Integrate into middle;
[0258] Step 42: Joint optimization strategy based on multi-agent deep reinforcement learning algorithm (MADRL):
[0259] (1) The MADDPG framework is integrated in a multi-agent collaborative environment. Due to the cooperation and competition between agents in the environment, isolated reinforcement learning algorithms are not feasible for a single agent. Therefore, a collaborative multi-agent framework is deployed to ensure a reasonable scheduling strategy among all agents. The DDPG algorithm carried by each agent integrates the advantages of policy gradient and deep Q network (DQN).
[0260] Each agent consists of an actor network and a comment Home Network Both are parameterized by deep neural networks; the actor network takes state observations as input and Output action; at the same time, the critic network uses the state-action function Evaluate the actions chosen by the actor network based on the current state;
[0261] (2) In the MADDPG framework, similar to the algorithm used for single agents, the MADDPG framework combines centralized training with decentralized execution. Each target network takes the global states of multiple agents as input to evaluate the quality of actions under these states; the target network promotes stable learning by generating temporarily delayed target policies and value estimates; in the MADDPG framework, the actor network observes the state of its local agent based on the target network’s local state. Determine local actions , the policy parameters of each agent are defined as ,in, Indicates The strategy parameters of each agent are Parameterized Continuous Strategy , is the actor network parameter of agent m; in the deterministic strategy framework, each agent adopts a continuous strategy ,in It is a random noise used to enhance the exploration capability of the system. In the MADDPG framework, the actual actions taken by the current agent can be perceived by other agents in the system after a certain amount of communication.
[0262] (3) The critic network transforms the state of the agent and actions As input, it is represented as ,in, is the network parameter of the critic network, which is used to evaluate the collective performance Perform global evaluation, whose input includes the actions and states of all agents; similar to DQN, MADDPG enhances system stability through the target network and experience replay buffer. The target actor network of the agent is defined as , the corresponding network parameters are defined as , the target network is expressed as , the corresponding network parameters are ; The update formulas for these parameters are defined as:
[0263]
[0264] Among them, the parameters represents a coefficient in the range of 0 to 1;
[0265] (4) The MADDPG algorithm consists of two stages: centralized offline training and decentralized execution. In the centralized offline training stage, each agent uses not only its own observations but also the observations and actions of other agents for strategy evaluation. Each agent can select an action based on its local observations and use the critic network to evaluate the selected action. By adopting the state-action value function Simplifying training, the function can predict the long-term expected reward based on the current observation-action pair. During the decentralized execution phase, each agent can only access its own state and decide its actions independently without knowing the situation of other agents.
[0266] During the centralized offline training process, the system collects the state information of the agent ,in Indicates the status is Execute action when The next state entered after that, this state information is stored in the experience playback buffer; is the reward value obtained by the current action; when collecting data into the experience replay buffer After that, random sampling is used in the policy evaluation process; the MADDPG algorithm follows the offline policy method and uses the sampled data to update and Parameters; the objective function update is expressed as:
[0267]
[0268] in, represents the discount factor, It is through the formula: We get the critic network updated by minimizing the loss function:
[0269]
[0270] in, [ ] means to seek expectation;
[0271] After updating the critic network, the actor network is parameterized by the following formula Update:
[0272]
[0273] in, Representation parameters The gradient of
[0274] The critic network then updates its parameters according to the gradient rule of the neural network, obtaining:
[0275]
[0276] in is the learning rate of the critic network, and then the actor network is updated using the policy gradient method to obtain:
[0277]
[0278] The actor network parameters are updated as:
[0279]
[0280] in, is the learning rate of the critic network; represents the actor network; Represents the actor network Ask about parameters The gradient of
[0281] Step 43. Standardized and adaptive learning rate for multi-agent deep reinforcement learning algorithm:
[0282] (1) During the initialization process of the multi-agent deep reinforcement learning algorithm, the environment state is standardized by storing the maximum and minimum values during training in temporary memory and calculating the scaling factor, which solves the magnitude difference problem of the input variables and improves the training efficiency through standardized calculation. The difference between the maximum and minimum values of each environment variable is used as the scaling factor; the standardized state information is obtained by subtracting the minimum value from the current state and dividing it by the obtained scaling factor;
[0283] (2) During the training process of reinforcement learning, due to the uncertainty of the environment, the sparsity of rewards and the complexity of the model itself, a fixed learning rate is often not enough to meet the needs of the entire training process. Therefore, a decaying learning rate strategy is used to adjust the learning rate according to the loss rate at different training stages:
[0284]
[0285] in, is the learning rate decay coefficient. When the monitoring indicator no longer improves, the system's learning rate can be reduced.
[0286] Step 5: Model deployment: deploy the trained neural network parameter model to the synaesthesia computing integrated system;
[0287] In this embodiment, step 5 specifically includes:
[0288] Step 51: synchronize the trained neural network parameters to each node in the CNN network at a set time interval according to the environmental dynamic information; the time interval depends on whether the new neural network parameter training process converges to a new state;
[0289] Step 52: Each node in the CNN network updates the unloading strategy and resource allocation strategy based on the obtained neural network parameters;
[0290] Step 53: Each node in the CNN network feeds back the updated strategy status to the integrated synaesthesia and computing system.
[0291] Step 6, decision implementation: Based on the real-time image task data, equipment status data and network status data, generate computing task segmentation decisions and computing resource allocation decisions, and transmit the decision result information to each node of the integrated synergy computing system, ultimately realizing the offloading of computing tasks.
[0292] In this embodiment, step 6 specifically includes:
[0293] Step 61: The sensory-computing integrated system obtains real-time network status data and device status data, and the sensing device obtains real-time image task data according to user needs;
[0294] Step 62: Based on the input network status data, device status data and image task data, according to the multi-agent Markov decision model determined by the neural network, generate a decision on the division of computing tasks and the allocation of computing resources;
[0295] Step 63, synchronizing the decision result information to each node in the integrated system of synergy and computing;
[0296] Step 64: Each node implements and executes a specific computing task.
[0297] like Figure 3 As shown, an embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned dynamic unloading method of mobile edge computing for CNN network is implemented.
[0298] like Figure 4 As shown, an embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned dynamic unloading method of mobile edge computing for CNN network is implemented.
[0299] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0300] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0301] The above descriptions are only some embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any equivalent device or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A dynamic offloading method for mobile edge computing of CNN network, characterized in that: The steps include: Step 1: According to the target detection requirements, image task data, device status data and network status data are obtained in the integrated system of synergy and computing to form a data set; Step 2: Based on the synaesthesia-computing integrated system and data set, a synaesthesia-computing integrated system model is constructed. Aiming at the computing task, the synaesthesia-computing integrated system model is transformed into the computing task segmentation and resource allocation problem; Step 3: Based on the integrated system model of synaesthesia and computing and the problem of computing task segmentation and resource allocation, a Markov decision model is constructed; Step 4: Based on the Markov decision model and the multi-agent deep reinforcement learning algorithm, a joint optimization strategy is obtained, and the neural network parameter model that meets the termination conditions is obtained through training through the data set; Step 5: Deploy the trained neural network parameter model to the synaesthesia computing integrated system; specifically including: Step 51: synchronize the trained neural network parameters to each node in the CNN network at a set time interval according to the environmental dynamic information; the time interval depends on whether the new neural network parameter training process converges to a new state; Step 52: Each node in the CNN network updates the unloading strategy and resource allocation strategy based on the obtained neural network parameters; Step 53, each node in the CNN network feeds back the updated strategy status to the integrated system of synaesthesia and computing; Step 6: Generate computing task segmentation decisions and computing resource allocation decisions based on real-time image task data, device status data, and network status data, and transmit the decision result information to each node of the integrated synergy computing system, ultimately realizing the offloading of computing tasks.
2. The mobile edge computing dynamic offloading method for CNN network according to claim 1, characterized in that: The step 1 specifically includes: Step 11: According to the task requirements submitted by the user, the integrated system of synergy and computing is composed of multiple mobile devices and sensing devices; the sensing devices in the integrated system of synergy and computing perform data sensing operations, collect and store image task data; the integrated system of synergy and computing simultaneously records the device status data and network status data of the mobile device; Step 12: Divide the acquired image task data, device status data and network status data based on time slots to form a time series data set within a set time period; the time slot refers to a time period with a set time length, and the data acquired within the time period are all marked as data of the time slot, and the time series data refers to data with time slot attributes and numbered in the order of time slots; Step 13: Divide the time series data set into a training set, a validation set, and a test set.
3. The method for dynamic offloading of mobile edge computing for CNN network as claimed in claim 2, characterized in that: The step 2 specifically includes: Step 21, the nodes with sensing, computing, communication and storage functions in the integrated system of synergy and computing are connected to each other through wireless and wired communication or wireless communication only to form an integrated system model of synergy and computing; the integrated system model of synergy and computing works according to the time slots divided in step 12; Step 22: In the synaesthesia-computing integrated system model based on time slot division, the user initiates a task request in a certain time slot; specifically: In the synaesthesia-computing integrated system model based on time slot division, the collection of mobile device MMDs , where device m represents a mobile device MMD selected from multiple mobile devices MMDs to implement task perception, and all of them are loaded with a computing module based on a CNN network. The CNN network contains multiple CNN models, and different CNN models correspond to different computing tasks; the collection of edge servers MESs , wherein device s represents selecting an edge server MES from multiple edge servers MESs to provide computing resources for surrounding mobile devices MMDs; the devices form a network of inter-sensing and computing integration through wired and wireless or only wireless means, the user terminal accesses the network of inter-sensing and computing integration through wired or wireless means, and initiates a task request to the network of inter-sensing and computing integration in a certain time slot; the request includes the type of data to be sensed and the required data processing requirements; Step 23: After receiving the task request, the integrated system model of synaesthesia and calculation implements the perception and calculation process; the calculation process consists of data transmission and data processing, and the perception and calculation process is an integrated process of synaesthesia and calculation; specifically: The mobile device MMDs are responsible for computing the implementation of the offload decision. First, the mobile device MMDs selects one of the devices Establish communication and offload part of the workload to MES; the entire integrated system model of synergy and computing runs in a time slot manner. As the entire running time of the synaesthesia computing integrated system, the entire running time is divided into time slots of equal length, the time slot duration is expressed as , Indicates rounding down, MMDs in time slot At the beginning, a target detection task is performed; the specific steps are as follows: in each time slot Initially, the device Different target detection tasks are generated according to the requirements submitted by users. The CNN models required for different types of tasks are heterogeneous. Need to choose the right equipment Establish connections and identify task offloading parts, each Receive multiple requests and allocate computing resources to them.
4. The method for dynamic offloading of mobile edge computing for CNN network as claimed in claim 3, characterized in that: The step 23 specifically includes: Step 231: partition and unload tasks according to the CNN network: Assuming that the synaesthesia computing integrated system has a total of For the CNN model CNN models, among which, ,use Indicates The total number of layers in the CNN model, Each layer of the CNN model The properties are ,in, , Respectively Enter the height, width and number of channels, Represent the channel value, kernel size and stride size respectively; the CNN network is mainly composed of convolutional layers and pooling layer The convolutional layer and the pooling layer are unified with express, The total computational effort of the layer is recorded as , the calculation formula is: In the reasoning process of the CNN network, the output of the previous layer is used as the input of the next layer. The output of the layer is equal to the output of the next layer The amount of input, so the amount of output data is recorded as , the calculation formula is: in, is the memory usage of unit data; Consider that there is only one optimal segmentation point for any CNN network on each MMD, assuming that the segmentation point is , Indicates partition point, the first layer to the second layer are executed locally. layer CNN network, and then execute The intermediate data after the layer is unloaded to the MES, and the MES executes layer until the last layer, and finally transmit the inference result back to MMD; When , all tasks are performed locally; when When the task is directly offloaded to MES; Step 232: Obtain the computation delay and transmission delay of the CNN network: (1) Execution time on MMD: Indicates the device The computing resources in the time slot Assigned to device The computing resources for local execution of applications are , ; Local execution time for: in, Indicates time slot Medium Equipment Above, first layer and partition layer The sum of the computational effort of all layers between Indicates CNN model The number of partition point layers; Indicates time slot Medium Equipment Previous The amount of computation of the layer; (2) Transmission time of intermediate output data: equipment To device The channel gain of the communication link between Modeled as: in, Represents the Euclidean distance between corresponding devices, Indicates the channel gain when the reference distance is 1m; the corresponding transmission rate Defined as: in, is the transmission bandwidth of the edge computing network, is the transmission power of the UAV, is the waveguide channel noise; let Indicates that during time slot t, device m is in the partition layer The size of the data transferred, get the device Offload to device Transmission delay for: Considering the computational delay of the target detection task in MES and the queuing delay between multiple tasks; let Indicates time slot Internal equipment Assign to device of computing resources, of which, , 'For equipment A collection of MMD devices that provide computing services at the same time. For equipment The total computing resources of Time slot, the execution delay of device m on device s for: in, Indicates transmission to the device The CNN network partition on The local computation amount of the part after the layer; Indicates time slot Medium Equipment Previous The amount of computation of the layer; Therefore, after dividing an object detection task, the device Total execution time It is equal to the sum of the local execution time on MMD, the intermediate data transmission time and the execution time on MES: Assume that each Each task must be completed within a single time slot, that is ; Step 233, calculate energy consumption: (1) Execution energy consumption on MMD: Assume Indicates the device The computing power of the device Local execution energy consumption for: (2) Equipment Energy consumption of intermediate output data transmission for: in, express The transmission power of (3) Equipment The energy consumed in executing the task is: in, Indicates the device Average power when performing tasks; (4) When the device Total energy consumption of executing some CNN network tasks It is the sum of local execution energy consumption, intermediate output data transmission energy consumption and task execution energy consumption: Step 234, calculate the quality of service: (1) Use and They represent the improvement levels of energy consumption and latency of MMDs respectively; Describes the energy saved by calculating the ratio of the local energy consumption to the energy consumption saved by offloading MMDs. It represents the reduction ratio of latency compared with local latency through MESs computing offloading; therefore, the energy consumption and improvement level of MMD are expressed as: and in, and Respectively represent the time slot Period equipment Latency and energy consumption when performing tasks entirely locally; (2) Introducing weight factors To measure the trade-off between energy saving and latency reduction in MMD, consider the remaining battery power of MMD and compare the remaining power to Further weighting factors are incorporated, where: in, and Respectively represent the current remaining power and current maximum battery capacity of the device; (3) The weighted sum of energy consumption improvement and delay improvement is expressed as: (4) Using exponential functions to characterize the trade-off model between energy consumption improvement and delay improvement and MMDs service quality The mapping relationship between them is defined as: in, Represents variable parameters, is a positive integer, then express Index of (5) Based on Derived , so it is deduced that The upper and lower bounds of the value are and ; Step 235: Define the target problem: The scheduling strategy mainly includes real-time dynamic adjustment of the weight ratio of delay and energy consumption according to the remaining battery power of the MMD, and minimizing system delay and energy consumption; each MMD makes a decision , and To implement computing tasks and optimize all MMDs in the system ,in, The decision of device m to select device s is as follows: Combining the equations, it can be inferred that the target problem is a mixed integer nonlinear programming problem, which is also a non-deterministic polynomial-hard problem; constraint Indicates the device Assign to device The computing resources must not exceed the maximum available local computing resources; Indicates that the partition point selection of the target detection model must be within the range of the number of layers of the target detection model; Indicates that the selection of MESs by MMD must be from the set Select from Indicates that the remaining battery power of the MMD does not exceed its maximum capacity.
5. The method for dynamic offloading of mobile edge computing for CNN network as claimed in claim 4, characterized in that: The step 3 specifically includes: Step 31: Each device m is regarded as an agent, and the computing task segmentation and resource allocation problem is further represented as a multi-agent Markov decision process, and a Markov decision model is constructed; Step 32: Adaptively determine the offloading decision of each device according to the environmental state to maximize the user service quality; the Markov decision model is expressed as ,in, is the state space, is the action space, is the state transition probability matrix, is the reward; the state, action, and reward of the Markov decision model are as follows: State: The state space consists of the states of MMD and MES, time slots The state space at time Defined as : in, For time slot Time equipment The CNN model category carried on, Indicates time slot Assigned to device The computing resources for local execution of applications, express Time slot equipment The remaining power, Indicates time slot Internal equipment Assign to device The computing resources are sent to each MMD in the form of broadcast; Indicates the device With equipment Uplink capacity between Action: All devices The action set is a mixed integer and non-integer offloading decision vector, where the device In time slot The action is expressed as: in, Indicates the offloading decision of the MMD offloading task, Represents the partition of the CNN network carried on MMD; Reward: In the Markov decision process, the reward is the Status The system benefits obtained by taking actions under Complete the device When the task is transferred, the device Generating Quality of Service Feedback, MESs collaborate to achieve service quality for all devices Maximization is the goal; therefore, the reward function is expressed as: 。 6. The method for dynamic offloading of mobile edge computing for CNN network as claimed in claim 5, characterized in that: The step 4 specifically includes: Step 41: Determine the network segmentation nodes based on the network node retention algorithm: (1) Select a CNN model j as the preprocessing target and calculate different bandwidths The execution time corresponding to all partition points , ≤ ;definition To preserve the set of partition layers; define For bandwidth The corresponding partition layer set; definition is the candidate point set corresponding to the current CNN model j; Calculate the same bandwidth The average execution time of all partition points , as shown below: (2) Order is a weight factor in the range [0,1], and each node is associated with the average execution time and weight factor The product of is compared; if Less than , then keep the node until ; Otherwise, eliminate the node; finally get the bandwidth Corresponding ; (3) Through the formula: renew , when taking different When the value is reached, repeat (1) to (3) in step 41 until all bandwidths are traversed. , and obtain the final ; (4) Candidate point sets for all types of CNN models Integrate into middle; Step 42: Joint optimization strategy based on multi-agent deep reinforcement learning algorithm: (1) The MADDPG framework is integrated in a multi-agent collaborative environment, where each agent consists of an actor network and a network of critics Both are parameterized by deep neural networks; the actor network takes state observations as input and Output action; at the same time, the critic network uses the state-action function Evaluate the actions chosen by the actor network based on the current state; (2) In the MADDPG framework, each target network takes the global states of multiple agents as input to evaluate the quality of actions under these states; the target network promotes stable learning by generating temporarily delayed target policies and value estimates; in the MADDPG framework, the actor network observes the state of its local agent based on the target network’s action quality. Determine local actions , the policy parameters of each agent are defined as ,in, Indicates The strategy parameters of each agent are Parameterized Continuous Strategy , is the actor network parameter of agent m; in the deterministic strategy framework, each agent adopts a continuous strategy ,in is the random noise used to enhance the system’s exploration capabilities; (3) The critic network transforms the state of the agent and actions As input, it is represented as ,in, is the network parameter of the critic network, which is used to evaluate the collective performance A global evaluation is performed, whose input includes the actions and states of all agents; the target actor network of the agents is defined as , the corresponding network parameters are defined as , the target network is expressed as , the corresponding network parameters are ; The update formulas for these parameters are defined as: Among them, the parameters represents a coefficient in the range of 0 to 1; (4) The MADDPG algorithm consists of two stages: centralized offline training and decentralized execution; During the centralized offline training process, the system collects the state information of the agent ,in Indicates the status is Execute action when The next state entered after that, this state information is stored in the experience playback buffer; is the reward value obtained by the current action; after collecting data into the experience replay buffer, random sampling is used for the policy evaluation process; the MADDPG algorithm follows the offline policy method and uses the sampled data to update and Parameters; the objective function update is expressed as: in, represents the discount factor, It is through the formula: Get, update the critic network by minimizing the loss function: in, [ ] means to seek expectation; After updating the critic network, the actor network is parameterized by the following formula Update: in, Representation parameters The gradient of The critic network then updates its parameters according to the gradient rule of the neural network, obtaining: in is the learning rate of the critic network, and then the actor network is updated using the policy gradient method to obtain: The actor network parameters are updated as: in, is the learning rate of the critic network; represents the actor network; Represents the actor network Ask about parameters The gradient of Step 43. Standardized and adaptive learning rate for multi-agent deep reinforcement learning algorithm: (1) During the initialization of the multi-agent deep reinforcement learning algorithm, the environment state is standardized by storing the maximum and minimum values during training in temporary memory and calculating a scaling factor. The difference between the maximum and minimum values of each environment variable is used as the scaling factor. The standardized state information is obtained by subtracting the minimum value from the current state and dividing it by the obtained scaling factor. (2) During the reinforcement learning training process, the decay learning rate strategy is used to adjust the learning rate according to the loss rate at different training stages: in, is the learning rate decay coefficient.
7. The method for dynamic offloading of mobile edge computing for CNN network according to claim 1, characterized in that: The step 6 specifically includes: Step 61, the integrated sensing and computing system obtains real-time network status data and device status data, and the sensing device obtains real-time image task data according to user needs; Step 62: Based on the input network status data, device status data and image task data, according to the multi-agent Markov decision model determined by the neural network, generate a decision on the division of computing tasks and the allocation of computing resources; Step 63, synchronizing the decision result information to each node in the integrated system of synergy and computing; Step 64: Each node implements and executes a specific computing task.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the dynamic unloading method of mobile edge computing for CNN network is implemented as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the dynamic unloading method of mobile edge computing for a CNN network is implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Network state self-adaptive target detection computing task unloading scheduling method
CN118484315A