A resource-aware intelligent edge gateway system and method of operation
By using a resource-aware intelligent edge gateway system and a multi-agent hybrid decision-making algorithm (MCMHD), the resource allocation of the wireless charging mobile edge computing system is optimized, solving the problems of uneven resource allocation and insufficient energy utilization, and achieving low latency and low power consumption for computing tasks.
Patent Information
- Application Number
- CN202411664799.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-11-20
AI Technical Summary
Wireless charging mobile edge computing systems suffer from uneven resource allocation and insufficient energy utilization, resulting in excessive latency and power consumption of computing tasks. In particular, when data volume surges, some devices are unable to compute or offload tasks due to insufficient energy.
A resource-aware intelligent edge gateway system is adopted, which combines deep reinforcement learning algorithms and optimizes the offloading decision and resource allocation of computing tasks through the multi-agent hybrid decision-making algorithm (MCMHD) to achieve multi-region collaborative computing and reduce latency and power consumption.
It minimizes the latency of computing tasks and the power consumption of user devices, meets user needs, and dynamically adjusts computing performance and energy consumption under different power conditions.
Smart Images

Figure CN119545435B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of edge computing, and specifically relates to a resource-aware intelligent edge gateway system and a running method. BACKGROUND
[0002] The latest development of Internet of Things technology has realized the intelligentization and autonomous control of many important industrial and commercial systems, such as intelligent power networks, unmanned farms and smart home automation, which greatly improves work efficiency and saves labor costs. However, Internet of Things devices are limited by strict device size and production cost, usually carrying limited capacity batteries and energy-saving low-performance processors, which makes them unable to support more and more new applications that require high-performance computing, and these applications also have the same characteristics that they will generate massive real-time data that need to be calculated, which brings serious delay and network congestion to the Internet of Things system. In order to solve the above problems, the academic and industrial circles begin to focus on mobile edge computing (MEC) technology, which integrates servers on wireless access points and base stations, thereby realizing a large number of computing tasks on large-scale low-power devices. On the other hand, wireless power transmission (WPT) technology solves the problem of limited battery capacity by deploying dedicated energy transmitters to broadcast energy in the form of wireless broadcast energy to charge devices equipped with remote energy collection, thereby solving the problem of limited battery capacity. At present, wireless charging mobile edge computing (WPMEC) combined with the above two technologies has become a new paradigm for realizing self-sustainable mobile computing for delay-sensitive and computationally intensive applications. In a typical wireless charging mobile edge computing system, a hybrid access point (AP) integrated with a mobile edge computing server and an energy transmitter can wirelessly charge a large number of low-power wireless devices such as sensors and wearable devices, and these devices rely on the harvested energy to perform local computing or offload tasks to the hybrid access point (AP) to perform computing tasks.
[0003] Current research on wireless charging edge computing mainly focuses on the task offloading and resource allocation of single wireless charging edge server and multiple user devices. Many solutions are proposed for indicators such as time delay and energy consumption, for example, user cooperation, unmanned aerial vehicle assistance and reflective surface assistance. However, there are still problems of uneven resource allocation and insufficient energy utilization in wireless charging mobile edge computing. In order to meet the task time delay requirement, a large number of tasks are offloaded to the edge server. Since there is no cooperation between single servers, the server load is too heavy, and even exceeds the maximum computing capacity, resulting in the phenomenon of discarding tasks, which will greatly affect the stability and efficiency of the whole system. On the other hand, due to the limitation of the half-duplex characteristics of user equipment, most of the current solutions mainly consider the balance between computing time delay and overall power consumption under the condition of complete energy consumption, in order to minimize the overall time delay index or power consumption index. Especially in the face of the situation of explosive data volume, some devices may not be able to perform local computing or task offloading due to insufficient energy acquisition. SUMMARY
[0004] In order to solve the above technical problems, the present application provides a resource-aware intelligent edge gateway system composed of multiple gateway devices and multiple users and a running method. The energy, computing and transmission resource information of user equipment and gateway devices are obtained through resource-aware technology, and a deep reinforcement learning algorithm is combined to apply it to the offloading decision and resource allocation of the system. The charging and offloading time, offloading strategy and resource allocation are continuously optimized through learning, the joint multi-region intelligent edge gateway collaborative computing is reduced, the computing task time delay and user equipment power consumption are reduced, and the user demand is met.
[0005] In order to achieve the above purpose, the present application is realized by the following technical scheme:
[0006] The present application is a resource-aware intelligent edge gateway system, which comprises a plurality of wireless charging intelligent edge gateways. The intelligent edge gateway comprises a user equipment service and protocol conversion unit, a computing resource awareness unit, a gateway device service unit, a wireless charging unit, a device management and control unit, a data storage unit, a computing task offloading decision unit and a task computing unit.
[0007] The user equipment service and protocol conversion unit is used to access the user equipment of multiple communication protocols, and to perform data identification, data analysis and data packaging on the original data input by the user equipment, convert the data into data of a specified type and structure of the intelligent edge gateway, and send the parameters of the computing task in the data to the computing resource awareness unit. At the same time, the user equipment service and protocol conversion unit sends the device information when the user equipment is accessed and the device power to the device management and control unit;
[0008] The computing resource perception unit is responsible for sending a gateway computing resource perception command, the computing resource usage and communication state information of other intelligent edge gateways in the intelligent edge gateway system, and device information and gateway information to the computing task offloading decision unit after receiving the computing task request of the first device;
[0009] The gateway device service unit is used for sending a gateway device information perception command and a computing result, and receiving a gateway device request and information, including a gateway device ID and a computing waiting time delay The transmission rate of the transmission gateway between the intelligent gateway in the system
[0010] The wireless charging unit is used for sending a wireless energy signal to provide wireless energy for user devices in the range;
[0011] The device management and control unit is used for storing device information sent by user devices when accessing, assigning a gateway device ID to the accessed user devices through the device management and control unit, and controlling the wireless charging unit to start and stop through the device management and control unit;
[0012] The data storage unit is used for storing computing task related data and computing resource perception data;
[0013] The computing task offloading decision unit is used for deciding whether to offload the computing task to the current edge computing gateway or an edge computing gateway other than the current edge computing gateway in the intelligent edge gateway system, that is, another regional intelligent gateway, not other types or models of gateways, and deciding the time allocation of user device task offloading and wireless charging, the size allocation of user device computing task local computing and offloading, the energy consumption allocation of user devices for local computing and offloading, and the CPU frequency of user devices for local computing;
[0014] The task computing unit is used for computing the offloaded computing task to obtain a computing result and send the task source to the user device service and protocol conversion unit or the gateway device service unit.
[0015] Further improvements of the application are that the computing task includes a user device ID, a computing task data volume d n , a maximum tolerable time delay T n,max , a CPU cycle number Φ required for computing 1-bit data, an energy size collected in the last frame , and a user device power
[0016] Further improvements of the application are that the device information includes a maximum computing frequency f n,maxMaximum transmit power p n,max Maximum battery capacity and equipment power The user equipment is powered by a battery, has a wireless energy harvesting module, and is connected to the edge computing gateway via a wireless link.
[0017] This invention provides a method for operating a resource-aware intelligent edge gateway system, specifically including the following steps:
[0018] Step 1: The user equipment generates a computing task and sends a data packet containing the user equipment ID and the amount of computing task data d to the intelligent edge gateway. n The number of CPU cycles required to compute 1 bit of data in the task, Φ; and the maximum tolerable latency T of the task. n,max The amount of energy collected in the previous frame and device power The system requests the uninstallation task and waits for a response from the smart edge gateway.
[0019] Step 2: The intelligent edge gateway receives the task offloading request sent by the user equipment, converts it into data of a specified type and structure through the user equipment service and protocol conversion unit, and identifies it as a task offloading request;
[0020] Step 3: The user equipment service and protocol conversion unit will convert the gateway device ID in the task offloading request and calculate the task data volume d. n The number of CPU cycles required to compute 1 bit of data in the task, Φ; and the maximum tolerable latency T of the task. n,max The amount of energy collected in the previous frame and device power The data is encapsulated into a data packet and passed to the computing resource awareness unit, which then starts listening for computing offload requests. The time interval between these listeners is recorded as a listening time frame.
[0021] Step 4: The computing resource awareness unit receives the data packet, generates a computing resource awareness command, and sends a computing resource awareness request to the remaining edge gateways in the intelligent edge gateway system through the gateway device service unit to obtain their computing wait latency. Gateways i To gateways j transmission rate Immediately after the perception is completed, the data is encapsulated into a data packet and transmitted to the computing task offloading decision unit, which then waits for the listening time frame to end. After the listening time frame ends, the computing resource perception unit will receive the data, which includes the gateway device ID and the computing task data volume d. n The number of CPU cycles required to compute 1 bit of data in the task, Φ; and the maximum tolerable latency T of the task. n,max The amount of energy collected in the previous frame and device power User Equipment ID and Channel Gain h n After the data packets are uniformly encapsulated, they are passed to the computing task unloading decision unit;
[0022] Step 5: The computational task offloading decision unit determines the amount of computational task data d of the user equipment. n Equipment power Channel gain h n And the computational latency of gateways within the intelligent edge gateway system. Transmission speed between gateways The optimal wireless charging time ratio coefficient τ for the objective function is calculated using the MCMHD algorithm. n The user calculates the CPU frequency f locally. n Signal transmission power p n and uninstallation destination gateway x n ;
[0023] Step 6: The computation task offloading decision unit encapsulates the completed offloading decision into two data packets containing different information, and sends them to different user devices and gateways through the user equipment service and protocol conversion unit and the gateway device service unit, respectively. After receiving the offloading decision, the user equipment service and protocol conversion unit of the remaining gateway devices adds the user equipment information to be offloaded to local computation as a temporary user device to the device management and control unit, and at the same time adds the computation task to the computation waiting queue of the task computation unit. Queuing is being conducted.
[0024] Step 7: After receiving the offloading decision from the wireless charging edge computing gateway, the user equipment adjusts its operating parameters and follows the offloading calculation part d″ in the calculation decision. n If the value is 0, determine whether the computation task needs to be unloaded. If so, split the computation task into local computation parts d′ according to the unloading decision. n and unloading calculation part d″ n If not needed, only the local computation part d′ is required. n Local computation part d′ n The computational processing and offloading of the computational components begin within the user equipment's own computing unit. n The process begins by offloading the wireless component to the associated smart gateway, and then proceeds to the offloading computing section d″. n Once uninstallation is complete, charging will begin.
[0025] Step 8: After receiving the user device's offload task data, the user equipment service and protocol conversion unit of the wireless charging edge computing gateway checks the destination gateway device ID. If it is its own gateway device ID, it adds it to the task calculation unit's calculation waiting queue. And cache to data storage unit, otherwise add to gateway device service unit, continue to forward to unloading destination gateway, at the same time, after the user equipment completes unloading within the time period stipulated in the unloading decision, start collecting wireless energy according to the time period stipulated in the unloading decision, and wireless charging is carried out;
[0026] Step 9, unloading calculation part d" n After the task calculation is completed, the source gateway device ID and the user ID are returned to the user equipment;
[0027] Step 10, after the user equipment receives the unloading part task calculation result, the local calculation result is merged;
[0028] Step 11, after the calculation result is merged, the wireless energy collection task continues to complete the decision distribution remaining time, after the wireless energy collection time period is completed, whether there is a calculation task is judged, if there is a task, steps 1 to 11 are repeated, and if there is not, whether the battery power is 100% is detected, if the battery power does not reach 100%, charging is continued, and lasts until the next calculation task is generated or the power reaches 100%.
[0029] Further improvement of the application is that in the step 1, the energy size Is defined as:
[0030]
[0031] Wherein, μ is the energy conversion coefficient, p s Is the energy emission power of the intelligent edge gateway wireless charging module, τ n Is the wireless charging time proportion coefficient, T f Is the length of a time frame, c n Is the distance of the user equipment from the edge gateway, and α is the energy conversion efficiency index.
[0032] Further improvement of the application is that in the step 5, the target function is:
[0033]
[0034] Wherein, Is the local calculation time delay of the user equipment, Is the task unloading time delay, Is the time delay of all tasks using local calculation, Is the local calculation power consumption of the user equipment, Is the task unloading power consumption, Is the energy collection size of all tasks using local calculation, N is the total number of user equipment requesting calculation task unloading, and n is the user equipment requesting calculation task unloading, Is the gateway processing time delay, including calculation waiting time delay Processing latency And computing task transmission latency For energy collection size, For all tasks, the energy consumption size of local computing is adopted.
[0035] A further improvement of the present application is that in the step 5, the wireless charging time proportion coefficient τ is obtained by using the MCMHD algorithm n The user local computing CPU frequency f n The signal transmission power p n And the training process of the offloading destination gateway x n The training process includes the following steps:
[0036] Step 5.1, initialize the Actor training network μ(S) for generating a set of continuous actions con_a;
[0037] Step 5.2, initialize the Critic training network Q(S, con_a*, dis_a) for generating discrete actions dis_a and estimating action value Q_values;
[0038] Step 5.3, initialize the Actor target network μ'(S) and the Critic target network Q'(S, con_a*, dis_a) as a copy of the training network to stabilize learning while participating in action value estimation;
[0039] Step 5.4, initialize the experience replay pool Rp for storing experience data;
[0040] Step 5.5, initialize the state state;
[0041] Step 5.6, check whether the termination condition is met, if the maximum number of iterations is reached, end the training; otherwise, continue;
[0042] Step 5.7, generate a set of continuous actions con_a using the Actor training network μ(S) and the current state state;
[0043] Step 5.8, select a discrete action dis_a using the Critic training network Q(S, con_a*, dis_a) in combination with the set of continuous actions con_a and the current state state;
[0044] Step 5.9, obtain the optimal continuous action con_a* in combination with the set of continuous actions con_a and the discrete action dis_a;
[0045] Step 5.10, interacting with the system environment Env of the MCMHD algorithm through the state state, the continuous action con_a*, the discrete action dis_a, and the system environment Env of the MCMHD algorithm to obtain the next state next_state and the reward reward, Env being a technical index class designed according to the system model of the MCMHD algorithm and the reward reward;
[0046] Step 5.11, storing the experience state, con_a*, dis_a, reward, next_state into the experience replay pool Rp;
[0047] Step 5.12, if the experience replay pool Rp accumulates enough data of a set minimum batch size, training a set number of data from the batch memory batch_memory;
[0048] Step 5.13, using the state in the batch_memory and the Actor target network μ'(S) to predict the continuous action set next_con_a of the next state;
[0049] Step 5.14, using the continuous action set next_con_a of the next state and the next_state in the batch_memory, and the Critic training network Q(S, con_a*, dis_a) and the Critic target network Q'(S, con_a*, dis_a) to calculate the action value set next_Q_values of the next state;
[0050] Step 5.15, obtaining the discrete action next_dis_a of the next state and the maximum action value next_Q_values* of the next state through the action value set next_Q_values of the next state;
[0051] Step 5.16, calculating the target action value targets_Q_values, which is the immediate reward reward plus the discounted next_Q_values*;
[0052] Step 5.17, using the Critic network Q(S, con_a*, dis_a) and the state state, the continuous action con_a set, and the discrete action dis_a in the batch_memory to calculate the predicted action value Q_values;
[0053] Step 5.18, calculating the loss of the Critic, which is the difference between the predicted action value Q_values and the target action value targets_Q_values, and performing backpropagation to update the Critic network;
[0054] Step 5.19, using the Critic network Q(S, con_a*, dis_a) to evaluate the state in the batch_memory and the action value predicted by the Actor network μ(S), calculating the loss of the Actor, backpropagating and updating the Actor network;
[0055] Step 5.20, using the soft update policy to update the parameters of the target network, so that the parameters of the target network are biased towards the parameters of the training network;
[0056] Step 5.21, the environment state is updated, and the state state is updated to next_state;
[0057] Step 5.22, repeat steps 5.6-5.21, repeat the whole process until the termination condition is met.
[0058] Further improvement of the application is that in the step 7, after task splitting, the local computing part d' n Immediately calculate locally on the user equipment, offload the computing part d" n At the same time, offload to the intelligent edge gateway in decision-making.
[0059] The beneficial effects of the application are:
[0060] The application combines wireless charging and edge computing technology to realize the offloading of computing tasks of different user equipment to the intelligent edge gateway for calculation, and at the same time complete a charging process;
[0061] The offloading strategy adopts the processing mode of cooperative calculation between the device and the intelligent edge gateway and the cooperative calculation between the intelligent edge gateways;
[0062] The application perceives the resources of the user equipment including energy, calculation and transmission, and proposes a calculation task for different user equipment, and combines the device wireless charging and the remaining power, adopts the multi-agent hybrid decision algorithm (MCMHD) of joint continuous action space and discrete action space, realizes the task offloading of multiple devices under the gateway management, and the resource allocation and offloading scheme of optimal wireless charging time balance and device calculation task allocation, minimizes the weighted sum of user equipment task delay and power consumption. BRIEF DESCRIPTION OF DRAWINGS
[0063] Figure 1 The figure is the architecture diagram of the intelligent edge gateway system of the application.
[0064] Figure 2 The figure is the composition block diagram of the intelligent edge gateway system of the application.
[0065] Figure 3This is a flowchart illustrating the overall system operation of the present invention.
[0066] Figure 4 This is a flowchart of the task unloading operation on the user equipment side of the present invention.
[0067] Figure 5 This is a flowchart of the task unloading operation on the edge gateway side of the present invention.
[0068] Figure 6 This is a flowchart of the MCMHD algorithm of the present invention.
[0069] Figure 7 The diagram shows the convergence effect of the MCMHD algorithm in the system of this invention.
[0070] Figure 8 This is a comparison of the convergence performance of the MADDPG algorithm, DDPG algorithm, and MADDPG algorithm in this system.
[0071] Figure 9 This is a comparison chart of task completion rate and device power failure rate between the MCMHD algorithm, DDPG algorithm, and MADDPG algorithm.
[0072] Figure 10 This is a comparison chart showing the impact of the average completion delay of the MCMHD algorithm, DDPG algorithm, and MADDPG algorithm of this invention.
[0073] Figure 11 This is a comparison chart showing the impact of the average completion energy consumption of the MCMHD algorithm, DDPG algorithm, and MADDPG algorithm of this invention.
[0074] Figure 12 This is a comparison chart showing the impact of task size on average latency between the present invention and systems without uninstallation, cloud uninstallation systems, and single-gateway uninstallation systems.
[0075] Figure 13 This is a comparison chart showing the impact of gateway computing capacity on the average computing latency of tasks compared to the present invention, a no-unloading system, a cloud-based unloading system, and a single-gateway unloading system. Detailed Implementation
[0076] The embodiments of the present invention will be disclosed below with reference to the drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential.
[0077] like Figure 1 As shown, this invention provides a resource-aware intelligent edge gateway system, which includes several wireless charging intelligent edge gateways, such as... Figure 2As shown, the intelligent edge gateway includes a user device service and protocol conversion unit, a computing resource awareness unit, a gateway device service unit, a wireless charging unit, a device management and control unit, a data storage unit, a computing task offloading decision unit, and a task computing unit for computing resource awareness, computing task offloading decision, task computing, data storage, wireless charging control, and device management.
[0078] The user device service and protocol conversion unit is used to access user devices with Ethernet, Wi-Fi, Zigbee, GPRS, GPS, 3G, 4G, 5G, REST / HTTP, MQTT, CoAP, DDS, AMQP, and other communication protocols, and to perform data recognition, data analysis, and data packaging on raw data input by user devices, converting them into data of a specified type and structure of the intelligent edge gateway, and sending parameters related to computing tasks in the data to the computing resource awareness unit. At the same time, the user device service and protocol conversion unit sends device information when the user device is connected and device power to the device management and control unit;
[0079] The computing resource awareness unit is responsible for sending gateway computing resource awareness commands to the gateway device service unit after receiving the first device's computing task request, as well as the computing resource usage and communication state information of other intelligent edge gateways in the intelligent edge gateway system, and forwarding the device information and gateway information to the computing task offloading decision unit. The computing task includes user device ID, computing task data volume d n , maximum tolerable time delay T n , CPU cycle number Φ required for computing task 1-bit data, energy size collected in the last frame , and user device power
[0080] The gateway device service unit is used to send gateway device information awareness commands and computing results, and to receive gateway device requests and information, including gateway device ID, computing waiting time delay , and transmission rate of the transmission gateway between the intelligent gateway in the system The device information includes maximum computing frequency f n,max , maximum transmission power p n,max , maximum battery capacity , and device power The user device is powered by a battery and has a wireless energy harvesting module, which accesses the edge computing gateway through a wireless link.
[0081] The wireless charging unit is used to send wireless energy signals to provide wireless energy for user devices within a certain range.
[0082] The device management and control unit is configured to store device information sent by the user equipment when accessing, assign a gateway device ID to the accessed user equipment through the device management and control unit, and control the wireless charging unit to turn on and off through the device management and control unit.
[0083] The data storage unit is configured to store computing task related data and computing resource perception data.
[0084] The computing task offloading decision unit is configured to make decisions on offloading a computing task to a current edge computing gateway or an edge computing gateway other than the current edge computing gateway in the intelligent edge gateway system, and on time allocation of user equipment task offloading and wireless charging, size allocation of user equipment computing task local computing and offloading, energy consumption allocation of user equipment for local computing and offloading, and CPU frequency of user equipment for local computing.
[0085] The task computing unit is configured to compute the offloaded computing task, obtain a computing result, and send a task source to the user equipment service and protocol conversion unit or the gateway device service unit.
[0086] The application also provides a running method of the resource-perception intelligent edge gateway system. Figure 3 As shown in the system overall flowchart, there are several intelligent edge gateways in the resource-perception intelligent edge gateway system, each of which controls the wireless charging process and the computing task offloading process of user equipment.
[0087] The user end running flowchart (including steps 1, 7, 10 and 11) is as shown in Figure 4 The gateway end running flowchart (including steps 2-6-9) is as shown in Figure 5 The number of intelligent edge gateways is determined by the specific system. The embodiment only serves to illustrate.
[0088] The running method of the application specifically includes the following steps:
[0089] Step 1, the user equipment u i generates a computing task, and sends an offloading task request containing a user equipment ID, a computing task data volume d n , a CPU cycle number Φ required by 1 bit of computing task data, a maximum tolerable time delay T n,max of the task, an energy size collected in the last frame, and a device power to the wireless charging intelligent edge gateway s j , and waits for a reply from the wireless charging intelligent edge gateway s , wherein the energy size is defined as:
[0090]
[0091] Where μ is the energy conversion coefficient, p s For the energy transmission power of the wireless charging module of the smart edge gateway, τ n T is the wireless charging time scaling factor. f c is the length of a time frame. n α represents the distance between the user equipment and the edge gateway, and α is the energy conversion efficiency index.
[0092] like Figure 5 As shown, step 2, intelligent edge gateway s j Upon receiving a task offloading request from a user equipment, the user equipment service and protocol conversion unit converts it into data of a specified type and structure, and identifies it as a task offloading request.
[0093] Step 3: The user equipment service and protocol conversion unit will convert the gateway device ID in the task offloading request and calculate the task data volume d. n The number of CPU cycles required to compute 1 bit of data in the task, Φ; and the maximum tolerable latency T of the task. n,max The amount of energy collected in the previous frame and device power The data is encapsulated into a data packet and passed to the computing resource awareness unit, which then starts listening for computing offload requests. The time interval between these listeners is recorded as a listening time frame.
[0094] Step 4: The computing resource awareness unit receives the data packet, generates a computing resource awareness command, and sends a computing resource awareness request to other remaining edge gateways in the intelligent edge gateway system through the gateway device service unit to obtain their computing wait latency. Gateways i To gateways j transmission rate (0 if there is no connection). After the perception ends, the data is immediately encapsulated into a data packet and transmitted to the computing task unloading decision unit, and then waits for the listening time frame to end. Since the listening time frame is very short, mainly to ensure that all task unloading requests can be received at the beginning of the time frame, it can be ignored when optimizing the objective function. After the listening time frame ends, the computing resource perception unit will receive the data including the gateway device ID and the computing task data volume d. n The number of CPU cycles required to compute 1 bit of data in the task, Φ; and the maximum tolerable latency T of the task. n,max The amount of energy collected in the previous frame and device power User Equipment ID and Channel Gain h n After the data packets are uniformly encapsulated, they are passed to the computing task unloading decision unit;
[0095] Step 5, the computing task offloading decision unit calculates the computing task data volume d of the user equipment n , the device power channel gain h n and the computing waiting delay of the gateway in the intelligent edge gateway system and the transmission speed between the gateways The MCMHD algorithm is used to calculate the optimal wireless charging time proportion coefficient τ for the objective function n , the user local computing CPU frequency f n , the signal transmission power p n and the offloading destination gateway x n , the wireless charging time proportion coefficient τ n , the wireless charging time proportion coefficient τ n and the signal transmission power p n are continuous variables, the offloading destination gateway x n is a discrete variable, and the objective function is:
[0096]
[0097] wherein, is the local computing delay of the user equipment, is the task offloading delay, is the delay of all tasks using local computing, is the local computing power consumption of the user equipment, is the task offloading power consumption, is the energy harvesting size of all tasks using local computing, N is the total number of user equipment requesting computing task offloading, and n is the user equipment requesting computing task offloading, is the gateway processing delay, including the computing waiting delay processing delay and the computing task transmission delay is the energy harvesting size, is the energy consumption size of all tasks using local computing.
[0098] Taking the device u1 in the intelligent edge gateway s1 region as an example, r i is the offloading rate of the user equipment, represents that u1 belongs to the intelligent edge gateway s1 and is allocated to the intelligent edge gateway s1, i.e., no secondary offloading to other gateways is needed, the computing waiting delay gateway task processing delay gateway task transmission delay delay of all tasks using local computing power consumption task offloading power consumption Energy collection size
[0099] Further, taking the device u2 in the smart edge gateway s1 area as an example, Representing that u2 belongs to the smart edge gateway s1, is allocated to the smart edge gateway s5, that is, needs to be unloaded to the gateway s5 twice, and the calculation waiting delay Gateway task processing delay Gateway task transmission delay Wherein Path is the best path of s1 reaching s5 in the routing table, and hop is the number of hops on the path.
[0100] Step 6, the calculation task unloading decision unit encapsulates the calculated unloading decision into two data packets containing different information, and sends them to different user devices and gateways through the user device service and protocol conversion unit and the gateway device service unit respectively. After the user device service and protocol conversion unit of the remaining gateway device receives the unloading decision, the user device information to be unloaded to local calculation is added to the device management and control unit as a temporary user device, and at the same time, the calculation task is added to the calculation waiting queue of the task calculation unit for queuing.
[0101] Step 7, after the user device receives the unloading decision of the wireless charging edge computing gateway, adjusts the working parameters of the device according to the flow shown in Figure 3 , and according to the unloading calculation part d″ n in the calculation decision, judges whether the calculation task needs to be unloaded or not. If it needs to be unloaded, the calculation task is split into local calculation part d′ n and unloading calculation part d″ n , if not, only the local calculation part d′ n , the local calculation part d′ n starts to calculate and process in the calculation unit of the user device itself, and the unloading calculation part d″ n starts to unload to the belonging smart gateway through the wireless component, and after the unloading calculation part d″ n is unloaded, starts to time the charging, after the task is split, the local calculation part d′ n is calculated immediately in the user device, and the unloading calculation part d″ n is unloaded to the smart edge gateway in the decision at the same time.
[0102] Step 8, after the user device service and protocol conversion unit of the wireless charging edge computing gateway receives the unloading task data of the user device, checks the destination gateway device ID. If it is the gateway device ID itself, it is added to the calculation waiting queue of the task calculation unit It is cached in the data storage unit; otherwise, it is added to the gateway device service unit and forwarded to the unloading destination gateway. Meanwhile, after the user device completes the unloading within the time specified by the unloading decision, it begins to collect wireless energy for wireless charging according to the time period specified by the unloading decision.
[0103] Step 9: Unload the computing component d″ n After the task calculation is completed, the data is sent back to the user device based on the source gateway device ID and the user ID.
[0104] Step 10: After receiving the calculation results of the unloading part of the task, the user equipment merges them with the local calculation results;
[0105] Step 11: After merging the calculation results, the wireless energy harvesting task continues to complete the decision allocation of the remaining time. After the wireless energy harvesting period is completed, it is determined whether there is a calculation task. If there is a task, steps 1 to 11 are repeated. If there is no task, it is checked whether the battery level is 100%. If the battery level is not 100%, it continues to charge until the next calculation task is generated or the battery level reaches 100%.
[0106] This invention employs a multi-agent hybrid decision-making algorithm (MCMHD) based on an improved MADDPG, which combines continuous and discrete action spaces. The invention utilizes a centralized training and distributed execution model for the multi-agent system, with each user device acting as an Actor module while sharing a common Critic module. This allows each user device to share continuous variables in the decision-making process while coordinating discrete variables. The state in the algorithm includes the amount of task data, d. n Equipment power Wireless link channel gain h n Energy collected from the user device in the previous time frame Computation latency of each smart edge gateway Actions include the continuous variable wireless charging time scaling factor τ. n User equipment CPU computing frequency f n and signal transmission power p n and selection of destination gateway for unloading discrete variables The constructed reward consists of four parts: the first part is the objective function reward r1, which is the weighted sum of normalized latency and energy consumption with respect to battery capacity; the second part is the task processing timeout penalty r2, which is the penalty for task processing latency exceeding the maximum tolerable latency; the third part is the task processing energy consumption exceeding penalty r3; and the fourth part is the low battery capacity penalty, which is the device's remaining battery power. A penalty r4 is applied when the battery level falls below the minimum battery threshold. The r3 part introduces information related to the device's battery level. S-shaped growth function The product of the energy collected by the user equipment in a time frame and the energy consumption threshold of the current computing task The energy consumption threshold makes the agent obtain different violation penalties at different battery levels, and combines the battery level of the equipment in r1 as a weight coefficient to adjust the weight of the time delay and the energy consumption in the objective function, so that the application can dynamically adjust the computing performance and energy consumption of the user equipment. The special reward composition makes the user equipment under the MCMHD algorithm pay more attention to the time delay when the battery level is good, and pay more attention to the energy consumption when the battery level is low. Under different conditions, the computing performance and energy consumption of the equipment are dynamically adjusted.
[0107] The application uses the MCMHD algorithm to process problems with continuous action and discrete action space, and the specific process is as shown in Figure 6 The application uses the MCMHD algorithm to process problems with continuous action and discrete action space, and the specific process is as shown in
[0108] Step 5.1, initializing the Actor training network μ (S) for generating a continuous action set con_a;
[0109] Step 5.2, initializing the Critic training network Q (S, con_a*, dis_a) for generating a discrete action dis_a and estimating the action value Q_values;
[0110] Step 5.3, initializing the Actor target network μ'(S) and the Critic target network Q'(S, con_a*, dis_a) as a copy of the training network to stabilize learning, and participating in action value estimation at the same time;
[0111] Step 5.4, initializing the experience replay pool Rp for storing experience data;
[0112] Step 5.5, initializing the state state;
[0113] Step 5.6, checking whether the termination condition is met, if the maximum number of iterations is reached, ending the training; otherwise, continuing;
[0114] Step 5.7, generating a continuous action set con_a using the Actor training network μ (S) and the current state state;
[0115] Step 5.8, using the Critic training network Q (S, con_a*, dis_a) to select a discrete action dis_a in combination with the continuous action set con_a and the current state state;
[0116] Step 5.9, obtaining the optimal continuous action con_a* in combination with the continuous action set con_a and the discrete action dis_a;
[0117] Step 5.10, interacting with the system environment Env of the MCMHD algorithm through the state state, the continuous action set con_a*, the discrete action set dis_a, and the system environment Env of the MCMHD algorithm to obtain the next state next_state and the reward reward, and the technical index class Env is designed according to the system model of the MCMHD algorithm and the reward reward;
[0118] Step 5.11, storing the experience state, con_a*, dis_a, reward, and next_state into the experience replay pool Rp; Step 5.12, if the experience replay pool Rp accumulates enough data of a set minimum batch size, training a set number of data from the batch data batch_memory;
[0119] Step 5.13, using the state in the batch_memory and the Actor target network μ'(S) to predict the continuous action set next_con_a of the next state;
[0120] Step 5.14, using the continuous action set next_con_a of the next state and the next_state in the batch_memory, and the Critic training network Q(S, con_a*, dis_a) and the Critic target network Q'(S, con_a*, dis_a) to calculate the action value set next_Q_values of the next state;
[0121] Step 5.15, obtaining the discrete action next_dis_a of the next state and the maximum action value next_Q_values* of the next state through the action value set next_Q_values of the next state;
[0122] Step 5.16, calculating the target action value targets_Q_values, which is the immediate reward reward plus the discounted next_Q_values*;
[0123] Step 5.17, using the Critic network Q(S, con_a*, dis_a) and the state state, the continuous action set con_a, and the discrete action dis_a in the batch_memory to calculate the predicted action value Q_values;
[0124] Step 5.18, calculating the loss of the Critic, which is the difference between the predicted action value Q_values and the target action value targets_Q_values, and performing backpropagation to update the Critic network;
[0125] Step 5.19, evaluate the state in the batch_memory and the action value predicted by the Actor network μ(S) using the Critic network Q(S, con_a*, dis_a), calculate the loss of the Actor, backpropagate and update the Actor network;
[0126] Step 5.20, update the parameters of the target network using the soft update policy, so that the parameters of the target network are slightly biased towards the parameters of the training network;
[0127] Step 5.21, the environment state is updated, and the state state is updated to next_state;
[0128] Step 5.22, repeat steps 5.6-5.21, repeat the entire process until the termination condition is met.
[0129] The above is the wireless charging time proportion coefficient τ of the MCMHD algorithm n , the user local computing CPU frequency f n , the signal transmission power p n and the decision x of the offloading destination gateway. n training process.
[0130] In order to verify the present application, the simulation experiment is provided as follows:
[0131] The simulation system includes 8 user devices and 5 intelligent gateway devices, one of which is the intelligent gateway device in the current area, responsible for transmitting wireless energy to the 8 user devices and making decisions on task offloading requests. The 8 user devices need to receive wireless energy transmitted by an intelligent gateway device and offload part of the generated computing task to the specified intelligent gateway in the decision. In the simulation, the computing tasks of the user devices use the same task size and maximum tolerable delay, and the movement range is specified within 10 to 20 meters from the current area intelligent gateway device. The simulation tests the performance of the system in continuously processing 100 times of the same computing task, and compares it with the system based on the DDPG algorithm and the MADDPG algorithm. Figure 7 The convergence effect diagram of the MCMHD algorithm in the system of the present application is shown in the figure, in which the reward value reaches 1.3 or more, close to the maximum reward value 1.5 set by the system. Figure 8 The convergence effect of the MADDPG algorithm and the DDPG algorithm and the MADDPG algorithm in the system is compared, and it can be seen that the improved MCMHD algorithm converges faster than the original MADDPG algorithm, and the explored reward value is higher. Although the DDPG algorithm converges faster, the maximum reward value explored is lower. Since the decision strategy with higher reward is found, Figure 9 In the figure, it can be seen that the MCMHD algorithm achieves higher task completion rate and lower device power failure rate. Figure 10 andFigure 11 It also shows that the MCMHD algorithm is more adaptable to the change of task size. Figure 12 and Figure 13 This shows the advantage of the MCMHD algorithm system compared with the non-unloading system, the cloud-based task unloading system and the single intelligent gateway system in terms of the average computing delay of tasks, where the vertical coordinate average delay ratio is defined as the average delay ratio of the MCMHD algorithm system to other mode systems.
[0132] The above merely illustrates the embodiments of the present application, and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. A resource-aware intelligent edge gateway system, characterized by: The intelligent edge gateway system comprises a plurality of wireless charging intelligent edge gateways, and the intelligent edge gateway comprises a user equipment service and protocol conversion unit, a computing resource awareness unit, a gateway equipment service unit, a wireless charging unit, an equipment management and control unit, a data storage unit, a computing task offloading decision unit, and a task computing unit. The user equipment service and protocol conversion unit is used for accessing user equipment of multiple communication protocols, and performing data identification, data analysis, and data encapsulation on original data input by the user equipment, converting the original data into data of a type and structure specified by the intelligent edge gateway, and sending parameters about a computing task in the data to the computing resource perception unit. Meanwhile, the user equipment service and protocol conversion unit sends device information when the user equipment is accessed and device power perceived at fixed times to the device management and control unit. to the device management and control unit. The computing resource awareness unit is responsible for sending a gateway computing resource awareness command to the gateway equipment service unit after receiving a computing task request from a first device, and the resource awareness unit is responsible for sending the computing resource usage and communication state information of other intelligent edge gateways in the intelligent edge gateway system to the computing task offloading decision unit. The gateway device service unit is used to send gateway device information awareness commands and calculation results, receive gateway device requests and information, including gateway device ID, calculation waiting time delay Transmission rate of the transmission gateway between the intelligent edge gateway in the system The wireless charging unit is used to send wireless energy signals to provide wireless energy for user equipment within a range. The equipment management and control unit is used to store device information sent by user equipment when accessing the gateway, and the gateway equipment ID is allocated to the user equipment through the equipment management and control unit. The data storage unit is used to store computing task related data and computing resource awareness data. The computing task offloading decision unit is used to decide whether to offload the computing task to the current intelligent edge gateway or to an intelligent edge gateway other than the current intelligent edge gateway in the intelligent edge gateway system, and to decide the time allocation of user equipment task offloading and wireless charging, the size allocation of user equipment computing task local computing and offloading, the energy consumption allocation of user equipment for local computing and offloading, and the CPU frequency of user equipment for local computing. The task computing unit is used to compute the offloaded computing task to obtain a computing result and send the task source to the user equipment service and protocol conversion unit or the gateway equipment service unit. The intelligent edge gateway system adopts a multi-agent hybrid decision algorithm MCMHD based on an improved MADDPG joint continuous action space and discrete action space, wherein a multi-agent centralized training and decentralized execution mode is adopted, each user equipment is an Actor module, and a Critic module is shared, so that each user equipment shares continuous variables in decision making while coordinating discrete variables.
2. The resource-aware intelligent edge gateway system of claim 1, wherein: The computing task includes a user equipment ID, a computing task data volume d n , a maximum tolerable time delay T of the task n,max , a number of CPU cycles Ф required for 1 bit of data of the computing task, an energy size collected in a previous frame , and a user equipment power 3. The resource-aware intelligent edge gateway system of claim 1, wherein: The device information includes a maximum calculation frequency f n,max , a maximum transmission power p n,max , a maximum battery capacity , and a device power level 4. The resource-aware intelligent edge gateway system of claim 1, wherein: The user equipment is powered by a battery and has a wireless energy harvesting module and accesses the intelligent edge gateway through a wireless link.
5. A method of operating a resource-aware intelligent edge gateway system as claimed in any of claims 1-4, characterized by: In the resource-aware intelligent edge gateway system, there are a plurality of intelligent edge gateways, and each intelligent edge gateway controls the wireless charging process and computing task offloading process of user equipment, and the secondary offloading of intelligent edge gateways realizes the cooperation between intelligent edge gateways. Step 1, the user equipment generates a computing task, sends a task offloading request containing the user equipment ID, the computing task data volume d n , the number of CPU cycles required by the computing task 1 bit data Ф, the maximum tolerable time delay T of the task n,max , the energy size collected in the last frame , and the device power , and waits for the reply of the intelligent edge gateway; Step 2: The intelligent edge gateway receives the task offloading request sent by the user equipment, converts it into data of a specified type and structure through the user equipment service and protocol conversion unit, and identifies it as a task offloading request. Step 3, the user equipment service and protocol conversion unit unloads the gateway device ID in the task offloading request, the calculation task data volume d n , the number of CPU cycles required by the calculation task 1 bit data Ф, the maximum tolerable time delay T of the task n,max , the energy size collected in the last frame and the device power Packaged as a data packet to the computing resource awareness unit, and start listening to the task offloading request, record the listening time gap as a listening time frame; Step 4, the computing resource awareness unit receives the data packet, generates a computing resource awareness command, and sends a computing resource awareness request to the remaining intelligent edge gateway in the intelligent edge gateway system through the gateway device service unit to obtain the computing waiting time delay thereof Transmission rate between each gateway After the awareness is completed, the data is immediately encapsulated into a data packet and delivered to the computing task offloading decision unit and waits for the end of the listening time frame. After the end of the listening time frame, the computing resource awareness unit receives the data packet containing the gateway device ID, the computing task data volume d n , the number of CPU cycles required for the computing task 1 bit data Ф, the maximum tolerable time delay T of the task n,max , the energy size collected in the last frame , and the device power , the user device ID, and the channel gain h n After the data packet is uniformly encapsulated, it is delivered to the computing task offloading decision unit; Step 5, the computing task offloading decision unit calculates the computing task data volume d of the user equipment n , the equipment power , the channel gain h n , the computing waiting delay of the gateway in the intelligent edge gateway system, and the transmission rate between the gateways The MCMHD algorithm is used to calculate the optimal wireless charging time proportion coefficient τ for the objective function n , the user local computing CPU frequency f n , the signal transmission power p n , and the offloading target gateway x n Step 6, the computing task offloading decision unit encapsulates the completed offloading decision into two data packets containing different information, and sends them to different user equipment and gateways through the user equipment service and protocol conversion unit and the gateway device service unit respectively. After the user equipment service and protocol conversion unit of the remaining gateway device receives the offloading decision, it adds the user equipment information to be offloaded to local computing to the device management and control unit as a temporary user equipment, and simultaneously adds the computing task to the computing waiting queue of the task computing unit for queuing. Step 7, after receiving the offloading decision of the wireless charging intelligent edge gateway, the user equipment adjusts the working parameters of the equipment, and carries out the offloading calculation part d" according to the offloading decision n whether it is 0, whether the calculation task needs to be offloaded, if it needs to be offloaded, the calculation task is split into a local calculation part d' n and an offloaded calculation part d" n if it does not need to be offloaded, only the local calculation part d' n the local calculation part d' n starts to perform calculation processing and the offloaded calculation part d" in the calculation unit of the user equipment itself n starts to offload to the intelligent edge gateway through the wireless component, and performs the offloaded calculation part d" n starts to count the charging time after the offloading is completed; Step 8, after the user equipment service and protocol conversion unit of the wireless charging intelligent edge gateway receives the offloading task data of the user equipment, it checks the destination gateway device ID. If it is the ID of the self gateway device, it is added to the calculation waiting queue of the task calculation unit and cached to the data storage unit, otherwise it is added to the gateway device service unit and continues to forward to the offloading destination gateway. At the same time, after the user equipment completes offloading within the time specified by the offloading decision, it starts to collect wireless energy and perform wireless charging according to the time period specified by the offloading decision. Step 9, unload the computing part d" n After the task is completed, the source gateway device ID and the user ID are returned to the user device. Step 10: After the user equipment receives the offloaded task computing result, the local computing result is combined. Step 11, after the calculation result is merged, the wireless energy collection task continues to complete the decision distribution of the remaining time, and after the wireless energy collection time period is completed, it is judged whether there is a calculation task, if there is a task, steps 1 to 11 are repeated, and if not, it is detected whether the battery power is 100%, if the battery power has not reached 100%, the charging continues until the next calculation task is generated or the power reaches 100%.
6. The method of Claim 5, wherein: In said step 1, the energy size is defined as: wherein μ is the energy conversion coefficient, p s is the energy transmission power of the intelligent edge gateway wireless charging module, τ n is the wireless charging time proportionality coefficient, T f is the length of a time frame, c n is the distance between the user equipment and the edge gateway, and α is the energy conversion efficiency index.
7. The method of Claim 5, wherein: The state state in the MCMHD algorithm includes the data volume d of the computing task n , the device power , the wireless link channel gain h n , the energy size collected by the user equipment in the last time frame , the computing waiting delay of each intelligent edge gateway The action action includes the continuous variable wireless charging time proportion coefficient τ n , the user local computing CPU frequency f n , the signal transmission power p n , and the discrete variable offloading target gateway selection The reward constructed by the agent is composed of four parts. The first part is the objective function reward r1, which is the weighted sum of the normalized delay and energy consumption with respect to the battery power. The second part is the task processing timeout penalty r2, which is the penalty obtained when the task processing delay exceeds the maximum tolerable delay of the task. The third part is the task processing energy consumption excess penalty r3. The fourth part is the battery power low penalty, which is the penalty r4 obtained when the device power is lower than the minimum power threshold. Among them, the task processing energy consumption excess penalty r3 introduces the product of the S-shaped growth function with respect to the device power and the energy size collected by the user equipment in the last time frame as the energy consumption threshold of the current computing task. The energy consumption threshold makes the agent obtain different violation penalties at different powers, and combines the device power as the weight coefficient in the objective function reward r1 to adjust the weights of the delay and energy consumption in the objective function.
8. The method of claim 5 or 7, wherein: In the step 5, the wireless charging time proportion coefficient τ is obtained by using the MCMHD algorithm n , the user local computing CPU frequency f n , the signal transmitting power p n and the offloading destination gateway x n The training process includes the following steps: Step 5.1, initialize the Actor training network μ(S) for generating a set of continuous actions con_a; Step 5.2, initialize the Critic training network Q(S, con_a*, dis_a) for generating discrete actions dis_a and estimating action values Q_values; Step 5.3, initialize the Actor target network μ'(S) and the Critic target network Q'(S, con_a*, dis_a) as copies of the training networks to stabilize learning while participating in action value estimation; Step 5.4, initialize the experience replay pool Rp for storing experience data; Step 5.5, initialize the state state; Step 5.6, check whether the termination condition is met, if the maximum number of iterations is reached, end the training; Otherwise, continue; Step 5.7, generate a set of continuous actions con_a using the Actor training network μ(S) and the current state state; Step 5.8, select a discrete action dis_a using the Critic training network Q(S, con_a*, dis_a) in combination with the set of continuous actions con_a and the current state state; Step 5.9, obtain the optimal continuous action con_a* in combination with the set of continuous actions con_a and the discrete action dis_a; Step 5.10, interact with the system environment Env of the MCMHD algorithm through the state state, con_a*, dis_a, and the system environment Env of the MCMHD algorithm to obtain the next state next_state and the reward reward, Env is a technical index designed according to the system model of the MCMHD algorithm and the reward reward; Step 5.11, store the experience data state, con_a * , dis_a, reward, next_state in the experience replay pool Rp; Step 5.12, if the experience replay pool Rp accumulates enough data of a set minimum batch size, train a set number of batches of experience data batch_memory extracted therefrom; Step 5.13, use the states in batch_memory and the Actor target network μ'(S) to predict the set of continuous actions next_con_a of the next state; Step 5.14, use the set of continuous actions next_con_a of the next state and the next_state in batch_memory, and the Critic training network Q(S, con_a*, dis_a) and the Critic target network Q'(S, con_a*, dis_a) to calculate the set of action values next_Q_values of the next state; Step 5.15, get the next discrete action next_dis_a and the maximum next state action value next_Q_values* from the next state action value set next_Q_values; Step 5.16, calculate the target action value targets_Q_values, which is the immediate reward reward plus the discounted next_Q_values*; Step 5.17, calculate the predicted action value Q_values using the Critic training network Q(S, con_a*, dis_a) and the state state, the continuous action con_a set and the discrete action dis_a in the batch_memory; Step 5.18, calculate the loss of the Critic, which is the difference between the predicted action value Q_values and the target action value targets_Q_values, and perform backpropagation to update the Critic training network; Step 5.19, evaluate the state in the batch_memory and the action value predicted by the Actor training network μ(S) using the Critic training network Q(S, con_a*, dis_a), calculate the loss of the Actor, perform backpropagation and update the Actor training network; Step 5.20, update the parameters of the target network using the soft update policy, so that the parameters of the target network are biased towards the parameters of the training network; Step 5.21, update the environment state, update the state state to next_state; Step 5.22, repeat steps 5.6-5.21, repeat the entire process until the termination condition is met. 9.The method of Claim 5, wherein: In said step 7, after task splitting, local computing part d' n Immediately compute locally on user device, offload computing part d" n Also offload to intelligent edge gateway in decision making.
Citation Information
Patent Citations
Industrial Internet of Things cloud edge collaborative unloading and resource allocation method based on DDPG-D3QN
CN116390125A
Federal learning and task collaboration method for automatic driving vehicle-mounted edge calculation
CN118689215A