Edge computing unloading decision-making system based on DQN algorithm

Through the edge computing offload decision system based on DQN algorithm, the task success rate and convergence speed of edge computing are improved, network delay and security risks of traditional cloud computing in the era of Internet of Things are solved, and efficient task offload decisions are achieved.

CN120264355APending Publication Date: 2025-07-04CHINA IND INTERNET RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510511690.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When the traditional cloud computing paradigm meets the challenges of the Internet of Things, network architectures have problems such as transmission bottlenecks brought about by massive device access, high bandwidth costs for remote data transmission, and data privacy leakage risks, and it is difficult to meet computing-intensive tasks and low latency requirements.

Method used

The edge computing offload decision system based on the DQN algorithm is adopted, and the exponential attenuation mechanism of the ε-greedy strategy and the Dueling Double DQN architecture are combined with the dynamic Gamma value adjustment strategy to build a state space and action decision mechanism to optimize the offload decision of the edge computing task.

Benefits of technology

It improves the convergence speed and task success rate of edge computing offload decisions, and improves the overall performance by 17.6%, effectively solves the network latency and security risks of traditional cloud computing, and adapts to task offloading in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264355A_ABST
    Figure CN120264355A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of edge computing, in particular to an edge computing unloading decision-making system based on a DQN algorithm. Comprising an environment modeling unit, a state space construction unit, an action decision-making unit, a reward feedback unit, a training and optimizing unit and a simulation verification unit, wherein the environment modeling unit is used for modeling the heterogeneity of terminal equipment where an edge computing unloading decision is located, the multidimensionality of task characteristics, the diversity of base station channels and the multidimensional factors of edge server resource management. According to the method, the convergence speed is improved through an exponential attenuation mechanism of an epsilon-greedy strategy, a Dueling Double DQN architecture is adopted, the problem of overhigh estimation is solved through separation of a value function and a dominant function, the convergence speed continues to be improved, and meanwhile, a dynamic Gamma value adjustment strategy is put forward, so that the task success rate is increased, and the final convergence effect becomes good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing, and specifically to an edge computing offloading decision-making system based on the DQN algorithm. Background Art

[0002] In the new era background of the co-evolution of the information technology revolution and the knowledge economy, the global digital transformation has fully entered the deep development stage of the Internet of Everything. The accelerated development of this process is significantly reflected in two major dimensions: on the one hand, the number of network edge devices has shown an explosive growth trend. According to the Cisco Global Internet Annual Report prediction, the number of mobile device accesses will jump from 8.8 billion in 2018 to 13.1 billion in 2023; on the other hand, the scale of data generated at the edge has exceeded the zettabyte level, forming an exponentially growing data torrent. It is worth noting that the emergence of emerging applications such as industrial Internet, augmented reality, and 8K real-time streaming media has posed a dual challenge to terminal devices - they need to handle computationally intensive tasks and must meet millisecond-level latency requirements, which has significantly exceeded the conventional computing capabilities and energy storage levels of current terminal devices.

[0003] Although the traditional three-layer cloud-edge-end architecture played an important role in the early development of the Internet of Things, its technical limitations have become increasingly prominent in new application scenarios. Currently, terminal devices are restricted by physical size and energy consumption limitations. Even if the performance is improved through hardware iteration, it is still difficult to independently complete complex computing tasks such as artificial intelligence inference and 3D modeling. The "terminal-cloud" binary transmission mode adopted by traditional cloud computing, due to the need to transmit data across a wide area network, not only causes the occupation of network bandwidth and the increase of transmission costs, but also leads to an end-to-end delay reaching the order of hundreds of milliseconds, which poses a fundamental constraint to application scenarios with strict real-time requirements such as autonomous driving. More alarmingly, centralized cloud data centers have become the main targets of network attacks, and the cloud storage mode of sensitive data such as medical and health data and industrial control information has significantly amplified the risk of privacy leakage.

[0004] In the prior art, the traditional cloud computing paradigm has shown systematic deficiencies in coping with the dual challenges of the Internet of Everything era: in the technical dimension, its network architecture is difficult to effectively resolve the transmission bottleneck brought about by the access of a large number of devices; in the economic dimension, the high bandwidth costs generated by remote data transmission restrict the sustainability of business models; in the security dimension, the single-point failure risk existing in the entire data life cycle poses a potential threat to critical information infrastructure.

[0005] Based on this, the present invention provides an edge computing offloading decision-making system based on the DQN algorithm to solve the above-mentioned technical problems. Summary of the Invention

[0006] The purpose of the present invention is to provide an edge computing offloading decision-making system based on the DQN algorithm. The present invention improves the convergence speed through the exponential decay mechanism of the ε-greedy strategy, and adopts the Dueling Double DQN architecture. By separating the value function and the advantage function, the overestimation problem is solved, and the convergence speed is further improved. At the same time, a dynamic Gamma value adjustment strategy is proposed to increase the task success rate and improve the final convergence effect.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The present invention provides an edge computing offloading decision-making system based on the DQN algorithm, including an environment modeling unit, a state space construction unit, an action decision-making unit, a reward feedback unit, a training and optimization unit, and a simulation verification unit, wherein:

[0009] The environment modeling unit: is used to model the heterogeneity of the terminal device, the multi-dimensionality of the task characteristics, the diversity of the base station channels, and the multi-dimensional factors of the edge server resource management in the edge computing offloading decision-making;

[0010] The state space construction unit: is used to construct a state space that can comprehensively reflect the current situation of the edge computing offloading scenario;

[0011] The action decision-making unit: is used to make an offloading decision for the edge computing task based on the DQN algorithm according to the current state space;

[0012] The reward feedback unit: is used to provide corresponding reward feedback for the system according to the execution result of the action decision-making to evaluate the quality of the decision;

[0013] The training and optimization unit: uses the reward feedback information to train and optimize the DQN algorithm;

[0014] The simulation verification unit: conducts a simulation experiment on the edge computing offloading decision-making system based on the DQN algorithm.

[0015] The environment modeling unit includes a terminal device modeling module, a task modeling module, a communication modeling module, and an edge server modeling module, wherein:

[0016] The terminal device modeling module is used to quantify and model the heterogeneous characteristics of the computing power, storage capacity, and battery power of the terminal device;

[0017] The task modeling module is used to model the attributes of the computing amount and data amount of the task in view of the diversity of the task;

[0018] The communication modeling module is used to establish a channel state model considering the time-varying nature of the base station channel, including signal strength, bandwidth, and interference situation;

[0019] The edge server modeling module is used to model the constraint conditions of the computing resources, storage resources, and bandwidth resources of the edge server, and evaluate the server's task processing ability and current load situation.

[0020] The state space construction unit includes a state parameter extraction module, a state vector generation module, and a state normalization module, where:

[0021] The state parameter extraction module is used to extract key parameters related to the edge computing offloading decision from the information provided by the environment modeling unit;

[0022] The state vector generation module is used to integrate the extracted state parameters into a multi-dimensional state vector;

[0023] The state normalization module is used to perform normalization processing on the generated state vector to unify parameters in different ranges to the same scale.

[0024] The state vector contains 16-dimensional parameters, and the specific composition is:

[0025]

[0026] In the formula, d n is the data volume of the task, unit: bit; c n is the computational volume of the task, unit: CPU cycle number; is the local computing ability of the terminal device from which the task comes, unit: CPU cycles per second is the energy consumption per computing cycle of the terminal device from which the task comes; P send,n is the power when the terminal device from which the task comes sends data; indicates whether the transmission power is occupied indicates whether the local computing power is occupied F left represents the available computing resources of the current MECS; Num is the serial number of the terminal device from which the task comes; Time is the current timestamp; matrix O N,M represents the channel occupancy:

[0027] N is the number of base stations, and its row index ranges from 0 to N - 1;

[0028] M is the number of channels of each base station, and its column index ranges from 0 to M - 1;

[0029] Matrix O N,M any element in has two possible values (0 or 1);

[0030] For O N,MFor the element (i, j), if the j-th channel of the i-th base station is occupied, its value is 1;

[0031] Matrix O N,M is flattened into a one-dimensional variable in the state vector x n in;

[0032] It should be noted that P send,n is time-invariant for the same UE but may vary for different UEs. 1(.) is a truth function that takes the value 1 when the event in the parentheses is true and 0 otherwise.

[0033] The action decision unit includes an action space definition module, a DQN network module, and an action selection module, where:

[0034] The action space definition module is used to clarify all possible offloading actions for edge computing tasks;

[0035] The DQN network module is used to implement the core network structure of the DQN algorithm, including an input layer and an output layer, and predicts the Q value of each action by learning the value function of the state-action pair;

[0036] The action selection module is used to select the optimal action in the current state according to the Q value output by the DQN network and use it as the offloading decision for edge computing tasks.

[0037] The Dueling Double DQN algorithm is adopted in the DQN network module, and its action value function is calculated as:

[0038]

[0039] In the formula, Q(s, a; θ) represents the action value function when taking action a in state s, and θ is the parameter of the entire network; V(s; θ V ) is the state value function, representing the expected long-term return that can be obtained in state s, and θ V is the network parameter for calculating the state value function; A(s, a; θ A ) is the action advantage function, measuring the degree of advantage of taking action a in state s compared to the average action, and θ A is the network parameter for calculating the action advantage function; |A| represents the number of actions in the action space A; ∑ a′ A(s, a′; θ A ) sums up the action advantage functions for all possible actions a′ in state s; is the average value of the action advantage function in state s, used to subtract the average advantage from the advantage function of action a to better distinguish the relative advantages between actions;

[0040] Among them, the specific formula for the target Q value is as follows:

[0041]

[0042] In the formula, r is the reward value immediately obtained after the agent executes an action in the current state; γ is the discount factor, and its value range is usually in [0, 1]. Q target is the action-value function calculated by the target network; s′ is the next state transferred to after executing the action; is the online network Q onine is the action that maximizes the Q value calculated in the state s′.

[0043] The expression for updating the dynamic exploration rate of the ε-greedy policy is as follows:

[0044] ε t = max(0.01, 0.99988 t ·ε0

[0045] In the formula, the initial value ε0 = 1, and t is the number of training steps.

[0046] The reward feedback unit includes a reward index definition module, a reward calculation module, and a reward feedback module, where:

[0047] The reward index definition module is used to determine the reward index for evaluating the quality of action decisions, such as task completion time, energy consumption, resource utilization rate, task success rate, and is reasonably set according to specific application scenarios and goals;

[0048] The reward calculation module is used to calculate the corresponding reward value according to the execution result of the action decision and in combination with the defined reward index;

[0049] The reward feedback module is used to feedback the calculated reward value to the training and optimization unit for updating the network parameters of the DQN algorithm.

[0050] The training and optimization unit includes an experience replay module, a target network update module, and a parameter optimization module, where:

[0051] The experience replay module is used to store information such as the state, action, reward, and next state obtained by the agent when executing actions in the environment in the experience replay buffer, and randomly sample samples from the buffer for training;

[0052] The target network update module is used to introduce a target network in order to improve the convergence and stability of the DQN algorithm;

[0053] The parameter optimization module is used to update and optimize the network parameters of the DQN algorithm according to the reward feedback information by using an optimization algorithm, and gradually improve the decision-making performance of the algorithm.

[0054] The simulation verification unit includes a simulation environment construction module, an experimental parameter setting module, and a performance evaluation module, where:

[0055] The simulation environment construction module is used to construct a simulated edge computing offloading scenario, including entities such as terminal devices, base stations, and edge servers, as well as a network communication environment and a task generation mechanism, to simulate the operation of the real world;

[0056] The experimental parameter setting module is used to set the relevant parameters of the simulation experiment;

[0057] The performance evaluation module is used to collect and analyze the performance indicators of the system during the simulation experiment, and verify the effectiveness and superiority of the edge computing offloading decision-making system based on the DQN algorithm by comparing the performance indicators under different algorithms or parameter settings.

[0058] Compared with the prior art, the beneficial effects of the present invention are:

[0059] 1. Through the exponential decay mechanism of the ε-greedy strategy, the present invention improves the convergence speed, and adopts the Dueling Double DQN architecture. By separating the value function and the advantage function, the problem of overestimation is solved, the convergence speed is further improved, and at the same time, a dynamic Gamma value adjustment strategy is proposed to increase the task success rate and improve the final convergence effect.

[0060] 2. By designing a triple state space, a binary action space and a joint reward mechanism of delay and energy consumption, experiments show that the standard DQN algorithm has a 17.6% comprehensive performance improvement compared with the traditional heuristic algorithm in a dynamic environment.

[0061] 3. By constructing an edge computing offloading decision-making system including a task model, a communication model and a computing model, the present invention transforms the task offloading problem in a dynamic environment into a time-series optimization problem through the Markov decision process, laying a mathematical model foundation for the application of reinforcement learning algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a system diagram of the edge computing offloading decision-making system based on the DQN algorithm of the present invention.

[0063] Figure 2 It is an explanatory diagram of the state space parameters of the edge computing offloading decision-making system based on the DQN algorithm of the present invention.

[0064] Figure 3This is the Q neural network structure diagram of the edge computing offloading decision-making system based on the DQN algorithm of the present invention.

[0065] Explanation of the reference numerals in the accompanying drawings:

[0066] 100. Environment modeling unit; 101. Terminal device modeling module; 102. Task modeling module; 103. Communication modeling module; 104. Edge server modeling module; 200. State space construction unit; 201. State parameter extraction module; 202. State vector generation module; 203. State normalization module; 300. Action decision-making unit; 301. Action space definition module; 302. DQN network module; 303. Action selection module; 400. Reward feedback unit; 401. Reward index definition module; 402. Reward calculation module; 403. Reward feedback module; 500. Training and optimization unit; 501. Experience replay module; 502. Target network update module; 503. Parameter optimization module; 600. Simulation verification unit; 601. Simulation environment construction module; 602. Experimental parameter setting module; 603. Performance evaluation module. Specific implementation manners

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0068] Embodiment:

[0069] As Figures 1 - 3 shown, this embodiment provides an edge computing offloading decision-making system based on the DQN algorithm, including an environment modeling unit 100, a state space construction unit 200, an action decision-making unit 300, a reward feedback unit 400, a training and optimization unit 500, and a simulation verification unit 600, wherein:

[0070] The environment modeling unit 100: is used to model the heterogeneity of the terminal device at the edge computing offloading decision-making, the multi-dimensionality of the task characteristics, the diversity of the base station channels, and the multi-dimensional factors of the edge server resource management;

[0071] The state space construction unit 200: is used to construct a state space that can comprehensively reflect the current situation of the edge computing offloading scenario;

[0072] The action decision-making unit 300: is used to make an offloading decision for the edge computing task based on the current state space and the DQN algorithm;

[0073] Reward Feedback Unit 400: It is used to provide corresponding reward feedback for the system according to the execution result of the action decision to evaluate the quality of the decision;

[0074] Training and Optimization Unit 500: It uses the reward feedback information to train and optimize the DQN algorithm;

[0075] Simulation Verification Unit 600: It conducts simulation experiments on the edge computing offloading decision-making system based on the DQN algorithm.

[0076] The Environment Modeling Unit 100 includes a Terminal Device Modeling Module 101, a Task Modeling Module 102, a Communication Modeling Module 103, and an Edge Server Modeling Module 104, where:

[0077] The Terminal Device Modeling Module 101 is used to quantify and model the heterogeneous characteristics of the computing power, storage capacity, and battery power of the terminal device;

[0078] The Task Modeling Module 102 is used to model the attributes of the computing amount and data amount of the task in view of the diversity of the task;

[0079] The Communication Modeling Module 103 is used to establish a channel state model considering the time-varying nature of the base station channel, including signal strength, bandwidth, and interference conditions;

[0080] The Edge Server Modeling Module 104 is used to model the constraint conditions of the computing resources, storage resources, and bandwidth resources of the edge server, and evaluate the server's ability to process tasks and the current load situation.

[0081] The State Space Construction Unit 200 includes a State Parameter Extraction Module 201, a State Vector Generation Module 202, and a State Normalization Module 203, where:

[0082] The State Parameter Extraction Module 201 is used to extract the key parameters related to the edge computing offloading decision from the information provided by the Environment Modeling Unit 100;

[0083] The State Vector Generation Module 202 is used to integrate the extracted state parameters into a multi-dimensional state vector;

[0084] The State Normalization Module 203 is used to perform normalization processing on the generated state vector to unify parameters in different ranges to the same scale.

[0085] The state vector contains 16-dimensional parameters, and the specific composition is:

[0086]

[0087] In the formula, d n is the data volume of the task, unit: bit; c nis the computational workload of the task, in CPU cycles; is the local computing power of the terminal device from which the task comes, in CPU cycles per second is the energy consumption per computing cycle of the terminal device from which the task comes; P send,n is the power when the terminal device from which the task comes sends data; indicates whether the transmission power is occupied indicates whether the local computing power is occupied F left indicates the available computing resources of the current MECS; Num is the serial number of the terminal device from which the task comes; Time is the current timestamp; matrix O N,M indicates the channel occupancy:

[0088] N is the number of base stations, and its row index ranges from 0 to N - 1;

[0089] M is the number of channels of each base station, and its column index ranges from 0 to M - 1;

[0090] matrix O N,M any element in has two possible values (0 or 1);

[0091] For the element (i, j) of O N,M if the j-th channel of the i-th base station is occupied, its value is 1;

[0092] matrix O N,M is flattened into a one-dimensional variable in the state vector x n in;

[0093] It should be noted that P send,n is invariant with time for the same UE, but may be different for different UEs. 1(.) is a truth function that takes the value 1 when the event in the parentheses is true and 0 otherwise.

[0094] The action decision unit 300 includes an action space definition module 301, a DQN network module 302, and an action selection module 303, where:

[0095] The action space definition module 301 is used to clarify the possible offloading actions of the edge computing task;

[0096] The DQN network module 302 is used to implement the core network structure of the DQN algorithm, including an input layer and an output layer, and predicts the Q value of each action by learning the value function of the state-action pair;

[0097] The action selection module 303 is used to select the optimal action in the current state according to the Q value output by the DQN network, as the offloading decision of the edge computing task.

[0098] The Dueling Double DQN algorithm is adopted in the DQN network module 302, and its action value function is calculated as:

[0099]

[0100] In the formula, Q(s,a;θ) represents the action value function when taking action a in state s, and θ is the parameter of the entire network; V(s;θ V ) is the state value function, representing the expected long-term return that can be obtained in state s, and θ V is the network parameter for calculating the state value function; A(s,a;θ A ) is the action advantage function, measuring the degree of advantage of taking action a compared to the average action in state s, and θ A is the network parameter for calculating the action advantage function; |A| represents the number of actions in the action space A; ∑ a′ A(s,a′;θ A ) sums up the action advantage functions of all possible actions a′ in state s; is the average value of the action advantage function in state s, used to subtract the average advantage from the advantage function of action a to better distinguish the relative advantages between actions;

[0101] Among them, the specific formula for the target Q value is:

[0102]

[0103] In the formula, r is the reward value immediately obtained after the agent executes the action in the current state; γ is the discount factor, and its value range is usually in [0,1], Q target is the action value function calculated by the target network; s′ is the next state transferred to after executing the action; is the online network Q onine is the action that maximizes the Q value calculated in state s′.

[0104] The expression for updating the dynamic exploration rate of the ε-greedy strategy is:

[0105] ε t = max(0.01, 0.99988 t ·ε0

[0106] In the formula, the initial value ε0 = 1, and t is the number of training steps.

[0107] The reward feedback unit 400 includes a reward metric definition module 401, a reward calculation module 402, and a reward feedback module 403, where:

[0108] The reward metric definition module 401 is used to determine the reward metrics for evaluating the quality of action decisions, such as task completion time, energy consumption, resource utilization rate, and task success rate, and set them reasonably according to specific application scenarios and goals;

[0109] The reward calculation module 402 is used to calculate the corresponding reward value according to the execution result of the action decision and the defined reward metrics;

[0110] The reward feedback module 403 is used to feedback the calculated reward value to the training and optimization unit 500 for updating the network parameters of the DQN algorithm.

[0111] The training and optimization unit 500 includes a replay buffer module 501, a target network update module 502, and a parameter optimization module 503, where:

[0112] The replay buffer module 501 is used to store information such as the state, action, reward, and next state obtained by the agent when performing actions in the environment in the replay buffer, and select samples from the buffer for training by random sampling;

[0113] The target network update module 502 is used to introduce a target network to improve the convergence and stability of the DQN algorithm;

[0114] The parameter optimization module 503 is used to update and optimize the network parameters of the DQN algorithm using an optimization algorithm according to the reward feedback information, and gradually improve the decision-making performance of the algorithm.

[0115] The simulation verification unit 600 includes a simulation environment construction module 601, an experimental parameter setting module 602, and a performance evaluation module 603, where:

[0116] The simulation environment construction module 601 is used to construct a simulated edge computing offloading scenario, including entities such as terminal devices, base stations, and edge servers, as well as a network communication environment and a task generation mechanism, to simulate the operation of the real world;

[0117] The experimental parameter setting module 602 is used to set the relevant parameters of the simulation experiment;

[0118] The performance evaluation module 603 is used to collect and analyze the performance metrics of the system during the simulation experiment, and verify the effectiveness and superiority of the edge computing offloading decision-making system based on the DQN algorithm by comparing the performance metrics under different algorithms or parameter settings.

[0119] Such as Figures 1 - 3As shown in the figure, this embodiment provides an edge computing offloading decision system based on the DQN algorithm. The specific method is as follows: First, the terminal device modeling module 101 is used to quantify and model the heterogeneous characteristics of the computing power, storage capacity, and battery power of the terminal device. The parameter range of the terminal device is as follows: CPU frequency: 1.0 - 2.5 GHz, battery capacity: 2000 - 5000 mAh. The task modeling module 102 is used to model the attributes of the computing amount and data amount of the task for the diversity of the task. The communication modeling module 103 is used to establish a channel state model considering the time-varying nature of the base station channel, including signal strength, bandwidth, and interference conditions. The edge server modeling module 104 is used to model the constraint conditions of the computing resources, storage resources, and bandwidth resources of the edge server, and evaluate the server's ability to process tasks and the current load situation. The state parameter extraction module 201 is used to extract the key parameters related to the edge computing offloading decision from the information provided by the environment modeling unit 100. The state vector generation module 202 is used to integrate the extracted state parameters into a multi-dimensional state vector. The state vector contains 16-dimensional parameters, and the specific composition is as follows:

[0120]

[0121] In the formula, d n is the data volume of the task, unit: bit; c n is the computing amount of the task, unit: number of CPU cycles; is the local computing power of the terminal device where the task comes from, unit: number of CPU cycles per second is the energy consumption per computing cycle of the terminal device where the task comes from; P send,n is the power when the terminal device where the task comes from sends data; indicates whether the transmission power is occupied indicates whether the local computing power is occupied F left represents the available computing resources of the current MECS; Num is the serial number of the terminal device where the task comes from; Time is the current timestamp; the matrix O N,M represents the channel occupancy situation: N is the number of base stations, and its row index ranges from 0 to N - 1; M is the number of channels of each base station, and its column index ranges from 0 to M - 1; any element in the matrix O N,M has two possible values (0 or 1); for the element (i, j) of O N,M , if the j-th channel of the i-th base station is occupied, its value is 1; the matrix O N,M is flattened into a one-dimensional variable in the state vector x n ; it should be noted that, P send,nIt remains unchanged for the same UE over time, but may be different for different UEs. 1(.) is a truth function that takes the value 1 when the event in the parentheses is true and 0 otherwise. The state normalization module 203 is used to normalize the generated state vector, unifying parameters in different ranges to the same scale. The action space definition module 301 is used to clarify the possible offloading actions for edge computing tasks; the DQN network module 302 is used to implement the core network structure of the DQN algorithm, including the input layer, hidden layer, and output layer. By learning the value function of the state-action pair, it predicts the Q value of each action; the Dueling Double DQN algorithm is adopted, and its action value function is calculated as follows:

[0122]

[0123] In the formula, Q(s,a;θ) represents the action value function when taking action a in state s, and θ is the parameter of the entire network; V(s;θ V ) is the state value function, representing the expected long-term return that can be obtained in state s, and θ V is the network parameter for calculating the state value function; A(s,a;θ A ) is the action advantage function, measuring the degree of advantage of taking action a in state s compared to the average action, and θ A is the network parameter for calculating the action advantage function; |A| represents the number of actions in the action space A; ∑ a′ A(s,a′;θ A ) sums the action advantage functions for all possible actions a′ in state s; is the average value of the action advantage function in state s, used to subtract the average advantage from the advantage function of action a to better distinguish the relative advantages between actions;

[0124] Among them, the specific formula for the target Q value is:

[0125]

[0126] In the formula, r is the reward value immediately obtained after the agent executes the action in the current state; γ is the discount factor, and its value range is usually in [0,1], Q target is the action value function calculated by the target network; s′ is the next state transferred to after executing the action; is the Q of the online network onine is the action that maximizes the Q value calculated in state s′. The action selection module 303 is used to select the optimal action in the current state according to the Q value output by the DQN network, using the ε-greedy strategy as the offloading decision for edge computing tasks. The expression for updating the dynamic exploration rate of the ε-greedy strategy is:

[0127] εt = max(0.01, 0.99988 t ·ε0

[0128] In the formula, the initial value ε0 = 1, and t is the number of training steps..... The reward input state dimension is 16, and the output action value dimension is 8. The network topology process is as follows:

[0129] Input layer: state (batch size 64 × 16)

[0130] Layer1: Gemm (weight matrix bias

[0131] Layer2: Gemm (weight matrix bias

[0132] Layer3: Gemm (weight matrix bias

[0133] Layer4: Gemm (weight matrix bias

[0134] Output layer: q_values (64 × 8)

[0135] Use a not-so-shallow neural network to fit the action value function, and at the same time use LeakyReLU to replace ReLU: By introducing the leakage coefficient α = 0.01, the gradient vanishing problem is alleviated. Previous experiments have shown that LeakyReLU improves the training stability compared to ReLU. The reward module 401 is used to determine the reward metrics for evaluating the quality of action decisions, such as task completion time, energy consumption, resource utilization rate, and task success rate, and is reasonably set according to specific application scenarios and goals; the reward calculation module 402 is used to calculate the corresponding reward value according to the execution results of the action decisions in combination with the defined reward metrics; the reward mechanism is constructed around the following key goals: Task success priority: Ensure the complete execution of the task and avoid system interruption caused by failure; Energy consumption and delay balance: Weigh the local computing energy consumption and edge offloading delay to adapt to dynamic environment constraints; Resource risk avoidance: Prevent base station channel overload or server overload to ensure the sustainable operation of the system; The immediate reward R of the agent at time t t is composed of the following multi-factor linear combination:

[0136]

[0137] Among them: E t and D tThey are the task energy consumption and latency respectively, and the coefficients α and β control their weights; F is the task completion flag, γ provides positive incentives, and a fixed penalty φ is triggered when it fails;

[0138] P i represents resource constraint events such as the base station channel being occupied and the server having no computing power.

[0139] Specifically:

[0140] 1. Task Completion Reward

[0141] When the task is successfully executed, the reward is jointly determined by the linear cost of energy consumption and latency and the completion incentive:

[0142] R fnish = αE + βD + γ

[0143] To avoid training instability caused by too low total reward in extreme scenarios, a lower threshold R min is set. If the calculated value is lower than R min , then the post-guaranteed incentive is given to maintain learning stability.

[0144] 2. Task Failure Penalty

[0145] When the task fails due to insufficient resources or timeout, a high fixed penalty φ is imposed, such as φ = -100. This design forces the agent to prioritize avoiding losses and thus improves system reliability.

[0146] 3. Operation Cost Penalty

[0147] Transmission cost: The communication overhead penalty when the task is offloaded, which inhibits the abuse of network bandwidth.

[0148] Local computing cost: The implicit cost of device CPU occupancy. This model assumes that local computing can only perform the calculation of one task simultaneously in a single thread. This results in limited local computing resources and requires the agent to feel it during training.

[0149] Server no-computing-power penalty: Guides the agent to perceive the load status of edge nodes and dynamically adjust the offloading strategy.

[0150] Base station channel occupied penalty: Guides the agent to perceive the busy status of the base station channel and dynamically utilize different base stations and their channels for uploading.

[0151] The reward feedback module 403 is used to feedback the calculated reward value to the training and optimization unit 500 for updating the network parameters of the DQN algorithm. The experience replay module 501 is used to store information such as the state, action, reward, and next state obtained by the agent when executing actions in the environment in the experience replay buffer, and samples are selected from the buffer for training by random sampling; the target network update module 502 is used to introduce a target network in order to improve the convergence and stability of the DQN algorithm; the parameter optimization module 503 is used to adopt an optimization algorithm to update and optimize the network parameters of the DQN algorithm according to the reward feedback information, and gradually improve the decision-making performance of the algorithm. The simulation environment construction module 601 is used to construct a simulated edge computing offloading scenario, including entities such as terminal devices, base stations, and edge servers, as well as a network communication environment and a task generation mechanism, to simulate the operation of the real world; the experimental parameter setting module 602 is used to set the relevant parameters of the simulation experiment; the relevant parameters include the task arrival rate, the number of terminal devices, and the resource configuration of the edge server, so as to evaluate the system performance under different scenarios. The performance evaluation module 603 is used to collect and analyze the performance indicators of the system during the simulation experiment, and verify the effectiveness and superiority of the edge computing offloading decision-making system based on the DQN algorithm by comparing the performance indicators under different algorithms or parameter settings. The performance indicators include the task completion time, energy consumption, resource utilization rate, and task success rate.

[0152] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0153] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art in the relevant technical field can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. An edge computing offloading decision system based on the DQN algorithm, characterized in that It includes an environment modeling unit (100), a state space construction unit (200), an action decision-making unit (300), a reward feedback unit (400), a training and optimization unit (500), and a simulation verification unit (600), where: The environment modeling unit (100): is used to model the heterogeneity of the terminal device where the edge computing offloading decision is located, the multi-dimensionality of task characteristics, the diversity of base station channels, and the multi-dimensional factors of edge server resource management; The state space construction unit (200): is used to construct a state space that can comprehensively reflect the current situation of the edge computing offloading scenario; The action decision-making unit (300): is used to make an offloading decision for edge computing tasks based on the current state space and the DQN algorithm; The reward feedback unit (400): is used to provide corresponding reward feedback to the system according to the execution result of the action decision to evaluate the quality of the decision; The training and optimization unit (500): uses the reward feedback information to train and optimize the DQN algorithm; The simulation verification unit (600): conducts simulation experiments on the edge computing offloading decision system based on the DQN algorithm.

2. The edge computing offloading decision system based on the DQN algorithm according to claim 1, wherein The environment modeling unit (100) includes a terminal device modeling module (101), a task modeling module (102), a communication modeling module (103), and an edge server modeling module (104), where: The terminal device modeling module (101) is used to quantify and model the heterogeneous characteristics of the computing power, storage capacity, and battery power of the terminal device; The task modeling module (102) is used to model the attributes of the computing amount and data amount of the task in view of the diversity of tasks; The communication modeling module (103) is used to establish a channel state model considering the time-varying nature of the base station channel, including signal strength, bandwidth, and interference situation; The edge server modeling module (104) is used to model the constraint conditions of the computing resources, storage resources, and bandwidth resources of the edge server, and evaluate the server's ability to process tasks and the current load situation.

3. The edge computing offloading decision system based on the DQN algorithm according to claim 1, characterized in that, The state space construction unit (200) includes a state parameter extraction module (201), a state vector generation module (202), and a state normalization module (203), where: The state parameter extraction module (201) is used to extract key parameters related to the edge computing offloading decision from the information provided by the environment modeling unit (100); The state vector generation module (202) is used to integrate the extracted state parameters into a multi-dimensional state vector; The state normalization module (203) is used to perform normalization processing on the generated state vector to unify parameters in different ranges to the same scale.

4. The edge computing offloading decision system based on the DQN algorithm according to claim 1, characterized in that, The state vector contains 16-dimensional parameters, and the specific composition is: where d n is the data volume of the task, unit: bit; c n is the computational volume of the task, unit: CPU cycle number; is the local computing power of the terminal device from which the task comes, unit: CPU cycles per second is the energy consumption per computing cycle of the terminal device from which the task comes; P send,n is the power when the terminal device from which the task comes sends data; indicates whether the transmission power is occupied indicates whether the local computing power is occupied F left represents the available computing resources of the current MECS; Num is the serial number of the terminal device from which the task comes; Time is the current timestamp; matrix O N,M indicates the channel occupancy: N is the number of base stations, and its row index ranges from 0 to N - 1; M is the number of channels of each base station, and its column index ranges from 0 to M - 1; Matrix O N,M Any element in it has two possible values (0 or 1); For O N,M For the element (i, j), if the j-th channel of the i-th base station is occupied, its value is 1; Matrix O N,M is flattened into a one-dimensional variable in n the state vector x; It should be noted that It is constant over time for the same UE, but may vary for different UEs. 1(.) is a truth function that takes the value 1 when the event in the parentheses is true and 0 otherwise.

5. The edge computing offloading decision system based on the DQN algorithm according to claim 4, characterized in that The action decision-making unit (300) includes an action space definition module (301), a DQN network module (302), and an action selection module (303), where: The action space definition module (301) is used to clarify all possible offloading actions for edge computing tasks; The DQN network module (302) is used to implement the core network structure of the DQN algorithm, including an input layer and an output layer, and predicts the Q-value of each action by learning the value function of the state-action pair; The action selection module (303) is used to select the optimal action in the current state according to the Q-value output by the DQN network, using the ε-greedy strategy as the offloading decision for the edge computing task.

6. The edge computing offloading decision system based on the DQN algorithm according to claim 5, wherein The Dueling Double DQN algorithm is adopted in the DQN network module (302), and its action value function is calculated as: Where, Q(s,a;θ) represents the action value function when taking action a in state s, and θ is the parameter of the entire network; V(s;θ V ) is the state value function, representing the expected long-term return that can be obtained in state s, and θ V is the network parameter for calculating the state value function; A(s,a;θ A ) is the action advantage function, measuring the degree of advantage of taking action a in state s compared to the average action, and θ A is the network parameter for calculating the action advantage function; |A| represents the number of actions in the action space A; ∑ a′ A(s,a′;θ A ) sums the action advantage functions for all possible actions a′ in state s; is the average value of the action advantage function in state s, used to subtract the average advantage from the advantage function of action a to better distinguish the relative advantages between actions; Among them, the specific formula for the target Q-value is: where r is the reward value immediately obtained by the intelligent agent after executing an action in the current state; γ is the discount factor, and its value range is usually in [0, 1], Q target is the action value function calculated by the target network; s′ is the next state transferred to after executing the action; is the online network Q onine is the action that maximizes the Q value calculated in the state s′.

7. The edge computing offloading decision system based on the DQN algorithm according to claim 5, characterized in that, The expression for updating the dynamic exploration rate of the ε-greedy strategy is: ε t = max(0.01, 0.99988 t ·ε0 In the formula, the initial value ε0 = 1, and t is the number of training steps.

8. The edge computing offloading decision-making system based on the DQN algorithm according to claim 1, characterized in that, The reward feedback unit (400) includes a reward metric definition module (401), a reward calculation module (402), and a reward feedback module (403), where: The reward metric definition module (401) is used to determine the reward metrics for evaluating the quality of action decisions, such as task completion time, energy consumption, resource utilization rate, task success rate, and is reasonably set according to specific application scenarios and objectives; The reward calculation module (402) is used to calculate the corresponding reward value according to the execution result of the action decision, in combination with the defined reward metrics; The reward feedback module (403) is used to feedback the calculated reward value to the training and optimization unit (500) for updating the network parameters of the DQN algorithm.

9. The edge computing offloading decision system based on the DQN algorithm according to claim 1, wherein The training and optimization unit (500) includes an experience replay module (501), a target network update module (502), and a parameter optimization module (503), where: The experience replay module (501) is used to store information such as the state, action, reward, and next state obtained by the agent when performing actions in the environment in the experience replay buffer, and randomly samples samples from the buffer for training; The target network update module (502) is used to introduce a target network in order to improve the convergence and stability of the DQN algorithm; The parameter optimization module (503) is used to adopt an optimization algorithm to update and optimize the network parameters of the DQN algorithm according to the reward feedback information, and gradually improve the decision-making performance of the algorithm.

10. The edge computing offloading decision-making system based on the DQN algorithm according to claim 1, characterized in that, The simulation verification unit (600) includes a simulation environment construction module (601), an experimental parameter setting module (602), and a performance evaluation module (603), where: The simulation environment construction module (601) is used to construct a simulated edge computing offloading scenario, including entities such as terminal devices, base stations, and edge servers, as well as a network communication environment and a task generation mechanism, to simulate the operation of the real world; The experimental parameter setting module (602) is used to set the relevant parameters of the simulation experiment; The performance evaluation module (603) is used to collect and analyze the performance metrics of the system during the simulation experiment, and verify the effectiveness and superiority of the edge computing offloading decision system based on the DQN algorithm by comparing the performance metrics under different algorithms or parameter settings.

Citation Information

Cited By

  • Human-machine cooperation assembly unit task intelligent scheduling method considering fatigue recovery

    CN121436592A