A task offloading method and system based on intelligent reflection surface node selection

By building an edge computing network with intelligent reflector node selection and utilizing signal-to-noise ratio optimization phase and deep reinforcement learning algorithms, the problem of insufficient flexibility in node selection and task offloading strategies in traditional edge computing systems is solved, and optimal decision-making for task offloading and balanced allocation of resources are achieved, thereby improving system efficiency and resource utilization.

CN117221322BActive Publication Date: 2025-09-26GUANGDONG POWER GRID CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311295684.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-08
Publication Date
2025-09-26
Estimated Expiration
2043-10-08

AI Technical Summary

Technical Problem

In traditional edge computing IoT systems, node selection and task offloading strategies lack flexibility and adaptability, resulting in the inability to optimize node selection and task offloading when the network environment changes dynamically, causing unbalanced task distribution, affecting system efficiency and resource utilization, and preventing high-priority tasks from responding in a timely manner. In addition, computing resource allocation does not fully consider task characteristics.

Method used

By constructing an edge computing network based on intelligent reflector node selection, using the signal-to-noise ratio to optimize the phase of the intelligent reflector node, and combining the deep reinforcement learning algorithm for resource allocation, the task offloading rate and the edge server computing power allocation rate are optimized, and the decomposition method is used to reduce the computational complexity and improve resource utilization.

Benefits of technology

It significantly improves the signal quality of edge computing networks, enhances data transmission efficiency, achieves optimal decision-making for task offloading, and improves resource utilization and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117221322B_ABST
    Figure CN117221322B_ABST
Patent Text Reader

Abstract

The present invention discloses a task offloading method and system based on intelligent reflector surface node selection. The method includes: constructing an edge computing network according to users to perform task offloading, preset edge servers and intelligent reflector surface nodes, calculating the signal-to-noise ratio of users performing task offloading through the intelligent reflector surface nodes, obtaining a first phase corresponding to the intelligent reflector surface node by maximizing the signal-to-noise ratio, allocating corresponding intelligent reflector surface nodes to users according to task priority, obtaining a task offloading rate and an edge server computing power allocation rate corresponding to each intelligent reflector surface node through a deep reinforcement learning algorithm, respectively allocating the user's offloading tasks to corresponding intelligent reflector surface nodes, and offloading the offloading tasks to the edge servers for calculation according to the first phase of the intelligent reflector surface node, the task offloading rate and the edge server computing power allocation rate, thereby improving the efficiency of data transmission and calculation and achieving the optimal decision for task offloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mobile communication technology, and in particular to a task offloading method and system based on intelligent reflection surface node selection. Background Art

[0002] With the rapid development of the Internet of Things (IoT), the widespread adoption of cloud computing, and the ubiquity of social media, global data volume is experiencing a dramatic surge in growth. This growth is driven by the surging demand for multimedia content, such as online gaming, high-definition video, and virtual reality. This has posed a significant challenge to network bandwidth. Against this backdrop, Mobile Edge Computing (MEC) has emerged. By deploying computing resources at the edge of the network, MEC significantly reduces communication latency while also alleviating the burden on traditional centralized data centers, enabling efficient delivery of computing services to mobile users. However, despite MEC's ​​significant advantages in improving network performance and quality, problems such as network congestion and unstable communications persist, leading to reduced data transmission and computing efficiency.

[0003] At the same time, traditional edge computing IoT systems typically use fixed node selection and task offloading strategies, lacking flexibility and adaptability. This makes it difficult for the system to adapt to dynamic changes in the network environment and the requirements of different tasks. It is impossible to optimize node selection and task offloading based on real-time channel status and task characteristics, thus limiting system performance. Furthermore, existing technologies rarely consider the proportional distribution of task offloading across different nodes. This results in some nodes carrying an excessive task load while others are idle. This unbalanced task distribution affects overall system efficiency and reduces resource utilization. Thirdly, in traditional edge computing systems, the allocation of computing resources to edge servers (ES) often fails to fully consider the characteristics and priorities of different tasks. This results in high-priority tasks not receiving timely responses or low-priority tasks occupying excessive computing resources, impacting overall system performance. Finally, traditional methods fail to fully consider complex system dynamics and task characteristics, making it impossible to achieve optimal decisions. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention discloses a task offloading method and system based on intelligent reflection surface node selection, which improves the efficiency of data transmission and calculation and realizes the optimal decision of task offloading.

[0005] To achieve the above objectives, in a first aspect, the present invention discloses a task offloading method based on intelligent reflective surface node selection, comprising:

[0006] An edge computing network is constructed based on a number of users to be task-offloaded, a preset edge server, and a number of intelligent reflective surface nodes;

[0007] Calculating, based on the channel parameter information within the edge computing network, a signal-to-noise ratio of the plurality of users performing task offloading through the plurality of smart reflective surface nodes;

[0008] Optimizing the phase of each of the plurality of smart reflective surface nodes by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each of the smart reflective surface nodes;

[0009] Obtaining a task priority corresponding to each of the plurality of users, and assigning a corresponding smart reflective surface node to each of the users according to the task priority;

[0010] According to the task offloading information corresponding to the user assigned to each smart reflective surface node and the parameter information of the edge computing network, a preset deep reinforcement learning algorithm is used to obtain the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node;

[0011] The offloading tasks of each of the plurality of users are respectively assigned to the corresponding intelligent reflecting surface node, and the offloading tasks are offloaded to the edge server for calculation according to the first phase, task offloading rate and edge server computing power allocation rate corresponding to each intelligent reflecting surface node.

[0012] The present invention discloses a task offloading method based on intelligent reflector node selection. The method first constructs a corresponding edge computing network based on a number of users corresponding to the tasks to be offloaded, a number of preset intelligent reflector nodes for task offloading, and a preset edge server for task computing, so as to perform task offloading. After the edge computing network is constructed, the signal-to-noise ratio of the edge computing network is calculated in real time according to the channel conditions of the edge computing network, so as to optimize and adjust the phases of the intelligent reflector nodes according to the signal-to-noise ratio to obtain the optimal phase corresponding to each intelligent reflector node, thereby significantly improving the signal quality of the edge computing network and further improving the efficiency of data transmission. After obtaining the optimal phase, corresponding nodes are allocated to the users according to their different task priorities, so as to effectively perform resource allocation during task offloading. During resource allocation, the task offloading rate corresponding to each node in the edge computing network and the edge server computing power allocation rate are optimized based on a preset reinforcement learning algorithm. The reinforcement learning algorithm enables the multiple intelligent reflector nodes to obtain balanced resource allocation, improves resource utilization, and achieves the optimal decision for task offloading.

[0013] As a preferred example, calculating, based on the channel parameter information in the edge computing network, a signal-to-noise ratio of the plurality of users performing task offloading through the plurality of intelligent reflective surface nodes includes:

[0014] Calculating a signal-to-noise ratio of each user for task offloading through the smart reflecting surface node based on phase reflection matrices corresponding to the plurality of smart reflecting surface nodes and allocation matrices corresponding to the plurality of users and the plurality of smart reflecting surface nodes;

[0015] Wherein, the phase reflection matrix is:

[0016] in, Indicates the phase of the smart reflector node that needs to be optimized; the K m represents the number of reflective elements equipped in each smart reflective surface node; m represents the number of smart reflective surface nodes; j represents an imaginary unit;

[0017] Wherein, the allocation matrix is:

[0018]

[0019] Wherein, n represents the number of users;

[0020] The calculation formula of the signal-to-noise ratio is:

[0021]

[0022] Among them, P n Represents user equipment S n The transmission power; σ 2 represents the additive white noise variance; in addition, h n,m and g n,d Represents user equipment S n To the intelligent reflector node R m and wireless channel between edge servers; h m Represents the intelligent reflector node R m Wireless channel to edge server.

[0023] The present invention calculates the signal-to-noise ratio so that the phase of the intelligent reflector node is obtained by subsequently optimizing the signal-to-noise ratio, thereby improving the communication quality of the edge computing network.

[0024] As a preferred example, optimizing the phase of each of the plurality of smart reflector nodes by maximizing the signal-to-noise ratio to obtain the phase corresponding to each smart reflector node includes:

[0025] According to the calculation formula of the signal-to-noise ratio, the transmission rate of the user equipment corresponding to each user is calculated; wherein the calculation formula of the transmission rate is: P n =Blog2(1+ζn); where B is the wireless bandwidth of the user equipment;

[0026] The transmission delay of the user equipment is calculated according to the transmission rate; the calculation formula of the transmission delay is: Among them, the u n Represents user equipment S n The task size is in Mbits. n Represents user equipment S n Uninstall rate;

[0027] According to the transmission delay, the user equipment S is calculated n The calculation delay at the edge server is calculated as follows: Among them, the is the computing power of the edge server, b n Represents the data allocated to user device S in the edge server n The computing power of ξ represents the number of CPU cycles required to process one bit of task;

[0028] The user equipment S is calculated according to the preset local delay calculation formula n The delay when processing tasks locally; the local delay is calculated as: Among them, the Indicates user equipment S n computing power;

[0029] An offload delay of the user equipment is obtained according to the transmission delay and the calculation delay; the offload delay is calculated as follows:

[0030] Based on the fact that the offloading process and the local calculation are performed in parallel, the task delay of the user equipment when performing task offloading is the maximum value of the offloading delay and the local delay; the calculation formula of the task delay is:

[0031] The total utility of the edge computing network is calculated based on the task delay and the task priority. The calculation formula of the total utility is:

[0032]

[0033] Among them, the U n For user equipment S n The utility is:

[0034]

[0035] Among them, the The utility corresponding to high priority tasks:

[0036]

[0037] Among them, the q n The priority of the task, represents the completion delay threshold of high priority tasks, Γ H represents a negative reward; said I(x) represents a 0-1 function;

[0038] described The utility corresponding to low-priority tasks:

[0039]

[0040] Among them, represents the completion delay threshold of low priority tasks, Γ L represents a negative reward, and α represents an exponential decay factor.

[0041] The present invention optimizes the phase of the intelligent reflector node by maximizing the signal-to-noise ratio. Specifically, the parameters to be optimized are obtained according to the calculation formula of the signal-to-noise ratio. Then, based on the parameters to be optimized, a series of preset optimization functions are used to finally obtain the utility expression of the system. Finally, the present invention obtains the optimal phase by optimizing the utility of the system.

[0042] As a preferred example, the step of obtaining the first phase corresponding to each smart reflector node further includes:

[0043] According to the calculation formula of the total utility, the calculation formula of the total utility is maximized through the preset constraints; wherein the constraints include: constraint C1 represents when l n,m =1, user equipment S n Through the only reflective surface node R m Offload tasks to edge servers, otherwise l n,m = 0; Constraint C2 means that a smart reflector node can only provide services to one user device at most; Constraint C3 means the phase shift of the smart reflector node belongs to the range of [0, 2π); constraints C4 and C5 correspond to the task offloading rate and the computing power allocation rate of the edge server respectively; constraint C6 represents that the computing power allocation ratio does not exceed the total computing power of the edge server; the optimization function of the calculation formula of the total utility determined according to the constraints is:

[0044]

[0045] According to the optimization function, the optimal phase shift configuration of the smart reflector node is optimized. The phase shift optimization objective function is:

[0046]

[0047] According to the triangle inequality, we can deduce that:

[0048] |g n,d +h m Θ k h n,m |≤|g n,d |+|h m Θ k h n,m |,

[0049] When arg(g n,d )=arg(h m Θ k h m,k ), assuming x m y m =h m Θ m h n,m ,at this time

[0050]

[0051]

[0052] Then the phase shift optimization objective function becomes:

[0053]

[0054] Therefore, we can get

[0055]

[0056] Therefore, the intelligent reflector node R m Corresponding user equipment S n The first phase is:

[0057]

[0058] in and They are h m and h n,m The kth element of ;

[0059] The intelligent reflective surface node R m The optimal phase shift matrix is ​​Φ * ={θ1 * ,θ2 *,…,θ M *} T ,in

[0060] The present invention can see from the optimization function corresponding to the calculation formula of total utility that directly solving the problem has a high complexity. In addition, it is observed that the signal-to-noise ratio during task offloading only involves the phase shift matrix Φ and the allocation matrix of the smart reflector node, and these two variables are independent of other variables. Therefore, a decomposition method is used to decompose the problem into two sub-problems, reducing the data processing volume and improving processing efficiency while maintaining the optimality of the solution.

[0061] As a preferred example, obtaining the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node by using a preset deep reinforcement learning algorithm includes:

[0062] Obtain parameter information of the edge computing network, establish a behavior network, and randomly initialize network parameters of the behavior network; the behavior network is a policy network used to derive an offloading policy;

[0063] Establishing a first main penalty network and a second main penalty network, and randomly initializing the network parameters of the first main penalty network and the second main penalty network; the first main penalty network and the second main penalty network are used to evaluate the action value function of the current strategy;

[0064] Establishing a first target neural network and a second target neural network and initializing network parameters of the first target neural network and the second target neural network; the target neural networks are used to update the parameters of the main neural network;

[0065] Initialize the experience pool, the current state of the environment, the current number of iterations, and the upper limit of the total number of iterations; the experience pool is used to store training samples;

[0066] Obtain an action based on the current policy and state, calculate the total utility of the edge computing network based on the action, and calculate the reward and next state brought about by this state change;

[0067] Storing the current state, the action based on the current strategy and state, the reward, and the next state in the experience pool;

[0068] Randomly extract a small batch of samples from the experience pool, calculate the loss function of the first main penalty network and the second main penalty network respectively, and update the network parameters of the first main penalty network and the second main penalty network by gradient descent method;

[0069] Use the loss function to update the current behavior network, and after several rounds of iterations, copy the parameters of the main neural network to the parameters of the target neural network;

[0070] After iterative training for the total number of iterations, the task offloading rate and the edge server computing power allocation rate are obtained.

[0071] In order to ensure the selection of the optimal intelligent reflector node while maximizing the effectiveness of the edge computing network, the present invention uses a reinforcement learning algorithm to make strategic decisions, improve resource utilization, and obtain the optimal decision for task offloading.

[0072] In a second aspect, the present invention discloses a task offloading system based on intelligent reflector node selection, the system comprising a network construction module, a signal-to-noise ratio module, a phase optimization module, a node allocation module, an offloading decision module, and an offloading calculation module;

[0073] The network building module is used to build an edge computing network based on a number of users to be task-offloaded, a preset edge server, and a number of intelligent reflective surface nodes;

[0074] The signal-to-noise ratio module is used to calculate the signal-to-noise ratio of the multiple users performing task offloading through the multiple smart reflective surface nodes based on the channel parameter information in the edge computing network;

[0075] The phase optimization module is configured to optimize the phase of each of the plurality of smart reflector nodes by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each of the smart reflector nodes;

[0076] The node allocation module is used to obtain the task priority corresponding to each of the plurality of users, and allocate a corresponding smart reflective surface node to each user according to the task priority;

[0077] The offloading decision module is used to obtain the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node through a preset deep reinforcement learning algorithm based on the task offloading information corresponding to the user assigned to each smart reflective surface node and the parameter information of the edge computing network;

[0078] The offloading calculation module is used to allocate the offloading tasks of each of the multiple users to the corresponding smart reflecting surface node, and offload the offloading tasks to the edge server for calculation according to the first phase, task offloading rate and edge server computing power allocation rate corresponding to each smart reflecting surface node.

[0079] The present invention discloses a task offloading system based on intelligent reflector node selection. The system first constructs a corresponding edge computing network based on a number of users corresponding to tasks to be offloaded, a number of preset intelligent reflector nodes for task offloading, and a preset edge server for task computing, so as to perform task offloading. After the edge computing network is constructed, the signal-to-noise ratio of the edge computing network is calculated in real time according to the channel conditions of the edge computing network, so as to optimize and adjust the phase of the intelligent reflector node according to the signal-to-noise ratio to obtain the optimal phase corresponding to each intelligent reflector node, thereby significantly improving the signal quality of the edge computing network and further improving the efficiency of data transmission. At the same time, after obtaining the optimal phase, the corresponding node is allocated to the user according to the different task priorities, so as to effectively perform resource allocation during task offloading. When performing resource allocation, the task offloading rate corresponding to each node in the edge computing network and the edge server computing power allocation rate are optimized based on a preset reinforcement learning algorithm. The reinforcement learning algorithm enables the plurality of intelligent reflector nodes to obtain balanced resource allocation, improves resource utilization, and achieves the optimal decision for task offloading.

[0080] As a preferred example, the signal-to-noise ratio module includes:

[0081] Calculating a signal-to-noise ratio of each user for task offloading through the smart reflecting surface node based on phase reflection matrices corresponding to the plurality of smart reflecting surface nodes and allocation matrices corresponding to the plurality of users and the plurality of smart reflecting surface nodes;

[0082] Wherein, the phase reflection matrix is:

[0083] in, Indicates the phase of the smart reflector node that needs to be optimized; the K m represents the number of reflective elements equipped in each smart reflective surface node; m represents the number of smart reflective surface nodes; j represents an imaginary unit;

[0084] Wherein, the allocation matrix is:

[0085]

[0086] Wherein, n represents the number of users;

[0087] The calculation formula of the signal-to-noise ratio is:

[0088]

[0089] Among them, P n Represents user equipment S n The transmission power; σ2 represents the additive white noise variance; in addition, h n,m and g n,d Represents user equipment S n To the intelligent reflector node R m and wireless channel between edge servers; h m Represents the intelligent reflector node R m Wireless channel to edge server.

[0090] The present invention calculates the signal-to-noise ratio so as to obtain the optimal phase of the intelligent reflector node by subsequently optimizing the signal-to-noise ratio, thereby improving the communication quality of the edge computing network.

[0091] As a preferred example, the uninstallation decision module includes a network construction unit, an initialization unit and an update iteration unit;

[0092] The network construction unit is used to obtain parameter information of the edge computing network, establish a behavior network, and randomly initialize the network parameters of the behavior network; the behavior network is a policy network, which is used to derive an unloading policy; establish a first main penalty network and a second main penalty network, and randomly initialize the network parameters of the first main penalty network and the second main penalty network; the first main penalty network and the second main penalty network are used to evaluate the action value function of the current policy; establish a first target neural network and a second target neural network and initialize the network parameters of the first target neural network and the second target neural network; the target neural network is used to update the parameters of the main neural network;

[0093] The initialization unit is used to initialize the experience pool, the current state of the environment, the current number of iterations, and the upper limit of the total number of iterations; the experience pool is used to store training samples; an action based on the current strategy and state is obtained, the total utility of the edge computing network is calculated based on the action, and the reward and next state brought about by the state change are calculated;

[0094] The update iteration unit is used to store the current state, the action based on the current strategy and state, the reward and the next state in the experience pool; randomly extract a small batch of samples from the experience pool, calculate the loss function of the first main penalty network and the second main penalty network respectively, and update the network parameters of the first main penalty network and the second main penalty network by gradient descent method; use the loss function to update the current behavior network, and after several rounds of iterations, copy the parameters of the main neural network to the target neural network parameters; after iterative training for the total number of iterations, obtain the task offloading rate and the edge server computing power allocation rate.

[0095] In order to ensure the selection of the optimal intelligent reflector node while maximizing the effectiveness of the edge computing network, the present invention uses a reinforcement learning algorithm to make strategic decisions, improve resource utilization, and obtain the optimal decision for task offloading.

[0096] In a third aspect, the present invention discloses an electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement a task offloading method based on intelligent reflection surface node selection as described in the first aspect when executing the program stored in the memory.

[0097] In a fourth aspect, the present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the task offloading method based on intelligent reflection surface node selection as described in the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] Figure 1 : A flowchart of a task offloading method based on intelligent reflective surface node selection disclosed in an embodiment of the present invention;

[0099] Figure 2 : A structural diagram of a task offloading system based on intelligent reflective surface node selection disclosed in an embodiment of the present invention;

[0100] Figure 3 : A schematic diagram of the architecture of a multi-user multi-intelligent reflector node selection system disclosed in an embodiment of the present invention;

[0101] Figure 4 :A graph showing the convergence of algorithms of different schemes in a Python simulation environment disclosed in an embodiment of the present invention;

[0102] Figure 5 : A schematic diagram of the system effects corresponding to different solutions under different ES computing capabilities in a Python simulation environment disclosed in an embodiment of the present invention;

[0103] Figure 6 : A schematic diagram of the system effects corresponding to different schemes under different user bandwidths in a Python simulation environment disclosed in an embodiment of the present invention;

[0104] Figure 7 : A schematic diagram of the system effects corresponding to different schemes under different delay thresholds of high-priority tasks in a Python simulation environment disclosed in an embodiment of the present invention;

[0105] Figure 8: This is a schematic diagram of the system effects corresponding to different schemes under different numbers of users in a Python simulation environment disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0106] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0107] Example 1

[0108] The embodiment of the present invention provides a task offloading method based on intelligent reflective surface node selection. For the specific implementation process of the offloading method, please refer to Figure 1 , mainly including steps 101 to 106, the steps including:

[0109] Step 101: Construct an edge computing network based on a number of users to be task-offloaded, a preset edge server, and a number of intelligent reflective surface nodes.

[0110] In this embodiment, this step includes: establishing an edge computing network according to a number of users who need to offload tasks, a number of intelligent reflective surface nodes, and an edge server.

[0111] Specifically, in this embodiment, a MEC network is established based on N users, M RIS nodes and an edge server. The MEC network is an edge computing network, the RIS node is an intelligent reflection surface node, and the user set can be expressed as S={S n |1≤n≤N}, the RIS set can be represented by R={R m |1≤m≤M}; N users offload tasks to ES through M RIS for computing, where ES is an edge server.

[0112] Step 102: Calculate the signal-to-noise ratio of the tasks unloaded by the plurality of users through the plurality of intelligent reflective surface nodes based on the channel parameter information within the edge computing network. The plurality of users unload the tasks to the edge server through the plurality of intelligent reflective surface nodes for calculation.

[0113] In this embodiment, this step includes: calculating a signal-to-noise ratio (SNR) of each user performing task offloading through the smart reflecting surface node based on phase reflection matrices corresponding to the plurality of smart reflecting surface nodes and allocation matrices corresponding to the plurality of users and the plurality of smart reflecting surface nodes;

[0114] Wherein, the phase reflection matrix is:

[0115] in, Indicates the phase of the smart reflector node that needs to be optimized; the K m represents the number of reflective elements equipped in each smart reflective surface node; m represents the number of smart reflective surface nodes; j represents an imaginary unit;

[0116] Wherein, the allocation matrix is:

[0117]

[0118] Wherein, n represents the number of users;

[0119] The calculation formula of the signal-to-noise ratio is:

[0120]

[0121] Among them, P n Represents user equipment S n The transmission power; σ 2 represents the additive white noise variance; in addition, h n,m and g n,d Represents user equipment S n To the intelligent reflector node R m and wireless channel between edge servers; h m Represents the intelligent reflector node R m Wireless channel to edge server.

[0122] Specifically, in this embodiment, the channel parameter information in the MEC network is obtained, and the signal-to-noise ratio (SNR) of all users unloading through all RIS is calculated, and the phase of RIS is optimized by maximizing the SNR. Specifically, each RIS is equipped with K m By properly adjusting the phase of the RIS, the user's signal can be reflected to the ES. The amplitude reflection coefficient of all reflection units is equal to 1, and the phase reflection matrix is ​​expressed as:

[0123]

[0124] Among them, the represents the scattering coefficient of m RIS units; j is an imaginary unit; where Indicates the RIS phase that needs to be optimized.

[0125] In addition, the allocation matrix of users and RIS is expressed as:

[0126] where l n,m ∈{0,1},l n,m =1 indicates user S n By RIS R m Offload the task to ES, otherwise l n,m = 0. We assume that each user is matched with a RIS. User S n The SNR at ES is given by

[0127]

[0128] The user equipment S n The transmission power is P n ,σ 2 is the additive white noise variance. In addition, from user S n To RISR m The wireless channels of the edge server (ES) are respectively and g n,d In addition, from RIS R m The wireless channel to ES is composed of express.

[0129] In this embodiment, this step calculates the signal-to-noise ratio so that the optimal phase of the smart reflective surface node is obtained by subsequently optimizing the signal-to-noise ratio, thereby improving the communication quality of the edge computing network.

[0130] Step 103: Optimizing the phase of each of the plurality of smart reflective surface nodes by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each of the smart reflective surface nodes.

[0131] In this embodiment, this step includes: calculating the transmission rate of the user equipment according to the calculation formula of the signal-to-noise ratio; wherein the calculation formula of the transmission rate is: P n =Blog2(1+ζn); wherein B is the wireless bandwidth of the user equipment; according to the transmission rate, the transmission delay of the user equipment is calculated; the calculation formula of the transmission delay is: Among them, the u n Represents user equipment S n The task size is in Mbits. n Represents user equipment S n Unloading rate; Calculate the user equipment S according to the transmission delay n The calculation delay at the edge server is calculated as follows: Among them, the is the computing power of the edge server, b n Represents the data allocated to user device S in the edge server n The computing power of ξ represents the number of CPU cycles required to process one bit of task;

[0132] The user equipment S is calculated according to the preset local delay calculation formula n The delay when processing tasks locally; the local delay is calculated as: Among them, the Indicates user equipment S n computing capability; obtaining an offload delay of the user equipment according to the transmission delay and the calculation delay; the calculation formula of the offload delay is:

[0133] Based on the fact that the offloading process and the local calculation are performed in parallel, the task delay of the user equipment when performing task offloading is the maximum value of the offloading delay and the local delay; the calculation formula of the task delay is: The total utility of the edge computing network is calculated based on the task delay and the task priority. The calculation formula of the total utility is:

[0134]

[0135] Among them, the U n For user equipment S n The utility is:

[0136]

[0137] Among them, the The utility corresponding to high priority tasks:

[0138]

[0139] Among them, the q n The priority of the task, represents the completion delay threshold of high priority tasks, Γ H represents a negative reward; said I(x) represents a 0-1 function;

[0140] described The utility corresponding to low-priority tasks:

[0141]

[0142] Among them, represents the completion delay threshold of low priority tasks, Γ L represents a negative reward, and α represents an exponential decay factor.

[0143] According to the calculation formula of the total utility, the calculation formula of the total utility is maximized through the preset constraints; wherein the constraints include: constraint C1 represents when l n,m =1, user equipment S n Through the only reflective surface node R m Offload tasks to edge servers, otherwise l n,m = 0; Constraint C2 means that a smart reflector node can only provide services to one user device at most; Constraint C3 means the phase shift of the smart reflector node belongs to the range of [0, 2π); constraints C4 and C5 correspond to the task offloading rate and the computing power allocation rate of the edge server respectively; constraint C6 represents that the computing power allocation ratio does not exceed the total computing power of the edge server; the optimization function of the calculation formula of the total utility determined according to the constraints is:

[0144]

[0145] According to the optimization function, the optimal phase shift configuration of the smart reflector node is optimized. The phase shift optimization objective function is:

[0146]

[0147] According to the triangle inequality, we can deduce that:

[0148] |g n,d +h m Θ k h n,m |≤|g n,d |+|h m Θ k h n,n |,

[0149] When arg(g n,d )=arg(h m Θ k h m,k ), assuming x m y m =h m Θ m h n,m ,at this time:

[0150]

[0151]

[0152] Then the phase shift optimization objective function becomes:

[0153]

[0154] Therefore, we can get

[0155]

[0156] Therefore, the intelligent reflector node R m Corresponding user equipment S n The first phase is:

[0157]

[0158] in and They are h m and h n,m The kth element of ;

[0159] The intelligent reflective surface node R m The optimal phase shift matrix is ​​Φ * ={θ1 * ,θ2 * ,…,θ M *} T ,in

[0160] Specifically, in this embodiment, it can be seen from P0 in the optimization function of the total utility that directly solving the problem has a high complexity. In addition, it can be observed that the SNR in the task offloading process only involves the phase shift matrix Φ and the RIS allocation matrix L, and these two variables are independent of other variables. Therefore, we use a decomposition method to decompose the problem into two sub-problems while maintaining the optimality of the solution. When performing phase shift optimization, if and only if arg(g n,d )=arg(h m Θ k h m,k ), the equation holds true. In this case, the phase of the signal reflected from the user-RIS and RIS-ES links is consistent with the phase of the user-ES direct link. We assume that x m y m =h m Θ m h n,m ,at this time:

[0161]

[0162]

[0163] In this embodiment, it can be seen from the optimization function corresponding to the calculation formula of the total utility that directly solving the problem is very complex. In addition, it is observed that the signal-to-noise ratio during the task offloading process only involves the phase shift matrix Φ and the allocation matrix of the smart reflector node, and these two variables are independent of other variables. Therefore, a decomposition method is used to decompose the problem into two sub-problems, reducing the data processing volume and improving processing efficiency while maintaining the optimality of the solution.

[0164] Step 104: Obtain a task priority corresponding to each of the plurality of users, and allocate a corresponding intelligent reflective surface node to each user according to the task priority.

[0165] In this embodiment, this step includes: after obtaining the optimal phase of the RIS, based on the different task priorities of each user, users with higher task priorities can prioritize RIS with higher SNR, and the selection process is from high to low. If there are users with the same task priority, a random strategy is adopted.

[0166] Step 105: Based on the task offloading information corresponding to the user assigned to each smart reflective surface node and the parameter information of the edge computing network, the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node are obtained through a preset deep reinforcement learning algorithm.

[0167] In this embodiment, the step includes: obtaining parameter information of the edge computing network, establishing a behavior network, and randomly initializing the network parameters of the behavior network; the behavior network is a policy network for deriving an unloading policy; establishing a first main penalty network and a second main penalty network, and randomly initializing the network parameters of the first main penalty network and the second main penalty network; the first main penalty network and the second main penalty network are used to evaluate the action value function of the current policy; establishing a first target neural network and a second target neural network and initializing the network parameters of the first target neural network and the second target neural network; the target neural network is used to update the parameters of the main neural network; initializing the experience pool, the current state of the environment, the current number of iterations and the upper limit of the total number of iterations; the experience pool The pool is used to store training samples; obtain an action based on the current strategy and state, calculate the total utility of the edge computing network according to the action, and calculate the reward and next state brought about by this state change; store the current state, the action based on the current strategy and state, the reward and the next state in the experience pool; randomly extract a small batch of samples from the experience pool, calculate the loss function of the first main penalty network and the second main penalty network respectively, and update the network parameters of the first main penalty network and the second main penalty network by the gradient descent method; use the loss function to update the current behavior network, and after several rounds of iterations, copy the parameters of the main neural network to the target neural network parameters; after iterative training for the total number of iterations, obtain the task offloading rate and the edge server computing power allocation rate.

[0168] Specifically, in this embodiment, in order to ensure the selection of the optimal RIS while maximizing the utility of the system, a reinforcement learning algorithm is used for policy decision-making. Furthermore, in order to solve the problem of system utility optimization, the present invention defines a Markov decision process. In the Markov decision process, the state space S of the system includes the SNR of each user's unloading process. The action space A of the system in the Markov decision process includes the unloading rate of each user's task and the ES computing power allocation rate. The strategy π in the Markov decision process is to output the probability distribution of an action under a given environmental state. The reward R in the Markov decision process is positive when the total system utility increases, 0 when the total system utility remains unchanged, and negative when the total system utility decreases. At time step τ, the agent is based on the state s of the RIS-assisted MEC network environment. τ ∈S and policy π selects action a τ ∈A. Then, action a τ Applied to the environment, resulting in a change from the current state s τ Transition to the next state s τ+1 The agent receives a reward r from the environment τ ∈R, the reward can be used to evaluate the effectiveness of its actions. According to the reward r receivedτ , the agent updates its policy π. The above iterative process allows the agent to gradually learn and optimize its policy by interacting with the environment. The specific steps of the SAC algorithm, or reinforcement learning algorithm, include:

[0169] 1) Obtaining information about the mobile edge computing network;

[0170] 2) Establishing an Actor network and randomly initializing its network parameter ν. The Actor neural network is a policy network used to derive an unloading strategy.

[0171] 3) Establish two main penalty (Critic) networks Q μ1 , Q μ2 , and randomly initialize its network parameters μ1 and μ2 to evaluate the action value function of the current strategy;

[0172] 4) Establish two target neural networks And initialize its network parameters The target neural network is used to update the parameters of the main neural network;

[0173] 5) Initializing the experience pool ER, which is used to store training samples;

[0174] 6) Initialize the current number of iterations e and obtain the upper limit of the total number of iterations E;

[0175] 7) Initialize the current state of the environment s τ ;

[0176] 8) Get the current strategy π and state s τ Action a τ ;

[0177] 9) According to action a τ To calculate the total utility U of the system total , based on which the reward r brought by this state change is calculated τ , and get the next state s τ+1 ;

[0178] 10)Tuple(s) τ ,a τ ,r τ ,s τ+1 ) is stored in the experience pool ER;

[0179] 11) Randomly extract a small batch of samples from the experience pool ER, calculate the loss function of the critic network, and update the parameters by gradient descent method

[0180] 12) Update the current Actor network using the loss function:

[0181] 13) After each Z iteration, where Z is the preset iteration value, the parameters μ1 and μ2 of the main neural network are copied to the parameters of the target neural network.

[0182] 14) The SAC network is trained after at most E rounds of iterative training or when it converges;

[0183] 15) Obtain effective task offloading and ES computing power allocation strategies.

[0184] In this embodiment, in order to ensure the selection of the optimal intelligent reflector node while maximizing the utility of the edge computing network, this step uses a reinforcement learning algorithm to make policy decisions, improve resource utilization, and obtain the optimal decision for task offloading.

[0185] Step 106: Allocate the offloading task of each of the plurality of users to the corresponding intelligent reflecting surface node, and offload the offloading task to the edge server for calculation according to the first phase, task offloading rate, and edge server computing power allocation rate corresponding to each intelligent reflecting surface node.

[0186] Furthermore, in an implementation method provided by an embodiment of the present invention, task offloading is performed by referring to the task offloading method based on intelligent reflective surface node selection provided in the embodiment. Specifically, in the implementation method, 1 ES, 10 users and 10 RIS are configured. Specifically, refer to Figure 3 , which is a system architecture diagram of a multi-user multi-intelligent reflector node provided by an embodiment of the present invention, refer to Figure 3 , the wireless channel in the network follows Rayleigh flat fading, where the average channel gain from the nth user to the ES is defined as (50+n) / 100, and the transmit power is set to 0.45W. If not specified, the computational power of the ES is set to 4×10 8 cycles / s, the wireless bandwidth of the network is 7MHz, the delay threshold of high-priority tasks is 6s, and the delay threshold of low-priority tasks is 8s. n The computing power follows The uniform distribution of , and the computation workload ξ is set to 6. Further, user S n The task size follows s n ~U(20,30)Mb is evenly distributed, and user S n The priority follows q n ~U(n / N,(n+1) / N) uniform distribution. In addition, the delay threshold of low priority tasks is set to 8s, the reward factor Γl Set to 1, the delay threshold of high priority tasks is set to 6s, and the penalty factor Γ h Set to -2. Based on the above configuration, in the Python simulation environment, simulations are performed using the task offloading method based on intelligent reflector node selection, random RIS selection scheme, fixed computing power, optimized offloading scheme, full offloading, optimized example scheme and full local computing scheme provided by the present invention. The convergence of the simulation algorithm can be referred to Figure 4 ,like Figure 4 As can be seen, the system utility convergence curves of the five schemes during the training process, with the number of training rounds ranging from 0 to 400. It can be clearly seen from the figure that, except for the all-local computing scheme, the utility of the four schemes increases with the increase of training epochs until convergence. This is because the all-local computing scheme only involves local computing without task offloading and computing power allocation, while the other four schemes can offload tasks to ES for synchronous computing, thereby improving the system utility. At the same time, the utility of the random RIS selection scheme is significantly lower than that of our proposed scheme, indicating that selecting RIS based on task priority can effectively guarantee the needs of users with different priorities, thereby improving the utility of the system. In addition, since the DRL algorithm explores effective task offloading and ES computing power allocation strategies, the system utility of our proposed scheme is significantly higher than that of the fixed computing power, optimized offloading and full offloading, optimized computing power schemes.

[0187] At the same time, based on the above network configuration and the preset Python simulation environment, simulations were conducted on the conditions of different ES computing capabilities, different user bandwidths, different delay thresholds for high-priority tasks, and different numbers of users, and the comparison between the solution proposed in the embodiment of the present invention and the other four solutions.

[0188] like Figure 5 The figure shows the comparison of the impact of the solution proposed by the present invention and other solutions on system utility under different ES computing capabilities in the Python simulation environment. Figure 5 As shown, the computing power of ES is from 2×10 8 cycles / s to 7×10 8cycles / s. As can be observed from the figure, except for the all-local computing solution, the utility of the other four solutions increases with the increase of ES computing power, but the upward trend becomes more gentle. This is because when the computing power of the ES is low, the computing latency at the ES is high, and thus the system utility is low. Therefore, the performance of the system improves significantly with the increase of ES computing power. At the same time, when the computing power of the ES is high, the ES can complete the task calculation faster, thereby reducing latency and reducing the impact on the system. In addition, it can be observed that the proposed solution is significantly better than the other four solutions due to its ability to optimize the offloading rate and the allocation of ES computing power.

[0189] Reference Figure 6 , Figure 6 The comparison diagram of the proposed solution and other solutions under the Python simulation environment with different user bandwidths is shown in the figure. Figure 6 , where the wireless bandwidth varies from 4MHz to 9MHz. Obviously, for all four schemes except the all-local computing scheme, the utility increases with the growth of wireless bandwidth. This performance is due to the fact that when the wireless bandwidth is limited, the task offloading transmission delay dominates the task computation delay, which leads to a significant improvement in system utility as the bandwidth increases. In addition, it can be observed that the utility of our proposed scheme exceeds that of the other four competing schemes, which demonstrates its superior ability in exploring effective task offloading and ES computing capacity allocation strategies using the DRL algorithm. The above results verify the superiority of our proposed method.

[0190] And refer to Figure 7 , Figure 7 The comparison diagram of the proposed scheme and other schemes under different delay thresholds of high-priority tasks in the Python simulation environment is as follows: Figure 7 , shows the impact of the delay threshold of high-priority tasks on the system utility of five different schemes, where the delay threshold changes from 4 seconds to 8 seconds. It can be observed that, except for the all-local computing scheme, the utility curves of the other four schemes show a steadily increasing trend. This trend is due to the increase in the delay threshold of high-priority users, and the other four schemes use the DRL algorithm to explore good strategies. In this case, more and more users' tasks can meet the delay requirements, thereby increasing the utility of the system. In addition, the utility of the all-local computing scheme remains unchanged because pure local computing cannot meet the needs of low-priority and high-priority users, so its utility remains unchanged. It is worth noting that since effective joint optimization strategies can be explored based on the DRL algorithm, the proposed scheme has higher utility than the other four schemes.

[0191] Last reference Figure 8 , Figure 8For the comparison between the solution proposed in the embodiment of the present invention and other solutions under different numbers of users, Figure 8 The effect of the number of users on the system utility of the five different schemes is shown, with the number of users ranging from 6 to 14. The figure shows that the system utility of our proposed fixed computing power, optimized offloading, full offloading, and optimized computing power schemes increases with the increase in the number of users, while the system utility of the random RIS selection and all-local computing schemes decreases with the increase in the number of users. The main reason behind this phenomenon is that the first three schemes rationally utilize RIS for signal enhancement, reducing task offloading delays and thus improving the overall effectiveness of the system. On the other hand, the random RIS selection and all-local computing schemes have inherent limitations. These limitations cause more users to experience task computation timeouts, which leads to higher negative penalties and thus reduces the system utility as the number of users increases. In addition, it is clear that our proposed scheme consistently outperforms the other four schemes. This is due to our proposed scheme's ability to effectively jointly optimize RIS allocation, task offloading, and ES computing power allocation.

[0192] On the other hand, the present invention also discloses a task offloading system based on intelligent reflective surface node selection. For the specific structure of the offloading system, please refer to Figure 2 The system includes a network construction module 201, a signal-to-noise ratio module 202, a phase optimization module 203, a node allocation module 204, an offloading decision module 205 and an offloading calculation module 206.

[0193] The network building module 201 is used to build an edge computing network based on a number of users to be task-offloaded, a preset edge server, and a number of intelligent reflective surface nodes.

[0194] The signal-to-noise ratio module 202 is configured to calculate the signal-to-noise ratio of the plurality of users performing task offloading through the plurality of intelligent reflective surface nodes based on the channel parameter information within the edge computing network.

[0195] The phase optimization module 203 is configured to optimize the phase of each of the plurality of smart reflective surface nodes by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each smart reflective surface node.

[0196] The node allocation module 204 is configured to obtain a task priority corresponding to each of the plurality of users, and allocate a corresponding smart reflective surface node to each user according to the task priority.

[0197] The offloading decision module 205 is used to obtain the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node through a preset deep reinforcement learning algorithm based on the task offloading information corresponding to the user assigned to each smart reflective surface node and the parameter information of the edge computing network.

[0198] The offloading calculation module 206 is used to assign the offloading tasks of each of the multiple users to the corresponding smart reflecting surface node, and offload the offloading tasks to the edge server for calculation based on the first phase, task offloading rate and edge server computing power allocation rate corresponding to each smart reflecting surface node.

[0199] As a preferred example, the signal-to-noise ratio module 202 includes:

[0200] According to the phase reflection matrices corresponding to the plurality of smart reflecting surface nodes and the allocation matrices corresponding to the plurality of users and the plurality of smart reflecting surface nodes, a signal-to-noise ratio of each user performing task offloading through the smart reflecting surface node is calculated.

[0201] Wherein, the phase reflection matrix is:

[0202] in, Indicates the phase of the smart reflector node that needs to be optimized; the K m represents the number of reflective elements equipped in each smart reflective surface node; m represents the number of smart reflective surface nodes; and j represents an imaginary unit.

[0203] Wherein, the allocation matrix is:

[0204]

[0205] Here, n represents the number of users.

[0206] The calculation formula of the signal-to-noise ratio is:

[0207]

[0208] Among them, P n Represents user equipment S n The transmission power; σ 2 represents the additive white noise variance; in addition, h n,m and g n,d Represents user equipment S n To the intelligent reflector node R m and wireless channel between edge servers; h m Represents the intelligent reflector node R m Wireless channel to edge server.

[0209] As a preferred example, the uninstallation decision module 205 includes a network construction unit, an initialization unit and an update iteration unit.

[0210] The network construction unit is used to obtain parameter information of the edge computing network, establish a behavior network, and randomly initialize the network parameters of the behavior network; the behavior network is a policy network, used to derive an unloading strategy; establish a first main penalty network and a second main penalty network, and randomly initialize the network parameters of the first main penalty network and the second main penalty network; the first main penalty network and the second main penalty network are used to evaluate the action value function of the current strategy; establish a first target neural network and a second target neural network and initialize the network parameters of the first target neural network and the second target neural network; the target neural network is used to update the parameters of the main neural network.

[0211] The initialization unit is used to initialize the experience pool, the current state of the environment, the current number of iterations, and the upper limit of the total number of iterations; the experience pool is used to store training samples; obtain actions based on the current strategy and state, calculate the total utility of the edge computing network based on the actions, and calculate the reward and next state brought about by this state change.

[0212] The update iteration unit is used to store the current state, the action based on the current strategy and state, the reward and the next state in the experience pool; randomly extract a small batch of samples from the experience pool, calculate the loss function of the first main penalty network and the second main penalty network respectively, and update the network parameters of the first main penalty network and the second main penalty network by gradient descent method; use the loss function to update the current behavior network, and after several rounds of iterations, copy the parameters of the main neural network to the target neural network parameters; after iterative training for the total number of iterations, obtain the task offloading rate and the edge server computing power allocation rate.

[0213] In addition to the above-mentioned method and system, an embodiment of the present invention also provides an electronic device and a computer-readable storage medium, wherein the electronic device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement a task offloading method based on intelligent reflective surface node selection according to an embodiment of the present invention when executing the program stored in the memory; the computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements a task offloading method based on intelligent reflective surface node selection according to an embodiment of the present invention.

[0214] Embodiments of the present invention disclose a task offloading method and system based on intelligent reflector node selection. The method first constructs a corresponding edge computing network based on a number of users corresponding to tasks to be offloaded, a number of preset intelligent reflector nodes for task offloading, and a preset edge server for task computing, to perform task offloading. After the edge computing network is constructed, the signal-to-noise ratio (SNR) of the edge computing network is calculated in real time based on the channel conditions of the edge computing network. The phases of the intelligent reflector nodes are optimized and adjusted based on the SNR to obtain the optimal phase corresponding to each intelligent reflector node, thereby significantly improving the signal quality of the edge computing network and, in turn, enhancing data transmission efficiency. After obtaining the optimal phase, corresponding nodes are assigned to users based on their task priorities to effectively allocate resources during task offloading. During resource allocation, the task offloading rate and edge server computing power allocation rate corresponding to each node in the edge computing network are optimized based on a preset reinforcement learning algorithm. The reinforcement learning algorithm ensures balanced resource allocation for the intelligent reflector nodes, improves resource utilization, and achieves optimal decision-making for task offloading.

[0215] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A task offloading method based on intelligent reflective surface node selection, characterized in that: include: An edge computing network is constructed based on a number of users to be task-offloaded, a preset edge server, and a number of intelligent reflective surface nodes; Calculating, based on the channel parameter information within the edge computing network, a signal-to-noise ratio of the plurality of users performing task offloading through the plurality of smart reflective surface nodes; wherein, based on a phase reflection matrix corresponding to the plurality of smart reflective surface nodes and an allocation matrix corresponding to the plurality of users and the plurality of smart reflective surface nodes, the signal-to-noise ratio of each user performing task offloading through the smart reflective surface node is calculated; The phase of each of the plurality of smart reflector nodes is optimized by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each smart reflector node; wherein, according to the signal-to-noise ratio calculation formula, the transmission rate of the user equipment corresponding to each user is calculated; according to the transmission rate, the transmission delay of the user equipment is calculated; according to the transmission delay, the calculation delay of the user equipment at the edge server is calculated; according to a preset local delay calculation formula, the delay of the user equipment when processing the task locally is calculated; according to the transmission delay and the calculation delay, the unloading delay of the user equipment is obtained; based on the unloading process and the local calculation being executed in parallel, the task delay of the user equipment when performing the task unloading is the maximum value of the unloading delay and the local delay; according to the task delay and the task priority, the total utility of the edge computing network is calculated; Obtaining a task priority corresponding to each of the plurality of users, and assigning a corresponding smart reflective surface node to each of the users according to the task priority; According to the task offloading information corresponding to the user assigned to each smart reflective surface node and the parameter information of the edge computing network, a preset deep reinforcement learning algorithm is used to obtain the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node; The offloading tasks of each of the plurality of users are respectively assigned to the corresponding intelligent reflecting surface node, and the offloading tasks are offloaded to the edge server for calculation according to the first phase, task offloading rate and edge server computing power allocation rate corresponding to each intelligent reflecting surface node.

2. The task offloading method based on intelligent reflective surface node selection according to claim 1, characterized in that: The signal-to-noise ratio of the tasks offloaded by the plurality of users through the plurality of intelligent reflective surface nodes is calculated based on the channel parameter information in the edge computing network. include: Wherein, the phase reflection matrix is: in, Indicates the phase of the smart reflector node that needs to be optimized; the K m represents the number of reflective elements equipped in each smart reflective surface node; m represents the number of smart reflective surface nodes; j represents an imaginary unit; Wherein, the allocation matrix is: Wherein, n represents the number of users; The calculation formula of the signal-to-noise ratio is: Among them, P n Represents user equipment S n The transmission power; σ 2 represents the additive white noise variance; in addition, h n,m and g n,d Represents user equipment S n To the intelligent reflector node R m and wireless channel between edge servers; h m Represents the intelligent reflector node R m Wireless channel to edge server.

3. The task offloading method based on intelligent reflective surface node selection according to claim 1, characterized in that: The optimizing the phase of each of the plurality of smart reflective surface nodes by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each smart reflective surface node includes: The calculation formula of the transmission rate is: P n =Blog2(1+ζn); where B is the wireless bandwidth of the user equipment; The calculation formula of the transmission delay is: Among them, the u n Represents user equipment S n The task size is in Mbits. n Represents user equipment S n Uninstall rate; The calculation formula for the calculation delay is: Among them, the is the computing power of the edge server, b n Represents the data allocated to user device S in the edge server n The computing power of ξ represents the number of CPU cycles required to process one bit of task; The local delay calculation formula is: Among them, the Indicates user equipment S n computing power; The calculation formula for the unloading delay is: The calculation formula for the task delay is: The calculation formula of the total utility is: Among them, the U n For user equipment S n The utility is: Among them, the The utility corresponding to high priority tasks: Among them, the q n The priority of the task, Represents the completion delay threshold of high priority tasks, represents a negative reward; said I(x) represents a 0-1 function; described The utility corresponding to low-priority tasks: Among them, Indicates the completion delay threshold of low priority tasks, represents a negative reward, and α represents an exponential decay factor.

4. The task offloading method based on intelligent reflective surface node selection according to claim 3, characterized in that: The obtaining of the first phase corresponding to each smart reflecting surface node further includes: According to the calculation formula of the total utility, the calculation formula of the total utility is maximized through the preset constraints; wherein the constraints include: constraint C1 represents when l n,m =1, user equipment S n Through the only reflective surface node R m Offload tasks to edge servers, otherwise l n,m = 0; Constraint C2 means that a smart reflector node can only provide services to one user device at most; Constraint C3 means the phase shift of the smart reflector node belongs to the range of [0, 2π); constraints C4 and C5 correspond to the task offloading rate and the computing power allocation rate of the edge server respectively; constraint C6 represents that the computing power allocation ratio does not exceed the total computing power of the edge server; the optimization function of the calculation formula of the total utility determined according to the constraints is: According to the optimization function, the optimal phase shift configuration of the smart reflector node is optimized, and the phase shift optimization objective function is: According to the triangle inequality, we can deduce that: |g n,d +h m Θ k h n,m |≤|g n,d |+|h m Θ k h n,m |, When arg(g n,d )=arg(h m Θ k h m,k ), assuming x m y m =h m Θ m h n,m ,at this time Then the phase shift optimization objective function becomes: Therefore, we can get Therefore, the intelligent reflector node R m Corresponding user equipment S n The first phase is: in and They are h m and h n,m The kth element of ; The intelligent reflective surface node R m The optimal phase shift matrix is in 5. The task offloading method based on intelligent reflective surface node selection according to claim 1, characterized in that: The method of obtaining the task offloading rate and edge server computing power allocation rate corresponding to each intelligent reflective surface node through a preset deep reinforcement learning algorithm includes: Obtain parameter information of the edge computing network, establish a behavior network, and randomly initialize network parameters of the behavior network; the behavior network is a policy network used to derive an offloading policy; Establishing a first main penalty network and a second main penalty network, and randomly initializing the network parameters of the first main penalty network and the second main penalty network; the first main penalty network and the second main penalty network are used to evaluate the action value function of the current strategy; Establishing a first target neural network and a second target neural network and initializing network parameters of the first target neural network and the second target neural network; the target neural networks are used to update the parameters of the main neural network; Initialize the experience pool, the current state of the environment, the current number of iterations, and the upper limit of the total number of iterations; the experience pool is used to store training samples; Obtain an action based on the current policy and state, calculate the total utility of the edge computing network based on the action, and calculate the reward and next state brought about by this state change; Storing the current state, the action based on the current strategy and state, the reward, and the next state in the experience pool; Randomly extract a small batch of samples from the experience pool, calculate the loss function of the first main penalty network and the second main penalty network respectively, and update the network parameters of the first main penalty network and the second main penalty network by gradient descent method; Use the loss function to update the current behavior network, and after several rounds of iterations, copy the parameters of the main neural network to the parameters of the target neural network; After iterative training for the total number of iterations, the task offloading rate and the edge server computing power allocation rate are obtained.

6. A task offloading system based on intelligent reflective surface node selection, characterized in that: The system includes a network construction module, a signal-to-noise ratio module, a phase optimization module, a node allocation module, an offloading decision module and an offloading calculation module; The network building module is used to build an edge computing network based on a number of users to be task-offloaded, a preset edge server, and a number of intelligent reflective surface nodes; The signal-to-noise ratio module is used to calculate the signal-to-noise ratio of the multiple users performing task offloading through the multiple smart reflective surface nodes based on the channel parameter information in the edge computing network; wherein the signal-to-noise ratio module is used to calculate the signal-to-noise ratio of each user performing task offloading through the smart reflective surface node based on the phase reflection matrix corresponding to the multiple smart reflective surface nodes and the allocation matrix corresponding to the multiple users and the multiple smart reflective surface nodes; The phase optimization module is used to optimize the phase of each of the several smart reflector nodes by maximizing the signal-to-noise ratio to obtain a first phase corresponding to each smart reflector node; wherein, according to the calculation formula of the signal-to-noise ratio, the transmission rate of the user equipment corresponding to each user is calculated; according to the transmission rate, the transmission delay of the user equipment is calculated; according to the transmission delay, the calculation delay of the user equipment at the edge server is calculated; according to a preset local delay calculation formula, the delay of the user equipment when processing the task locally is calculated; according to the transmission delay and the calculation delay, the unloading delay of the user equipment is obtained; based on the unloading process and the local calculation being executed in parallel, the task delay of the user equipment when performing the task unloading is the maximum value of the unloading delay and the local delay; according to the task delay and the task priority, the total utility of the edge computing network is calculated; The node allocation module is used to obtain the task priority corresponding to each of the plurality of users, and allocate a corresponding smart reflective surface node to each user according to the task priority; The offloading decision module is used to obtain the task offloading rate and edge server computing power allocation rate corresponding to each smart reflective surface node through a preset deep reinforcement learning algorithm based on the task offloading information corresponding to the user assigned to each smart reflective surface node and the parameter information of the edge computing network; The offloading calculation module is used to allocate the offloading tasks of each of the multiple users to the corresponding smart reflecting surface node, and offload the offloading tasks to the edge server for calculation according to the first phase, task offloading rate and edge server computing power allocation rate corresponding to each smart reflecting surface node.

7. The task offloading system based on intelligent reflective surface node selection according to claim 6, characterized in that: The signal-to-noise ratio module also include: Wherein, the phase reflection matrix is: in, Indicates the phase of the smart reflector node that needs to be optimized; the K m represents the number of reflective elements equipped in each smart reflective surface node; m represents the number of smart reflective surface nodes; j represents an imaginary unit; Wherein, the allocation matrix is: Wherein, n represents the number of users; The calculation formula of the signal-to-noise ratio is: Among them, P n Represents user equipment S n The transmission power; σ 2 represents the additive white noise variance; in addition, h n,m and g n,d Represents user equipment S n To the intelligent reflector node R m and wireless channel between edge servers; h m Represents the intelligent reflector node R m Wireless channel to edge server.

8. The task offloading system based on intelligent reflective surface node selection according to claim 6, characterized in that: The uninstallation decision module includes a network construction unit, an initialization unit and an update iteration unit; The network building unit is used to obtain parameter information of the edge computing network, establish a behavior network, and randomly initialize the network parameters of the behavior network; the behavior network is a policy network, which is used to derive an unloading policy; Establishing a first main penalty network and a second main penalty network, and randomly initializing the network parameters of the first main penalty network and the second main penalty network; the first main penalty network and the second main penalty network are used to evaluate the action value function of the current strategy; establishing a first target neural network and a second target neural network and initializing the network parameters of the first target neural network and the second target neural network; the target neural network is used to update the parameters of the main neural network; The initialization unit is used to initialize the experience pool, the current state of the environment, the current number of iterations, and the upper limit of the total number of iterations; The experience pool is used to store training samples; Obtain an action based on the current policy and state, calculate the total utility of the edge computing network based on the action, and calculate the reward and next state brought about by this state change; The update iteration unit is used to store the current state, the action based on the current strategy and state, the reward and the next state in the experience pool; A small batch of samples are randomly drawn from the experience pool, and the loss functions of the first main penalty network and the second main penalty network are calculated respectively. The network parameters of the first main penalty network and the second main penalty network are updated by the gradient descent method; the current behavior network is updated with the loss function, and after each round of iteration, the parameters of the main neural network are copied to the parameters of the target neural network; after iterative training for the total number of iterations, the task offloading rate and the edge server computing power allocation rate are obtained.

9. An electronic device, characterized in that: The invention comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement a task offloading method based on intelligent reflecting surface node selection as described in any one of claims 1 to 5 when executing the program stored in the memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the task offloading method based on intelligent reflecting surface node selection according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Unloading decision-making method of mobile edge computing system based on intelligent reflecting surface assistance

    CN113543176A