Task unloading and resource management method for ultra-dense heterogeneous network oriented to sympathetic calculation fusion
By building the Dec-POMDP model and using the GAT-MADDPG algorithm to optimize service cache, channel allocation and task offloading, the problem of inefficient resource management in ultra-intensive heterogeneous networks is solved, and energy consumption is minimized and performance improvement is achieved under sensing accuracy and delay constraints.
Patent Information
- Application Number
- CN202510419019.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
In super-intensive heterogeneous networks, traditional single agent optimization methods are difficult to cope with high-dimensional non-convex optimization problems, cannot take into account the comprehensive performance of communication, perception and computing requirements, and the interference problems between base stations are prominent, resulting in inefficient resource management.
Build a system model, including service caching, resource division, communication and perception models, which are transformed into Dec-POMDP problems, and use the GAT-MADDPG algorithm for solving them, optimize service cache decisions, channel allocation and task offload decisions, and use graph attention network to capture the spatial correlation between base stations and optimize resource management.
Under the constraints of sensing accuracy and delay, reduce system energy consumption, improve system performance, improve resource management efficiency, reduce interference between base stations, and achieve green and efficient network resource management.
Smart Images

Figure CN120264313A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and relates to a method for task offloading and resource management in an ultra-dense heterogeneous network for integrated communication, sensing and computing. Background Art
[0002] With the rapid development of the sixth-generation mobile communication technology, the integrated technology of sensing, communication and computing (ISCC) is becoming an important direction to meet the diverse requirements of future wireless communication networks. For the increasing number of Internet of Things (IoT) terminals, the wireless connection function achieved solely by relying on the cellular wireless network centered on macro base stations (MBS) cannot meet their needs. To solve this problem, ultra-dense heterogeneous networks (UDHN) are considered a promising solution. Integrating UDHN into the ISCC system can further shorten the distance between IoT terminals and the computing center. The dense deployment of small base stations not only provides higher network capacity and coverage for IoT terminals to achieve precise sensing, but also brings high energy consumption and serious interference problems. In addition, the sensing requirements and delay sensitivity of IoT terminals increase, making it an urgent problem to reasonably manage resources under multi-dimensional resource constraints. Therefore, how to achieve green and efficient network resource management while meeting the diverse requirements of IoT terminals has become one of the hot issues concerned by the current academic and industrial communities.
[0003] In recent years, the integrated communication and sensing technology (ISAC) has become one of the key technologies for next-generation wireless networks by sharing radio and hardware resources between communication systems and radar systems. Mobile edge computing (MEC), as an efficient computing architecture, allows IoT terminals with limited computing resources to offload computationally intensive tasks to edge servers for execution. The joint design of ISAC and MEC systems to achieve ISCC has become a new research hotspot. In addition, in 6G and future wireless networks, the ultra-dense heterogeneous network architecture has become one of the typical examples and has received a lot of attention. To achieve efficient coordination of communication, sensing and computing tasks, it is necessary to solve the resource optimization problem in multi-agent systems.
[0004] However, due to the interaction and mutual influence among base stations, especially in ultra-dense heterogeneous networks, the interference problem among base stations becomes particularly prominent, and traditional single-agent optimization methods are difficult to handle high-dimensional non-convex optimization problems and difficult to balance the comprehensive performance of communication, sensing, and computing requirements. Summary of the Invention
[0005] To solve the above-mentioned problems of the prior art, the present invention adopts a task offloading and resource management method for ultra-dense heterogeneous networks oriented to communication, sensing, and computing integration, including:
[0006] S1. Construct a system model; the system model includes: a macro base station, small base stations, and their associated Internet of Things terminals;
[0007] S2. Construct a service cache model, a resource partitioning model, a communication model, a sensing model, and a task offloading model according to the system model;
[0008] S3. Construct a joint optimization problem according to the service cache model, the resource partitioning model, the communication model, the sensing model, and the task offloading model;
[0009] S4. Transform the joint optimization problem into a Dec-POMDP; where Dec-POMDP is a distributed partially observable Markov decision process;
[0010] S5. Solve the joint optimization problem by using the GAT-MADDPG algorithm based on the Dec-POMDP to obtain the optimal service cache decision, channel allocation decision, transmission power of the Internet of Things terminal, and task offloading decision; where GAT is graph attention, and MADDPG is a multi-agent deep deterministic policy gradient algorithm.
[0011] Beneficial Effects:
[0012] 1. The present invention proposes a joint optimization problem of service caching, task offloading, channel allocation, and power control in the scenario of a communication-sensing-computation integrated ultra-dense heterogeneous network, aiming to minimize the energy consumption of the system under service caching, sensing accuracy, and latency constraints, thereby improving the system performance. 2. The present invention models the joint optimization problem as a distributed partially observable Markov decision process (Dec-POMDP) and uses the GAT-MADDPG algorithm to solve it. The GAT-MADDPG algorithm effectively captures the spatial correlation between base stations through a graph attention network, enabling each base station to better understand the states of its neighboring base stations, and thus optimizing resource management decisions. 3. There will be interference when IoT terminals select the same sub-channel to transmit data, and the service caching states of adjacent base stations will affect the task offloading strategy. Therefore, the present invention considers the service caching state information and channel gain to construct the node features of agent m, so that the node features of each base station contain important state information, enabling the base station to better understand the states of its neighboring base stations, and thus optimizing resource management decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 FIG. is a flowchart of a task offloading and resource management method for a communication-sensing-computation integrated ultra-dense heterogeneous network provided by an embodiment of the present invention;
[0014] Figure 2 FIG. is a schematic diagram of an ultra-dense heterogeneous network under ISCC provided by an embodiment of the present invention;
[0015] Figure 3 FIG. is a schematic diagram of a resource allocation model provided by an embodiment of the present invention;
[0016] Figure 4 FIG. is a framework diagram of the GAT-MADDPG algorithm provided by an embodiment of the present invention;
[0017] Figure 5 FIG. is a framework diagram of a critic network provided by an embodiment of the present invention;
[0018] Figure 6 FIG. is a comparison schematic diagram between the method of the present invention and a comparative algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] As Figure 1As shown in the figure, the embodiment of the present invention adopts a method for task offloading and resource management in a hyperscale heterogeneous network oriented to the integration of communication, sensing, and computing, including:
[0021] S1. Construct a system model;
[0022] The system model of the present invention is as Figure 2 shown. Considering a hyperscale heterogeneous network oriented to the integration of communication, sensing, and computing, it consists of a macro base station and a number of small base stations; define the set of small base stations SBS as If the index of the macro base station MBS is represented by 0, then the set of all base stations is Within the coverage area of each small base station, a number of Internet of Things (IoT) terminals are randomly deployed. Under base station m, define the set of IoT terminals as K m is the number of IoT terminals under base station m. Each base station is equipped with a mobile edge computing (MEC) server with a certain computing capacity. The system operates in discrete time slots, and define the time set as T is the number of time. All IoT terminals are equipped with integrated sensing and communication technology, and perform sensing and communication through integrated signals. The IoT terminals obtain the surrounding environment information through sensors, generate tasks (tasks are indivisible), and further perform calculations locally or by the base station. The small base stations are connected to the macro base station through wired links.
[0023] For IoT terminal k under base station m m , the task generated in time slot t is defined as where is the data volume size of the task, with the unit of bit; is the number of CPU cycles required to calculate 1 bit of the task, is the service type required to execute the task; is the maximum tolerable delay for executing the task, is the minimum value to meet the sensing accuracy.
[0024] S2. Construct a service cache model, a resource partitioning model, a communication model, a sensing model, and a task offloading model according to the system model;
[0025] Construct a service cache model:
[0026] When the terminal offloads the task to the MEC server of the base station for processing, the MEC server can only provide computing services for the terminal if it caches the services required for the computing task; there are C different types of services in the network, and define the service type set Service cache decision Indicate whether the server of base station m caches service c at time slot t; specifically, when the MEC server of base station m caches service c, ψ c,m (t) is 1, otherwise it is 0. The cache space constraint can be expressed as:
[0027]
[0028] where z c represents the size of the storage space occupied by caching service c; Z m is the size of the storage space of base station m.
[0029] Building a resource partitioning model includes:
[0030] Considering the dense deployment of base stations makes the network environment increasingly complex, and thus makes the interference problem extremely serious. To effectively solve the interference problem, an effective interference management mechanism is introduced, that is, the spectrum is segmented as Figure 3 shown. To eliminate the interference between MBS and SBS, the frequency band W is divided into two parts, namely W = W MBS +W SBS , W MBS =λW is the spectrum of the macro base station, W SBS =(1 - λ)W is the spectrum of the small base station, λ is the frequency band division factor, and satisfies 0≤λ≤1. Given that the mutual interference generated between SBSs is relatively weak, in order to greatly improve the spectrum efficiency, small base stations are allowed to share and reuse spectrum resources. This spectrum allocation strategy insightfully reveals the essential differences between macro base stations and small base stations, thereby being able to reduce the interference level between base stations and significantly enhance the spectrum utilization efficiency of the entire network. In addition, by moderately multiplexing spectrum resources between small base stations, the potential of the spectrum is further explored, enabling the network to more flexibly and efficiently adapt to the communication requirements of different regions, and comprehensively improving the communication quality and network comprehensive performance.
[0031] If the bandwidth of any channel n is B, then the number of available channels N m of MBS and SBS are λW / B and (1 - λ)W / B respectively, then define as the set of available channels for all base stations, as the available channel resources of the base station. When m = 0, it represents the available channel resources of the macro base station, and when m≠0, it represents the available channel resources of the small base station.
[0032] Building a communication model includes:
[0033] In the synaesthesia-computing fusion ultra-dense heterogeneous network environment, when the Internet of Things (IoT) terminal offloads tasks to the base station, the communication process is as follows: The IoT terminal acquires idle channel resources, adjusts the transmission power, and sends the tasks to the edge server of the base station through the wireless channel for processing. Considering that the calculation result data is much smaller than the calculation input data, the result backhaul delay is negligible.
[0034] Construct the available channel allocation decision for each base station Denote the channel selection variable of IoT terminal k at time slot t m If the base station allocates channel n to IoT terminal k m , then Otherwise
[0035] Then the interference received from other IoT terminals at base station m at time slot t Can be expressed as:
[0036]
[0037] Where, Denote the transmission power of IoT terminal k m At time slot t, Denote the channel gain when IoT terminal k m Transmits tasks to base station m through sub-channel n, Denote that in addition to IoT terminal k m Using channel n, the mobile IoT terminal k covered by other base stations m' m' Also transmits data through channel n, causing interference to the transmission of IoT terminal k m .
[0038] The signal-to-noise ratio when IoT terminal k m Uploads tasks to base station m through sub-channel n can be expressed as:
[0039]
[0040] Where, Denote the co-channel interference. Since small base stations share spectrum resources, when IoT terminal k m Transmits data to the base station through sub-channel n, other IoT terminals using this channel will generate co-channel interference. Macro base stations use separate spectrum resources and there is no channel interference.
[0041] According to Shannon's formula, the uplink transmission rate received at base station m from IoT terminal k m Can be expressed as:
[0042]
[0043] Constructing the perception model includes:
[0044] In the perception model, it is assumed that each Internet of Things (IoT) terminal integrates an OFDM waveform through its dual-functional transmitter. The integrated waveform is reflected by the target for radar monitoring and is also decoded by the base station to achieve the transmission of perception results. On sub-channel n, IoT terminal k m The continuous integration of S OFDM symbols can be expressed as:
[0045]
[0046] where is the center frequency of sub-channel n, is the amplitude of the integrated waveform of IoT terminal k m on sub-channel n, is the phase code of the waveform of the modulated symbol l of IoT terminal k m on channel n, T s is the duration of a single OFDM symbol with a cyclic prefix, and rect[x] is the rectangular function whose value is 1 when 0 ≤ x ≤ 1 and 0 otherwise.
[0047] Here, considering the impulse response m of IoT terminal k on channel n is a wide-sense stationary Gaussian random process, then the reflected signal m received by IoT terminal k can be expressed as:
[0048]
[0049] where w(t) is additive white Gaussian noise with a mean of 0 and a variance of σ 2 .
[0050] The present invention uses the conditional mutual information (MI, mutual information) between the echo signal and the impulse response to measure the perception performance, which can be expressed as:
[0051]
[0052] where is the radar signal-to-noise ratio of IoT terminal k m on channel n, and its expression is:
[0053]
[0054] where is the impulse response m of IoT terminal k At the frequency f of sub-channel n n The Fourier transform at represents the channel gain on sub-channel n of the radar receiver from the Internet of Things (IoT) terminal k m to the IoT terminal k m′ . Let \(h_{k,k}^{n}\) be the channel gain, \(P_{k}\) and \(P_{k'}\) be the transmission powers of the IoT terminals k and k m respectively, and \(\sigma\) m′ be the noise 2 .
[0055] Build a task offloading model
[0056] Build the task offloading decision of each IoT terminal k m , where \(x_{k,m}\in\{0, 1\}\) indicates whether the task of the IoT terminal k is processed by the base station m m .
[0057] Task processed at the local base station: The transmission delay of the task offloaded by the IoT terminal k m to the base station m at time slot t can be expressed as
[0058]
[0059] Correspondingly, the communication energy consumption of the IoT terminal k m transmitting the task can be expressed as
[0060]
[0061] The computing delay of the task at the local base station can be expressed as
[0062]
[0063] where F m represents the maximum available computing frequency of the MEC server of the base station m
[0064] The computing energy consumption of the MEC server processing the task can be expressed as
[0065]
[0066] where \(\epsilon\) m represents the effective energy coefficient related to the edge server chip architecture
[0067] Task Processing at Adjacent Base Stations: When the local base station does not have the service required to process the task, it is first considered to be processed by the adjacent base station. The task execution delay consists of three parts: the delay of transmitting the task to the local base station, the delay of forwarding it to the adjacent base station, and the computing delay. At time slot t, the Internet of Things (IoT) terminal k m The transmission delay of unloading the task to the adjacent base station m' can be expressed as:
[0068]
[0069] where R m,m' (t) represents the average transmission rate of the task from BSm to BSm'.
[0070] The Internet of Things (IoT) terminal k m The transmission energy consumption of uploading the task to the adjacent base station can be expressed as:
[0071]
[0072] where ξ represents the power consumption of the task from BSm to BSm'.
[0073] The computing delay of the task at the adjacent base station can be expressed as:
[0074]
[0075] In the formula, F m' represents the maximum available computing frequency of the MEC server of base station m'.
[0076] Correspondingly, the computing energy consumption of the task at the adjacent base station can be expressed as:
[0077]
[0078] In summary, the total execution delay and energy consumption of the task executed by the small base station
[0079]
[0080] where is the total delay of the task executed by the local base station, is the total delay of the task executed by the adjacent base station, is the total execution energy consumption of the task executed by the local base station, is the total execution energy consumption of the task executed by the adjacent base station.
[0081] Task Processing at Macro Base Station:
[0082] When the local base station and the adjacent base stations do not have the services required for caching and processing the task, the task is offloaded to the macro base station for processing. The task execution delay includes the transmission time and the computing time. In time slot t, the Internet of Things (IoT) terminal k m The transmission delay of offloading the task to the macro base station can be expressed as:
[0083]
[0084] Correspondingly, for the IoT terminal k m The transmission energy consumption of uploading the task to the macro base station can be expressed as:
[0085]
[0086] The computing delay of the task at the macro base station can be expressed as:
[0087]
[0088] where F0 represents the maximum available computing frequency of the MEC server at the macro base station.
[0089] Correspondingly, the computing energy consumption of the task at the macro base station can be expressed as:
[0090]
[0091] In summary, the total execution delay and energy consumption of the task executed by the macro base station can be respectively expressed as:
[0092]
[0093] For the IoT terminal k m The task execution delay and energy consumption can be respectively expressed as:
[0094]
[0095] S3. Construct a joint optimization problem according to the service caching model, resource partitioning model, communication model, sensing model, and task offloading model;
[0096] The present invention aims to improve the system performance by optimizing the total energy consumption of all IoT terminal tasks completed in the system. Specifically, under multiple constraints such as computing resources, delay, and sensing accuracy, a joint optimization method is proposed, covering key factors such as service caching, task offloading, sub-channel allocation, and transmission power, to minimize the system energy consumption. The optimization problem is as follows:
[0097]
[0098] Among them, the constraint C1 indicates that the task completion delay of the IoT terminal cannot exceed the maximum delay of task execution. The constraints C2 and C3 indicate that any IoT terminal can only be associated with one channel. The constraint C4 indicates the constraint of channel allocation decision. The constraint C5 indicates the upper and lower bounds of the transmission power of the IoT terminal. C6 represents the minimum requirement to ensure the sensing accuracy of the IoT terminal. The constraints C7 and C8 indicate that the IoT terminal can select at most one base station to remotely process the task. The constraint C9 indicates that the cached service cannot exceed the cache capacity of each base station. For IoT terminal k m is the maximum transmission power.
[0099] S4. Transform the joint optimization problem into a Dec-POMDP; among them, Dec-POMDP is a distributed partially observable Markov decision process;
[0100] The optimization objective can be solved by finding the optimal decision variables b, x, p, and ψ in all time slots. However, due to the large number of states, this solution process is extremely complex. In addition, considering the dynamic distributed edge computing environment, the base station has no way to obtain complete environmental information. Therefore, it is very difficult to solve this problem using traditional optimization methods. The present invention abstracts the optimization objective into a distributed partially observable Markov decision process (Dec-POMDP) and uses a reinforcement learning algorithm to solve this problem.
[0101] Specifically, the present invention takes each base station as an agent. Considering the highly dynamic network environment and distributed computing offloading strategy, the base station cannot obtain complete environmental state information, so the optimization problem is abstracted into a Dec-POMDP. This process can be represented by the tuple where, represents the global state space, represents the local observation space of the agent, represents the global action space of the agent, represents the reward function.
[0102] The global state space includes: at time slot t, the global environmental state contains detailed information of all IoT terminal tasks and the wireless channel state in the network. Specifically, it can be represented as the global information where, ψ(t - 1) is the state information of the base station cached service, represents the set of task information of all IoT terminals, represents the set of instantaneous channel gains between the base station and all IoT terminals.
[0103] Observation Space: In the ultra-dense heterogeneous network environment with the integration of communication and sensing, at time t, the agent m can only observe the task details of the Internet of Things (IoT) terminals within the coverage of the current base station and the set of instantaneous channel gains between this base station and other IoT terminals. Specifically, it can be expressed as where, ψ(t - 1) is the service cache information of base station m, represents the set of task information of the IoT terminals within the coverage of base station m, represents the set of instantaneous channel gains between base station m and all IoT terminals in the environment.
[0104] Action Space: The agent m executes actions according to the observed environmental information o m (t) and the current policy π m . The actions of the agent at this moment include service cache decision, task offloading decision, channel allocation decision, and power management decision. Specifically, it can be expressed as action where, ψ m (t) represents the service cache decision of base station m, b m (t) represents the task offloading decision of the IoT terminals under base station m, x m (t) represents the available channel allocation decision under base station m, represents the set of transmission powers of the IoT terminals under base station m.
[0105] Reward Function: The reward function is the reward obtained by the agent for choosing a specific action in each state. By maximizing the expected total reward, the agent is guided to learn to execute actions that are more beneficial to the overall performance of the system. The optimization objective of this chapter is to minimize the energy consumption of the system under the constraints of sensing accuracy and latency. Therefore, this chapter gives the negative value of the system energy consumption as the reward to the agent, and gives an additional reward to the agent whose decision meets the latency and sensing accuracy requirements. The reward function of agent m can be expressed as:
[0106]
[0107] In the formula, represents the energy consumption of the IoT terminals within the coverage of base station m; represents whether the agent's decision meets the sensing accuracy requirements of the IoT terminals; represents whether the agent's decision meets the latency requirements of the IoT terminals. Among them, H(·) is the Heaviside step function, and η1 and η2 are reward coefficients. When the latency and sensing accuracy requirements are met, Y m (t) = η1, S m (t) = η2, otherwise, Y m (t) = 0, S m (t) = 0.
[0108] S5. Solve the joint optimization problem using the GAT-MADDPG algorithm based on Dec-POMDP to obtain the optimal channel allocation decision and task offloading decision; where GAT-MADDPG is the multi-agent deep deterministic policy gradient algorithm.
[0109] The algorithm framework is as Figure 4 shown. Each base station is an agent configured with an actor-critic network. A graph attention network is embedded in the critic network, and the agent specifically focuses on the state information of adjacent agents, mines the potential spatial correlation, and thus makes a better decision. This algorithm uses the CTDE method for learning; where centralized training can obtain the historical information of the distributed edge network environment state, use this information to guide the update of network parameters, and thus learn a better policy. In the distributed execution stage, the agent only decides to execute actions based on its local observation data and the current policy.
[0110] Specifically, the actor-critic network includes: Actor network, GAT-based Critic network, target Actor network, and GAT-based target Critic network; the agent uses an architecture of distributed execution and centralized training to learn and obtain the optimal service caching decision, channel allocation decision, and task offloading decision;
[0111] The agent uses an architecture of distributed execution and centralized training to learn, including:
[0112] In the distributed execution stage, each agent observes the global state space S at the current moment to obtain the local observation information o at the current moment m ; each agent inputs the local observation information o m into its own Actor network to obtain the action a selected by each agent at the current moment m ; each agent executes the action a it selects m to obtain the local observation information o' of each agent at the next moment m and the reward r m ; combine the local observation information o m of all agents, the action a m , the local observation information o' at the next moment m and the reward r m into an experience sample Store the experience sample in the experience replay pool;
[0113] In the centralized training stage, sample experience samples from the experience replay pool Train the actor-critic network of each agent according to the sampled empirical samples to obtain the trained actor-critic network of each agent; where u represents the serial number of the empirical samples, and U represents the number of sampled empirical samples;
[0114] Use π = {π1, π2,..., π M} and Q = {Q1, Q2,..., Q M} to represent the actor network and the critic network respectively, and the corresponding parameters are represented by θ = {θ1, θ2,..., θ M} and w = {w1, w2,..., w M} respectively; Training the actor-critic network of the agent according to the sampled empirical samples includes:
[0115] S51. Input the local observation information into the Actor network to obtain the action
[0116] S52. Input the action a u and the local observation information into the GAT-based Critic network to obtain the action value function
[0117] The GAT-based Critic network includes: a GAT network and a Critic network; as Figure 5 shown, the GAT-based Critic network processes the action a u and the local observation information including:
[0118] Model the multi-agent environment as an undirected graph G(V, E); where V represents the set of agents, and E represents the set of edges between agents;
[0119] Considering that there will be interference when terminals under different SBSs select the same sub-channel to transmit data, and the cache status of other base stations will affect the task offloading strategy. Base stations can share cache information, which can improve the task execution hit rate of IoT terminals, thereby improving resource utilization. Channel gain is a key factor affecting interference between base stations. Base stations adjust their own resource allocation strategies through the wireless channel information of other base stations to reduce interference and optimize power control or sub-channel selection. So when the agent m is the central node, according to the service cache status information ψ m′ (t - 1) (0 or 1) composed vector and the IoT terminal k m' within the coverage of other agents m' The composed vector Construct the node feature h of the agent m m (t);
[0120] Input the node feature h m (t) and the undirected graph G(V, E) into the GAT network to obtain the potential feature vector of the agent m Input the potential feature vector of the agent m Action a u And the task information Input into the Critic network to obtain the action value function
[0121] Construct the node feature h of the agent m m (t) includes:
[0122] Perform a linear transformation on the service cache status information vector And the channel gain vector :
[0123]
[0124] Among them, W ψ Is the cache feature projection matrix, W g Is the channel feature projection matrix, β ψ And β g Are the cache feature bias and the channel feature bias respectively. MLP is a feedforward neural network, Relu and Leaky Relu are activation functions. MLP is a feedforward neural network. The cache status information is 0 or 1, and the data has sparsity and non-negativity, so the Relu function is used for activation; while the channel gain is continuous data, and using the Leaky Relu function for activation can avoid the vanishing gradient of the channel features.
[0125] Perform weighted fusion on the above two transformed node features:
[0126] h m (t) = αψ′ m (t - 1)+(1 - α)g′ m (t)
[0127] α = σ(W α [ψ′ m (t - 1); g′ m (t)]+β α )
[0128] In the formula, α ∈ (0, 1), α is the weight of the cache node feature, W α Is the weight generation vector, σ(·) is the Sigmoid function, β αGenerate a bias for the weight. When the interference between base stations intensifies (at this time, α approaches 1). As α → 1, the model depends more on the cache state; when the interference between base stations decreases, at this time α → 0, the model pays more attention to the channel quality.
[0129] When the interference between base stations intensifies, the channel gain will decrease. The channel gain may become negative. After being projected by leakyRelu, a low-quality channel representation close to 0 is output. The interference does not affect the service cache state, and the service cache feature remains a stable positive value; the overall value inside the sigmoid function decreases, and the saturation characteristic of the sigmoid function makes α tend to 1. W α Learn the positive weight of the service cache and the negative weight of the channel gain.
[0130] The GAT network processes the node feature h m (t) and the undirected graph G(V,E), including:
[0131] As shown in the appendix Figure 4 As shown, by adding a graph attention network to the critic network, the attention coefficients of other agents can be obtained, and the feature information can be aggregated according to these coefficients, guiding the agent to focus on the feature information of adjacent base stations in a targeted manner and select a better action.
[0132] In the graph attention network, the attention coefficient between agent m and agent m' can be expressed as:
[0133] e mm' =att(Wh m (t),Wh m' (t))
[0134] In the formula, att(·) represents the attention mechanism, W is the corresponding learnable weight matrix, and the attention coefficient e mm' represents the importance of the state information of agent m' to agent m. The features of each agent and its adjacent agents are different, so the attention coefficient is asymmetric, that is, e mm' ≠e m'm . To facilitate comparison between different agents, according to the undirected graph G(V,E), define as the set of neighbor agents of agent m. The neighbor base station is assumed to be the other base station closest to this base station. Use the soft update function to normalize the attention coefficient to obtain the corresponding attention weight δ mm' :
[0135]
[0136] To stabilize the learning process of the algorithm, a multi-head attention mechanism is adopted. Given the normalized attention coefficients, for an agent m with P independent attention mechanisms, its potential feature vector can be expressed as:
[0137]
[0138] where σ is a non-linear function; the symbol || represents the concatenation operation; p is the serial number of the attention mechanism, and P is the number of heads.
[0139] Through the GAT network, the agent can specifically extract the feature information of adjacent base stations, improve its own strategy, and thus execute better actions.
[0140] S53. Input the local observation information into the target Actor network to obtain the action
[0141] S54. Input the action a′ u and the local observation information into the GAT-based target Critic network to obtain the target action value function
[0142] S55. Update the Actor network according to the action a u and the action value function Soft-update the target Actor network according to the parameters of the updated Actor network; according to the action value function the target action value function and the reward function calculate the loss function value, update the GAT-based Critic network according to the loss function value, and soft-update the GAT-based target Critic network according to the parameters of the updated GAT-based Critic network; when the loss function value converges to stability (i.e., the minimum), obtain the trained actor-critic network.
[0143] The actor network updates the policy π through the gradient of the expected reward m :
[0144]
[0145] where U is the number of samples randomly sampled in a small batch, u is the serial number of the sample, is the action value function.
[0146] The GAT-based Critic network updates the parameter ω by minimizing the loss function of the agent m , and the loss function It can be expressed as:
[0147]
[0148] where, π′ m is the target actor network, is the target critic network, and γ is the discount factor;
[0149] The parameters θ' m and w' m of the target network are updated by soft update: θ′ m = ζθ m +(1 - ζ)θ′ m 、ω′ m = ζω m +(1 - ζ)ω′ m where ζ is the network update rate;
[0150] Based on the trained actor - critic network of each agent, the optimal service caching decision, channel allocation decision, transmission power of the IoT terminal, and task offloading decision are obtained.
[0151] The convergence curves of the average rewards of the GAT - MADDPG, MADDPG, and DDPG algorithms as the number of training episodes increases are as Figure 6 shown. It can be seen from the figure that the average reward values of the algorithms all show an upward trend and finally tend to be stable, indicating that the agents continuously optimize their strategies through continuous learning and finally achieve convergence. Specifically, the GAT - MADDPG algorithm tends to be stable after 400 training episodes, and its convergence speed is better than that of the MADDPG algorithm that converges after 550 training episodes and the DDPG algorithm that converges after 600 training episodes. In addition, the reward value after the GAT - MADDPG algorithm converges is better than that of the MADDPG and DDPG algorithms because the graph attention network can effectively capture the state information of adjacent base stations and use valuable information to help the agents select better action strategies, thereby improving the average reward value and convergence speed. Although the MADDPG algorithm considers cooperation among agents, the agents cannot obtain the state information of adjacent nodes, resulting in inaccurate decisions and performance inferior to that of the GAT - MADDPG algorithm. The DDPG algorithm has the worst performance. The fundamental reason is that the agents adopt an independent learning mechanism, do not consider cooperation with adjacent agents, and cannot make full use of the information of other agents, resulting in low learning efficiency and thus the worst performance. In summary, compared with the comparison algorithms, the GAT - MADDPG algorithm has the best performance, thus verifying the effectiveness and superiority of the adopted algorithm.
[0152] The above-described embodiments have further elaborated on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-described embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A task offloading and resource management method for a hyperscale heterogeneous network oriented to the integration of communication and sensing, characterized in that Including: S1. Construct a system model; The system model includes: a macro base station, a small base station, and their associated Internet of Things (IoT) terminals; S2. Construct a service caching model, a resource partitioning model, a communication model, a sensing model, and a task offloading model according to the system model; S3. Construct a joint optimization problem based on the service caching model, the resource partitioning model, the communication model, the sensing model, and the task offloading model; S4. Transform the joint optimization problem into a Dec-POMDP; where Dec-POMDP is a distributed partially observable Markov decision process; S5. Solve the joint optimization problem using the GAT-MADDPG algorithm based on Dec-POMDP to obtain optimal service caching decisions, channel allocation decisions, transmission powers of IoT terminals, and task offloading decisions; where GAT is graph attention and MADDPG is a multi-agent deep deterministic policy gradient algorithm.
2. The method for task offloading and resource management of a hyperscale heterogeneous network for integrated communication, sensing, and computing according to claim 1, wherein Constructing a service caching model, a resource partitioning model, a communication model, a sensing model, and a task offloading model according to the system model includes: Building a service cache model includes: making service cache decisions where ψ c,m (t) ∈ {0, 1} indicates whether base station m caches service type c at time slot t, is the set of base stations, is the set of time, is the set of service types; Building a resource partitioning model includes: dividing the total spectrum resource W into the spectrum W of macro base stations MBS and the spectrum W of small base stations SBS ; according to the spectrum W of macro base stations MBS and the spectrum W of small base stations SBS build the available channel set of each base station Building a communication model includes: making a decision on the available channel allocation for each base station Among them, indicates whether base station m allocates channel n to IoT terminal k at time slot t m ; building the transmit power of each IoT terminal k m at time slot t Calculating the channel gain for each IoT terminal k m to transmit tasks to its base station m through channel n According to the channel allocation decision Transmit power And the channel gain Calculate the uplink transmission rate at which base station m receives tasks on channel n from IoT terminal k m Building a perception model includes using the conditional mutual information between the echo signal and the impulse response of the Internet of Things (IoT) terminal k m as the perception accuracy of the IoT terminal Building a task offloading model includes: making a task offloading decision Among them, indicates whether the task of the Internet of Things terminal k m is processed by the base station m, and is the set of Internet of Things terminals within the coverage of the base station m.
3. The method for task offloading and resource management in an ultra-dense heterogeneous network for communication, sensing, and computing integration according to claim 2, wherein The sensing accuracy of the IoT terminal is: where S is the Internet of Things (IoT) terminal k m the number of OFDM symbols for continuous integration, T s is the duration of a single OFDM symbol with a cyclic prefix, B represents the bandwidth of the channel, is the IoT terminal k m the radar signal-to-noise ratio on channel n, is the IoT terminal k m the impulse response on channel n at the frequency f of channel n n the Fourier transform at that point, represents from the IoT terminal k m to the IoT terminal k m′ the channel gain on channel n, are respectively the IoT terminal k m 、k m′ transmission powers, σ 2 is the noise, is the set of small base stations, is the set of IoT terminals of base station m'.
4. The method for task offloading and resource management in an ultra-dense heterogeneous network for communication, sensing, and computing integration according to claim 3, wherein The joint optimization problem is: Among them, the constraint C1 indicates that the task execution delay of the IoT terminal cannot exceed the maximum delay of task execution. The constraints C2 and C3 indicate that any IoT terminal can only be associated with one channel. The constraint C4 represents the constraint of channel allocation decision. The constraint C5 represents the upper and lower bounds of the transmission power of the IoT terminal. C6 represents the minimum requirement to ensure the sensing accuracy of the IoT terminal. The constraints C7 and C8 indicate that the IoT terminal can choose at most one base station to remotely process the task. The constraint C9 represents that the cached service cannot exceed the cache capacity of each base station; E km (t) is the total energy consumption of the task execution of IoT terminal k m . km (t) is the total task execution delay of IoT terminal k m . is the maximum tolerable delay of the task execution of IoT terminal k m . is the maximum transmission power of IoT terminal k m . is the minimum sensing accuracy of the IoT terminal, z c is the size of the storage space occupied by the cached service c, Z m is the size of the storage space of base station m.
5. The method for task offloading and resource management of a hyperscale heterogeneous network for integrated communication, sensing, and computing according to claim 4, wherein The total latency and total energy consumption of task execution are: Among them, is the total delay for the task to be executed by the macro base station, is the total delay for the task to be executed by the small base station, is the total execution energy consumption for the task to be executed by the macro base station, is the total execution energy consumption for the task to be executed by the small base station, is the total delay for the task to be executed by the local base station, is the total delay for the task to be executed by the neighboring base station, is the total execution energy consumption for the task to be executed by the local base station, is the total execution energy consumption for the task to be executed by the neighboring base station, is the service type required to execute the task.
6. The method for task offloading and resource management of a hyperscale heterogeneous network oriented to the integration of communication, sensing and computing according to claim 4, wherein, Transforming the joint optimization problem into a Dec-POMDP includes: Construct the global state space Global information Construct a partially observable space Local observation information Construct the action space Action Construct the reward function Indicates whether the agent's action meets the sensing accuracy requirements of the IoT terminal, Indicates whether the agent's action meets the latency requirements of the IoT terminal; Among them, represents the set of task information of all IoT terminals, is the set of task information of IoT terminals within the coverage of base station m, represents the set of channel gains between all base stations and all IoT terminals, g m (t) is the set of channel gains between base station m and all IoT terminals, H(·) is the Heaviside step function, and η1 and η2 are reward coefficients.
7. The method for task offloading and resource management of a hyperscale heterogeneous network for integrated communication, sensing, and computing according to claim 6, wherein Solving the joint optimization problem using the GAT-MADDPG algorithm based on Dec-POMDP includes: regarding each base station as an agent configured with an actor-critic network; the actor-critic network includes: an Actor network, a GAT-based Critic network, a target Actor network, and a GAT-based target Critic network; the agents learn using a distributed execution and centralized training architecture to obtain optimal service caching decisions, channel allocation decisions, and task offloading decisions; The agents learning using a distributed execution and centralized training architecture includes: During the distributed execution phase, each agent observes the global state space S at the current moment to obtain the local observation information o at the current moment m ; Each agent takes the local observation information o m as the input to its own Actor network to obtain the action a selected by each agent at the current moment m ; Each agent executes the action a it has selected m , to obtain the local observation information o' of each agent at the next moment m and the reward r m ; Combine the local observation information o m , action a m , local observation information o' at the next moment m and reward r m of all agents into an experience sample Store the experience sample in the experience replay pool; During the centralized training phase, sample experience samples from the experience replay pool Train the actor-critic network of each agent according to the sampled experience samples to obtain the trained actor-critic network of each agent; where u represents the serial number of the experience sample, and U represents the number of sampled experience samples; Obtaining optimal service caching decisions, channel allocation decisions, transmission powers of IoT terminals, and task offloading decisions according to the trained actor-critic network of each agent.
8. The method for task offloading and resource management in an ultra-dense heterogeneous network for communication, sensing, and computing integration according to claim 7, wherein Training the actor-critic network of the agents according to the sampled experience samples includes: Input the local observation information into the Actor network to obtain an action where π m is the Actor network; Apply action a u and local observation information to the GAT-based Critic network to obtain the action value function where is the GAT-based Critic network; Input the local observation information into the target Actor network to obtain the action a′ u ; Input the action a′ u and the local observation information into the GAT-based target Critic network to obtain the target action value function Update the Actor network according to the action a u and the action value function , and softly update the target Actor network according to the parameters of the updated Actor network; According to the action value function the target action value function and the reward function calculate the loss function value, update the GAT-based Critic network according to the loss function value, and softly update the GAT-based target Critic network according to the parameters of the updated GAT-based Critic network; When the loss function value is minimized, obtain the trained actor-critic network.
9. The method for task offloading and resource management of a hyperscale heterogeneous network for communication, sensing, and computing fusion according to claim 8, wherein The GAT-based Critic network includes: a GAT network and a Critic network; the GAT-based Critic network processes the action a u and local observation information and the processing includes: Modeling the multi-agent environment as an undirected graph G(V, E); where V represents the set of agents and E represents the set of edges between agents; When the agent m is the central node, according to the service cache status information ψ m′ of other agents m' at time (t - 1), the vector formed and the Internet of Things terminals k within the coverage of other agents m' m′ and the channel gain between the agent m form a vector to construct the node feature h m (t) of the agent m; where t represents the current time; Input the node feature h m (t) and the undirected graph G(V, E) into the GAT network to obtain the potential feature vector of the agent m The potential feature vector of agent m Action a u And task information Are input into the Critic network to obtain the action value function 10. The method for task offloading and resource management of a hyperscale heterogeneous network for communication, sensing, and computing integration according to claim 9, characterized in that, Construct the node feature h of the agent m m (t) includes: h m h(t) = αψ'(t - 1) + (1 - α)g'(t) m h(t) = αψ'(t - 1) + (1 - α)g'(t) m h(t) It should be noted that there may be some inaccuracies in the original formula expression. The correct expression in the translated content is adjusted according to the general mathematical formula writing norms. If there are specific requirements for the formula writing style, it may need to be further optimized according to the actual situation. α = σ(W α [ψ′ m (t - 1); g′ m (t)] + β α ) Among them, W ψ is the service cache feature projection matrix, W g is the channel gain feature projection matrix, W α is the weight projection matrix, β ψ 、β g 、β α are the service cache feature bias, channel gain feature bias, and weight bias respectively. MLP is a feed-forward neural network, Relu and Leaky Relu are activation functions, and α is the weight of the service cache node feature.
Citation Information
Cited By
Internet of vehicles multi-hop unloading and resource allocation method based on multi-agent reinforcement learning
CN122179841A