A Computation Offloading Method and Device for Integrated Communication and Sensing
By establishing relevant models on the terminal device and training, dynamically adjusting the channel and radio frequency transmission power, the problems of mutual interference between the sensing beam and the communication beam and dynamic changes in the channel conditions in the terminal device are solved, and the effectiveness of calculation offloading and the energy efficiency of the terminal device are realized.
Patent Information
- Application Number
- CN202210186961.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-02-28
AI Technical Summary
The perceived beam and the communication beam of the terminal device interfere with each other, and due to the mobility of the terminal, the conditions of the communication channel and the perceived channel are highly dynamically changed, making it difficult to realize computing offloading in the integrated communication and perception technology.
Establish a relevant model for computing offloading by the terminal, and use the terminal's to be computed tasks, uplink communication channel gain, perceived impulse response, and the angle difference between the communication beam and the perceived beam as inputs, and train and learn the related model to obtain the terminal's unloading parameters for the to be computed, including computing task offloading decisions and unloading RF transmission power decisions.
By dynamically adjusting the channel and RF transmission power, an effective trade-off between network-level perception accuracy and calculation task processing time is achieved, ensuring the timeliness of communication data and perceived data of the terminal, while reducing the terminal's data processing pressure and energy consumption.
Smart Images

Figure CN114554548B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to a computing offloading method and device for communication and sensing integration. Background Art
[0002] The high-frequency bands used by wireless communication networks are gradually approaching and even overlapping with the radio sensing bands. Communication and sensing integration technologies have been introduced in different scenarios (such as intelligent transportation and smart factories), enabling wireless networks to provide high-precision positioning services while realizing wireless communication functions, effectively improving resource utilization and information processing efficiency. In addition, with the continuous evolution and development of new services such as autonomous driving and immersive extended reality, due to the limitation of terminal computing capabilities, it is impossible to process a large amount of data in a short time. Therefore, mobile edge computing technology has emerged, and the terminal can offload computing tasks to the mobile edge through the uplink communication beam, effectively reducing the energy consumption of the terminal. However, there is mutual interference between the sensing beam and the communication beam of the terminal, and at the same time, due to the mobility of the terminal device, the communication channel and the sensing channel conditions change highly dynamically, resulting in the difficulty of realizing computing offloading in communication and sensing integration technology. Summary of the Invention
[0003] The purpose of the present invention is to provide a computing offloading method and device for communication and sensing integration, so as to solve the problem in the prior art that there is mutual interference between the sensing beam and the communication beam of the terminal, and at the same time, due to the mobility of the terminal, the communication channel and the sensing channel conditions change highly dynamically, resulting in the difficulty of realizing computing offloading in communication and sensing integration technology.
[0004] To achieve the above purpose, an embodiment of the present invention provides a computing offloading method for communication and sensing integration, including:
[0005] Establish a relevant model for the terminal for computing offloading;
[0006] Taking the computing task to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as inputs, training and learning the relevant model to obtain the offloading parameters for the computing task to be calculated by the terminal; the computing task includes the communication data to be calculated and the sensing data to be calculated, and the offloading parameters include the computing task offloading decision and the offloading radio frequency transmission power decision;
[0007] Offload the computing task to be calculated to the edge side according to the offloading parameters.
[0008] Further, the relevant model includes:
[0009] An interference model between beams, a sensing model, an uplink communication model of the terminal, a task model of the computing task to be calculated, a computing model of the terminal, and a computing model of the edge side.
[0010] Further, the inter-beam interference model is a sector antenna model, which is used to characterize the beam interference gain between the communication beam and the sensing beam.
[0011] Further, the sensing model of the terminal includes:
[0012] A model for the terminal to send orthogonal frequency division multiplexing sensing signals and a model for the echo signals received by the edge side, and the output of the sensing model is the conditional mutual information between the target impulse response and the received signal.
[0013] Further, the uplink communication model of the terminal is used to characterize the uplink transmission rate and the uplink communication channel gain between the terminal and the edge side.
[0014] Further, the task model is used to characterize the number of the to-be-computed tasks within a preset time period, the size of each to-be-computed task, the number of CPU cycles required to execute one bit of the to-be-computed task, and the delay threshold of each to-be-computed task.
[0015] Further, the computing model of the terminal is used to characterize the delay of the terminal in processing the to-be-computed task and the sensing performance of the terminal.
[0016] Further, the computing model of the edge side is used to characterize the delay of the edge side in processing the to-be-computed task and the sensing performance of the terminal;
[0017] Among them, the delay of the edge side in processing the to-be-computed task includes: the first delay for the to-be-computed task to be transmitted from the terminal to the edge side, the second delay for the edge side to complete processing the to-be-computed task, and the third delay for the edge side to send the processed to-be-computed task to the terminal.
[0018] Further, the correlation model further includes a joint optimization model;
[0019] The joint optimization model is an optimization model regarding the delay of processing the to-be-computed task and the sensing performance of the terminal;
[0020] The joint optimization model includes: constraint conditions for the total task running delay, constraint conditions for the sensing performance of the terminal, constraint conditions for the radio frequency transmission power of the communication beam of the terminal, constraint conditions for the radio frequency transmission power of the sensing beam of the terminal, constraint conditions for the relationship between the sensing beam and the communication beam, and constraint conditions for the offloading parameters.
[0021] Further, training and learning the correlation model to obtain the offloading parameters of the terminal for the to-be-computed task includes:
[0022] Taking the computing task to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as the state space, and the computing task offloading decision and the offloading radio frequency transmission power decision as the action space, and establishing a reward function to obtain the offloading parameters;
[0023] Among them, taking the computing task to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, the communication beam, the angle difference between the sensing beam, and the reward function as inputs, through a reinforcement learning module based on Multi-DQN, the computing offloading decision in the task offloading strategy is obtained;
[0024] Taking the computing task to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, the communication beam, the angle difference between the sensing beam, and the reward function as inputs, through a reinforcement learning module based on TD3, the radio frequency transmission power decision in the task offloading strategy is obtained.
[0025] The beneficial effects of the above technical solutions of the present invention are as follows:
[0026] The communication-aware integrated computing offloading method according to the embodiment of the present invention establishes a relevant model for the terminal for computing offloading, and takes the computing task to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as inputs, and trains and learns the relevant model to obtain the offloading parameters for the computing task to be calculated by the terminal; the computing task includes the communication data to be calculated and the sensing data to be calculated, and the offloading parameters include the computing task offloading decision and the offloading radio frequency transmission power decision; according to the offloading parameters, offloading the computing task to be calculated to the edge side can ensure the timeliness of the terminal's processing of communication data and sensing data while reducing the terminal's data processing pressure and energy consumption. Description of the Drawings
[0027] Figure 1 It is a schematic diagram of the communication-aware integrated computing offloading method according to the embodiment of the present invention;
[0028] Figure 2 It is one of the schematic diagrams of the beam interference of the terminal antenna according to the embodiment of the present invention;
[0029] Figure 3 It is another schematic diagram of the beam interference of the terminal antenna according to the embodiment of the present invention;
[0030] Figure 4 It is a schematic diagram of the training and learning steps of the communication-aware integrated computing offloading method according to the embodiment of the present invention. Detailed Embodiments
[0031] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0032] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.
[0033] In various embodiments of the present invention, it should be understood that the magnitudes of the serial numbers of the following processes do not mean the order of execution is prior or subsequent. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0034] In addition, the terms "system" and "network" are often used interchangeably in this article.
[0035] In the embodiments provided in this application, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0036] As Figure 1 shown, a communication-aware integrated computing offloading method according to an embodiment of the present invention includes the following steps:
[0037] Step 101, establish a relevant model for the terminal for computing offloading.
[0038] Optionally, the establishment of the relevant model for the terminal for computing offloading needs to be established in a wireless network environment based on communication-aware integrated technology. Define the key elements in the wireless network environment based on communication-aware integrated technology, that is, the set of terminal devices I = {1, 2,..., i}, and the set of time slots K = {1, 2,..., k}. At the same time, define the total radio frequency transmission power of each terminal device as:
[0039] where, is the uplink communication beam transmission power of terminal device i, is the sensing beam transmission power of terminal device i. In an embodiment of the present invention, during the movement of the terminal device, a forward horizontal sensing beam and an uplink communication beam are simultaneously transmitted.
[0040] Step 102: Use the to-be-computed task of the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as inputs to train and learn the relevant model, and obtain the offloading parameters of the terminal for the to-be-computed task; the computing task includes the to-be-computed communication data and the to-be-computed sensing data, and the offloading parameters include the computing task offloading decision and the offloading radio frequency transmission power decision.
[0041] Optionally, use the set s i (k) of the to-be-computed tasks, the uplink communication channel gain the sensing impulse response and the angle difference Δζ i (k) between the communication beam and the sensing beam as inputs, train the relevant model in a wireless network environment based on communication-sensing integrated technology, and use the respective deep neural networks and embedded algorithms of the reinforcement learning module based on Multi-DQN and the reinforcement learning module based on TD3 for training.
[0042] In an embodiment of the present invention, the computing task offloading decision includes the communication data and sensing data in the to-be-computed task that need to be offloaded to the edge side, and also includes the communication channel and sensing channel conditions for transmitting the communication data and sensing data to the edge side. According to the computing task offloading decision
[0043] Step 103: Offload the to-be-computed task to the edge side according to the offloading parameters.
[0044] In an embodiment of the present invention, according to the offloading parameters, determine the target to-be-computed task that the terminal needs to offload to the edge side, and the target to-be-computed task includes the to-be-computed communication data and the to-be-processed sensing data; according to the offloading parameters, determine the radio frequency generating power for transmitting the target to-be-computed task.
[0045] The communication-sensing integrated computing offloading method of the embodiments of the present invention can, by establishing a relevant model for the terminal for computing offloading, using the to-be-computed task of the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as inputs to train and learn the relevant model, obtain the offloading parameters of the terminal for the to-be-computed task; the computing task includes the to-be-computed communication data and the to-be-computed sensing data, and the offloading parameters include the computing task offloading decision and the offloading radio frequency transmission power decision; and offload the to-be-computed task to the edge side according to the offloading parameters, ensure the timeliness of the terminal's processing of communication data and sensing data while reducing the terminal's data processing pressure and energy consumption.
[0046] Optionally, the relevant model includes:
[0047] The inter-beam interference model, the sensing model, the uplink communication model of the terminal, the task model of the task to be calculated, the computing model of the terminal, and the computing model of the edge side.
[0048] In the communication and sensing integrated computing offloading method according to an embodiment of the present invention, by establishing relevant models for the terminal's computing offloading in a communication and sensing integrated wireless network environment and training and learning the relevant models, the offloading parameters of the task to be calculated are obtained. It can dynamically adjust the channel and the radar beam transmission power of the terminal according to the time-varying communication channel and sensing channel, so as to achieve an effective trade-off between the network-level sensing accuracy and the computing task processing time.
[0049] Optionally, the inter-beam interference model is a sector antenna model, which is used to characterize the beam interference gain between the communication beam and the sensing beam.
[0050] In an embodiment of the present invention, the inter-beam interference model can be approximately expressed as a sector antenna model, that is, for the integrated communication and sensing system carried by the terminal device, the antenna gains of the transmitter and the receiver can be respectively described as:
[0051] And
[0052]
[0053] Where And Are respectively the main lobe gains of the transmitter and receiver antennas, And Are respectively the side lobe gains of the transmitter and receiver antennas. And Are the transmission beam width and the reception beam width.
[0054] As Figure 2 Shown, the dotted line represents the communication beam, and the solid line represents the sensing beam. When , the beam interference between the antennas can be expressed as:
[0055]
[0056] Where Is the overlapping angle between the communication beam and the sensing beam, And Are the offset angles of the communication beam and the sensing beam with respect to the horizontal direction. At this time, the beam interference is divided into two parts: main lobe interference and side lobe interference.
[0057] As Figure 3 Shown, the dotted line represents the communication beam, and the solid line represents the sensing beam. When , the beam interference between the antennas can be expressed as: At this time, the interference between beams is only the interference between sidelobes.
[0058] In the solution of the present invention, by determining the beam interference gain between the communication beam and the sensing beam, the radar beam transmission power of the terminal is dynamically adjusted.
[0059] Optionally, the sensing model of the terminal includes:
[0060] The model of the orthogonal frequency division multiplexing sensing signal sent by the terminal and the model of the echo signal received on the edge side, and the output of the sensing model is the conditional mutual information between the target impulse response and the received signal.
[0061] In an embodiment of the present invention, the performance of radar detection, that is, the sensing performance of the terminal, is measured by the conditional mutual information between the target impulse response and the received signal. The greater the conditional mutual information, the smaller the degree of reduction in the uncertainty of the a priori after measurement. Therefore, the parameters of the detection target can be accurately estimated.
[0062] The OFDM (Orthogonal Frequency Division Multiplexing) orthogonal frequency division multiplexing sensing signal sent by the terminal device i at time slot k can be expressed as:
[0063]
[0064] where f c is the center frequency, N s is the number of consecutive OFDM symbols sent per time slot, N c is the number of subcarriers, is the amplitude of the OFDM signal on subcarrier n, is the signal amplitude on subcarrier n. c n,l (k) is the phase encoding of the terminal device i on symbol l of subcarrier n, T = 1 / Δf is the duration of the basic OFDM symbol. In addition, T s is the duration of the complete symbol, and it satisfies T s = T + T g , T g is the cyclic prefix, rect[t / T s is the rectangular function.
[0065] The sensing target impulse response g n (t) can be expressed as a Gaussian random process. Therefore, the echo signal received by the base station can be expressed as:
[0066]
[0067] where, is additive white Gaussian noise with a power spectral density of At this time, for the terminal device i, the conditional mutual information between the perceived target impulse response and the echo signal can be expressed as:
[0068]
[0069] where T p = N s T s is the duration of the OFDM signal, and f n = f c + nΔf is the frequency of subcarrier n. is the value of the Fourier transform in time slot k. In addition,
[0070] Furthermore, the conditional mutual information (mutual information rate) per unit time is defined in this method as:
[0071]
[0072] The solution of the present invention determines the sensing performance of the terminal according to the conditional mutual information between the target impulse response and the received signal, so that while ensuring the communication data transmission and processing of the terminal, the sensing performance of the terminal is ensured.
[0073] Optionally, the uplink communication model of the terminal is used to characterize the uplink transmission rate and the uplink communication channel gain between the terminal and the edge side.
[0074] In an embodiment of the present invention, the terminal device uploads the computing task to the mobile edge server tightly coupled with the edge side base station through orthogonal frequency division multiple access. According to the Shannon formula, the uplink transmission rate between the terminal device i and the edge side base station can be expressed as:
[0075]
[0076] where B is the uplink bandwidth allocated to each terminal device, and σ 2 is the noise power of the communication channel. is the time-varying communication channel gain. In addition, is the main lobe gain of the antenna of the edge side base station and can be expressed as:
[0077]
[0078] where is the beam width of the receiver of the edge side base station, and the path loss (dB) between the terminal device i and the edge side base station is:
[0079] L(d i (k)) = 40(1 - 4×10-3 D hb )log 10 d i (k)-18log 10 D hb +21log 10 f(k)+80;
[0080] Wherein, f(k) is the carrier frequency, and D hb is the height of the antenna. At this time, the uplink communication channel gain is:
[0081]
[0082] In an embodiment of the present invention, the channel information of the communication data and the sensing data can be randomly assigned through a function, and the function can be a random function and a Gaussian function.
[0083] Optionally, the task model is used to characterize the number of the to-be-computed tasks, the size of each to-be-computed task, the number of CPU cycles required to execute one bit of the to-be-computed task, and the delay threshold of each to-be-computed task within a preset time period.
[0084] In an embodiment of the present invention, each terminal has a batch of tasks (such as the original sensing information to be processed) that need to be processed at time slot k. At this time, this batch of tasks can be further described as a triple s i (k)=n i (k)l i (k) is the size of these tasks, n i (k) is the number of tasks to be processed at time slot k, l i (k) is the size of each task, c i (k) is the number of CPU cycles required to execute one bit of the task, is the maximum delay that can be tolerated for processing these tasks.
[0085] In an embodiment of the present invention, all the to-be-computed tasks of the terminal, the size of each to-be-computed task, the number of CPU cycles required to execute one bit of the to-be-computed task, and the delay threshold of each to-be-computed task are obtained, and then the target to-be-computed tasks to be offloaded to the edge side are determined according to the computing power of the terminal.
[0086] Optionally, the computing model of the terminal is used to characterize the delay of the terminal in processing the to-be-computed tasks and the sensing performance of the terminal.
[0087] In addition, β i (k) is the computing offloading decision factor. When β iWhen β(k)=1, the computing task is offloaded to the mobile edge server for processing, while when β i (k)=0, the task is processed on the terminal device side. At time slot k, all the computing tasks to be calculated are processed on the terminal. The computing model of terminal device i (β i (k)=0) mainly consists of one stage: the local server on the terminal device processes the computing task. At this time, the processing delay of the computing task on the terminal device side is:
[0088]
[0089] Among them, is the CPU frequency of the local server (CPU cycles per second). When the tasks are all processed on the local server, the radio frequency transmission power of terminal device i is all used for sensing. At this time, the mutual information rate used by terminal device i to measure the sensing performance can be expressed as:
[0090]
[0091] Among them,
[0092] Optionally, the computing model on the edge side is used to characterize the processing delay of the computing task to be calculated on the edge side and the sensing performance of the terminal;
[0093] Among them, the processing delay of the computing task to be calculated on the edge side includes: the first delay for the computing task to be transmitted from the terminal to the edge side, the second delay for the edge side to complete the processing of the computing task to be calculated, and the third delay for the edge side to send the processed computing task to the terminal.
[0094] In an embodiment of the present invention, at time slot k, the computing model of terminal device i (β i (k)=1) mainly consists of three stages: terminal device i uploads the computing task to the edge side base station; the computing task is processed on the mobile edge server side; the processing result is sent down to terminal device i. At this time, the processing delay for the computing task from terminal device i is:
[0095]
[0096] Among them, is the uplink transmission delay of the computing task; is the processing delay of the computing task on the mobile edge server; is the downlink transmission delay of the computing task. Since the processed result information is relatively small, the downlink transmission delay can be ignored.
[0097] Furthermore, the first delay can be expressed as:
[0098]
[0099] The second time delay can be expressed as:
[0100]
[0101] where is the CPU frequency of the mobile edge server (CPU cycles per second). At this time, a part of the radio frequency power of the terminal device i is used to upload the computing task, and the other part is used for sensing. Therefore, the mutual information rate used by the terminal device i to measure the sensing performance can be expressed as:
[0102]
[0103] where
[0104] Optionally, the correlation model further includes a joint optimization model;
[0105] The joint optimization model is an optimization model for the time delay of processing the to-be-computed task and the sensing performance of the terminal;
[0106] The joint optimization model includes: a constraint condition for the total time delay of task operation, a constraint condition for the sensing performance of the terminal, a constraint condition for the radio frequency transmission power of the communication beam of the terminal, a constraint condition for the radio frequency transmission power of the sensing beam of the terminal, a constraint condition for the relationship between the sensing beam and the communication beam, and a constraint condition for the offloading parameter.
[0107] The objective of the embodiment of the present invention is to jointly optimize the processing time delay of the computing task and the sensing performance of the terminal device. Therefore, the following joint optimization model is constructed:
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115] where τ(k) is the total processing time delay of the computing task at the network level and can be expressed as:
[0116]
[0117] I(k) is the total mutual information rate at the network level and can be expressed as:
[0118]
[0119] α(k) and δ(k) are the weights of the total task processing delay and the total mutual information rate respectively, and α(k)+δ(k)=1. τ max is the maximum total task processing delay, and I max is the maximum mutual information rate. C1 indicates that the total task running delay cannot exceed the maximum constraint delay; C2 restricts the minimum sensing performance of the terminal device; C3 and C4 constrain the transmission powers of the sensing beam and the communication beam of the terminal device; C5 gives the relationship between the sensing beam and the communication beam; C6 restricts the duality of the offloading decision variable.
[0120] The solution of the present invention constructs a joint optimization model, takes the processing delay of the computing task and the sensing performance of the terminal as optimization metrics, and obtains the offloading decision and the radio frequency generation power for the to-be-computed task while ensuring the task processing delay and the sensing performance of the terminal.
[0121] Optionally, training and learning the related model to obtain the offloading parameters of the terminal for the to-be-computed task includes:
[0122] Taking the to-be-computed task of the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as the state space, the computing task offloading decision and the offloading radio frequency transmission power decision as the action space, and establishing a reward function to obtain the offloading parameters;
[0123] Among them, taking the to-be-computed task of the terminal, the uplink communication channel gain, the sensing impulse response, the communication beam, the angle difference between the sensing beam, and the reward function as inputs, through the reinforcement learning module based on Multi-DQN, the computing offloading decision in the task offloading strategy is obtained;
[0124] Taking the to-be-computed task of the terminal, the uplink communication channel gain, the sensing impulse response, the communication beam, the angle difference between the sensing beam, and the reward function as inputs, through the reinforcement learning module based on TD3, the radio frequency transmission power decision in the task offloading strategy is obtained.
[0125] In an embodiment of the present invention, the optimization model is transformed into a Markov decision process to realize data interaction with different modules. The relevant elements of this Markov decision process are specifically represented as follows:
[0126] (1) State space: At each time slot k, there is a batch of computing tasks to be processed on the terminal device side. At the same time, due to the continuous movement of the terminal device, the mobility of the terminal device can be further mapped into a time-varying uplink communication channel, a time-varying sensing target impulse response, and a time-varying angle difference between the communication beam and the sensing beam. Specifically expressed as:
[0127]
[0128] (2) Action space: At each time slot k, there are two types of decisions in this method, namely the computing offloading decision β i (k), and the transmit power decision That is:
[0129]
[0130] To better match the reinforcement learning module based on Multi-DQN and the reinforcement learning module based on TD3, this method further divides the action space A into A = A o ∪A p , where
[0131] (3) Reward: Based on the joint optimization model, the reward of this Markov decision process can be expressed as:
[0132]
[0133] As Figure 4 shown, in an embodiment of the present invention, the establishment of the relevant model in the communication-aware integrated computing offloading method is based on the wireless network environment based on communication-aware integration. The experience cache pool module is mainly to overcome the correlation and non-stationarity problems of the experience data in the reinforcement learning network training process. The specific operation is to store the data tuples (s k , a k , r k , s k+1 ) from the environment module at each time slot into the cache pool, and at the same time, the data that was first added to the cache pool will be removed; in each training, N batches of data will be sampled.
[0134] The two-level deep reinforcement learning policy module is composed of a reinforcement learning sub-module based on Multi-DQN and a reinforcement learning sub-module based on TD3. As Figure 4 shown, the two reinforcement learning modules simultaneously accept N batches of data from the experience cache pool module as input, and use their respective deep neural networks and embedded algorithms for training. After the model is trained, the optimal decision can be achieved.
[0135] (1) Reinforcement learning sub-module based on Multi-DQN
[0136] The action space in the Markov decision process of the environment module is Among them, the computing offloading decision β i (k) is a discrete action. Therefore, the DQN algorithm is embedded in this module to output the optimal computing offloading decision. In an actual wireless network based on the integrated sensing and communication technology, the density of terminal devices is large, and the action space of a single DQN network architecture will increase exponentially with the traffic flow. Therefore, to reduce the dimension of the action space, this module introduces the Multi-DQN architecture, that is, the discrete action of each terminal device is trained and output by a corresponding DQN network. Each DQN network adopts the ε-greedy strategy, that is, with a probability of ε, a random action is selected from the action space, or with a probability of 1 - ε, the highest Q value is selected, which is expressed as:
[0137]
[0138] Among them, ε ∈ (0, 1), θ i is the weight of the i-th DQN unit, At this time, this action and the action output by the reinforcement learning module based on TD3 are input to the environment module together, and then the environment module transfers from s k to s k+1 . Further, based on the randomly sampled data tuple (s k , a k , r k , s k+1 ) provided by the experience buffer pool, the target Q value generated by the target Q network of the i-th DQN unit is:
[0139]
[0140] Among them, is the weight of the target Q network in the i-th DQN unit. At this time, the Q network in the i-th DQN unit can be trained by minimizing the following loss function:
[0141]
[0142] The weight in the target Q network of the i-th DQN unit can be updated by copying the parameters of the Q network every G time slots.
[0143] (2) Reinforcement learning sub-module based on TD3
[0144] The action space in the Markov decision process of the environment module is Among them, the transmit power decision β i(k) is a continuous action, so this module embeds the TD3 algorithm to output the optimal transmission power decision.
[0145] In this module, six deep neural networks are introduced. Among them, three are evaluation networks (one is the evaluation actor network to output actions, and the other two are evaluation critic networks to evaluate Q-values), and the other three are target networks (one is the target actor network, and two are target critic networks to generate target values for the corresponding evaluation networks). μ is the weight of the evaluation actor network, and η1 and η2 are the weights of the evaluation critic networks respectively; is the weight of the target actor network, and are the weights of the target critic networks respectively.
[0146] This module adopts the actor-critic method, that is, the evaluation actor network outputs the deterministic transmission power decision based on the deterministic policy gradient theory while the evaluation critic network evaluates the output actions. In addition, to explore new actions in the action space, Gaussian noise is introduced as follows:
[0147]
[0148] At this time, this action and the action output by the reinforcement learning module based on TD3 are input to the environment module together, and then the environment module transfers from s k to s k+1 . According to then the two evaluation critic networks can output the corresponding Q-values Q(s,a,η1) and Q(s,a,η2). Therefore, the weight μ can be updated according to the following formula:
[0149]
[0150] Based on the randomly sampled data tuples (s k ,a k ,r k ,s k+1 ) provided by the experience buffer pool, the above formula can be approximated as:
[0151]
[0152] The target Q value can be updated according to the target actor network and the target critic network:
[0153]
[0154] The loss functions of the two evaluation critic networks can be updated according to the following formula:
[0155]
[0156]
[0157] Slowly update the weights based on the following formula and to improve the stability of the training process:
[0158]
[0159]
[0160]
[0161] Furthermore, it should be noted that the terminals described in this specification include, but are not limited to, smart phones, tablet computers, etc., and many of the functional components described are called modules to more particularly emphasize the independence of their implementation methods.
[0162] In the embodiments of the present invention, the module can be implemented by software so as to be executed by various types of processors. For example, an identified executable code module can include one or more physical or logical blocks of computer instructions. For example, it can be constructed as an object, a process, or a function. Nevertheless, the executable code of the identified module does not need to be physically located together, but can include different instructions stored in different locations. When these instructions are logically combined together, they constitute the module and achieve the specified purpose of the module.
[0163] When the module can be implemented by software, considering the level of existing hardware technology, for the modules that can be implemented by software, without considering the cost, those skilled in the art can build the corresponding hardware circuits to implement the corresponding functions. The hardware circuits include conventional very large scale integration (VLSI) circuits or gate arrays, as well as existing semiconductors such as logic chips, transistors, or other discrete components. The module can also be implemented by programmable hardware devices, such as field programmable gate arrays, programmable array logic, programmable logic devices, etc.
[0164] The above exemplary embodiments are described with reference to these drawings. Many different forms and embodiments are possible without departing from the spirit and teachings of the present invention. Therefore, the present invention should not be construed as being limited to the exemplary embodiments presented herein. Rather, these exemplary embodiments are provided so that the present invention will be complete and full, and will convey the scope of the present invention to those skilled in the art. In these drawings, component sizes and relative sizes may be exaggerated for clarity. The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. As used herein, unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are intended to include the plural forms as well. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Unless otherwise indicated, when stating a value range, the range includes the upper and lower limits thereof and any sub-ranges therebetween.
[0165] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention described above, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A communication-aware integrated computing offloading method, characterized in that, Including: Establish relevant models for the terminal's computing offloading; Taking the computing tasks to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as inputs, training and learning the relevant models to obtain the offloading parameters for the terminal's computing tasks to be calculated; the computing tasks include communication data to be calculated and sensing data to be calculated, and the offloading parameters include computing task offloading decisions and offloading radio frequency transmission power decisions; Offloading the computing tasks to be calculated to the edge side according to the offloading parameters; The relevant models include: Inter-beam interference model, sensing model, uplink communication model of the terminal, task model of the computing tasks to be calculated, computing model of the terminal, computing model of the edge side, and joint optimization model; The inter-beam interference model is a sector antenna model, which is used to characterize the beam interference gain between the communication beam and the sensing beam; The sensing model of the terminal includes: The model of the orthogonal frequency division multiplexing sensing signal sent by the terminal and the model of the echo signal received by the edge side, and the output of the sensing model is the conditional mutual information between the target impulse response and the received signal; The uplink communication model of the terminal is used to characterize the uplink transmission rate and uplink communication channel gain between the terminal and the edge side; The task model is used to characterize the number of computing tasks to be calculated within a preset time period, the size of each computing task to be calculated, the number of CPU cycles required to execute one bit of the computing task to be calculated, and the delay threshold of each computing task to be calculated; The computing model of the terminal is used to characterize the delay of the terminal in processing the computing tasks to be calculated and the sensing performance of the terminal; The computing model of the edge side is used to characterize the delay of the edge side in processing the computing tasks to be calculated and the sensing performance of the terminal; Among them, the delay of the edge side in processing the computing tasks to be calculated includes: the first delay of the computing tasks to be calculated transmitted from the terminal to the edge side, the second delay of the edge side in completing the processing of the computing tasks to be calculated, and the third delay of the edge side in sending the processed computing tasks to be calculated to the terminal; The joint optimization model is an optimization model regarding the delay of processing the computing tasks to be calculated and the sensing performance of the terminal; The joint optimization model includes: constraint conditions for the total task running delay, constraint conditions for the sensing performance of the terminal, constraint conditions for the radio frequency transmission power of the terminal's communication beam, constraint conditions for the radio frequency transmission power of the terminal's sensing beam, constraint conditions for the relationship between the sensing beam and the communication beam, and constraint conditions for the offloading parameters.
2. The communication-aware integrated computing offloading method according to claim 1, characterized in that, The training and learning of the relevant models to obtain the offloading parameters for the terminal's computing tasks to be calculated includes: Taking the computing tasks to be calculated by the terminal, the uplink communication channel gain, the sensing impulse response, and the angle difference between the communication beam and the sensing beam as the state space, the computing task offloading decision and the offloading radio frequency transmission power decision as the action space, and establishing a reward function to obtain the offloading parameters; Among them, taking the task to be calculated of the terminal, the uplink communication channel gain, the sensing impulse response, the communication beam, the angular difference between the sensing beam, and the reward function as inputs, through a reinforcement learning module based on Multi-DQN, the calculation offloading decision in the task offloading strategy is obtained; Taking the task to be calculated of the terminal, the uplink communication channel gain, the sensing impulse response, the communication beam, the angular difference between the sensing beam, and the reward function as inputs, through a reinforcement learning module based on TD3, the radio frequency transmission power decision in the task offloading strategy is obtained.