Edge-oriented computation offloading method based on user scheduling for cell-free networks
By dividing orthogonal sub-channels and dynamically scheduling user equipment in decellularized networks, and combining deep reinforcement learning to optimize user scheduling and resource allocation, the offloading delay problem caused by multi-user interference and pilot pollution in MEC-assisted decellularized networks is solved, achieving latency reduction and improved resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2026-07-03
AI Technical Summary
In MEC-assisted decellularized networks, multi-user interference and pilot pollution increase computation offloading latency, which is unfair to user devices with heavy computational loads and poor channel conditions, especially in IIoT scenarios. Existing technologies have failed to effectively optimize this.
By dividing the uplink channel into multiple orthogonal sub-channels, dynamically scheduling user equipment for offloading, and combining deep reinforcement learning training to optimize user scheduling and computing resource allocation, a JCSCS model is constructed to mitigate the impact of multi-user interference and pilot pollution.
It significantly reduces offload latency, improves fairness and system performance for user devices, reduces total latency, and makes full use of the computing resources of the MEC server.
Smart Images

Figure CN122340552A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technology in the field of wireless communication, specifically a computational offloading method based on user scheduling for edge decellularization networks. Background Technology
[0002] In traditional 5G cellular network architectures, Industrial Internet of Things (IIoT) User Equipment (UEs) at the cell edge face challenges such as inter-cell interference and low-reliability transmission links, which limit the performance of Mobile Edge Computing (MEC). Compared to cellular MEC networks, decellularized networks can suppress inter-cell interference by having multiple distributed access points (APs) collaboratively process UE signals, significantly reducing latency caused by compute offloading and supporting more stringent latency requirements. While cell edge issues are mitigated in decellularized MEC networks, multi-user interference and pilot pollution remain major factors limiting system performance. Pilot pollution, in particular, due to the limited availability of orthogonal pilot resources, can significantly increase offload latency.
[0003] Most studies on computational offloading in MEC-assisted decellularized networks either ignore the offloading delay caused by pilot contamination or merely model it formally without optimization. Recently, segmenting frequency bands and controlling the number of users scheduled on sub-channels has proven effective in mitigating pilot contamination in decellularized downlink networks. However, this approach has not been studied in MEC-assisted decellularized uplink networks. Most related work lacks a user scheduling strategy during offloading, instead scheduling all UEs simultaneously across the entire frequency band. In IIoT offloading scenarios, this approach significantly increases offloading delay. Furthermore, this method is unfair to IIoT user equipment with high computational demands and poor channel conditions. Summary of the Invention
[0004] To address the aforementioned problems in existing technologies, this invention proposes a computational offloading method based on user scheduling for edge decellularization networks. This method divides the uplink channel into multiple orthogonal sub-channels and dynamically schedules user equipment to the sub-channels for offloading, thereby mitigating the impact of multi-user interference and pilot pollution on offloading delay.
[0005] This invention is achieved through the following technical solution:
[0006] This invention relates to a computation offloading method based on user scheduling. After modeling a MEC-assisted decellularized communication model and a parallel computation offloading model to construct an uplink network simulation environment, a joint optimization computation offloading, user scheduling, and computation resource allocation strategy (JCSCS) model considering limited orthogonal pilot resources and edge computing resources is constructed and trained in the uplink network simulation environment using deep reinforcement learning based on near-end policy optimization (PPO). Finally, the optimal computation offloading strategy, user scheduling strategy, and computation resource allocation strategy are obtained through PPO.
[0007] The JCSCS model has a network structure with multiple action outputs, including: an environment module, a policy generation module, a policy evaluation module, and an experience replay module. The environment encoding module generates feature vectors based on state information reflecting the characteristics of the environment. The policy generation module outputs a policy consisting of a set of multiple future actions based on the feature vectors. The policy evaluation module evaluates and estimates the current policy and feeds it back to the policy generation module to adjust its policy model. The experience replay module is used to improve the model training efficiency and performance.
[0008] The uplink network simulation environment described herein is implemented using, but is not limited to, Python. Technical effect
[0009] Compared to traditional MEC-assisted decellularization uplink methods that offload computation across the entire channel and schedule all user equipment simultaneously without sub-channel division or specific user scheduling algorithms, this invention fully considers the impact of multi-user interference and pilot contamination on offload delay. It divides the channel into multiple orthogonal sub-channels and dynamically schedules user equipment for offload on these sub-channels. Experimental results show that, compared to the computational offload method used in traditional MEC-assisted decellularization uplinks, this invention significantly reduces offload delay caused by multi-user interference and pilot contamination, substantially reduces the total delay for all user equipment, and ensures a certain degree of user fairness. Attached Figure Description
[0010] Figure 1 This is a flowchart of the present invention;
[0011] Figure 2 Schematic diagram of a cellless environment assisted by MEC;
[0012] Figure 3 The network architecture diagram of the JCSCS model;
[0013] Figure 4 A graph showing the relationship between the number of users and the maximum latency under different algorithms;
[0014] Figure 5 A graph showing the relationship between total computing resources and maximum latency under different algorithms;
[0015] Figure 6 This is a graph showing the relationship between bandwidth and maximum latency under different algorithms. Detailed Implementation
[0016] like Figure 1 As shown in this embodiment, a computation offloading method based on user scheduling for edge decellularization networks includes the following steps:
[0017] Step 1, construct as follows Figure 2 The model shown is a MEC-assisted decellular communication model and a parallel computing offloading model in an edge decellular network simulation scenario containing M single-antenna user equipment (UE) and N single-antenna access points (AP).
[0018] The aforementioned decellular communication model includes the channel coefficients estimated during the uplink training phase and the closed-loop uplink rate during the uplink data transmission phase, specifically:
[0019] A) Channel coefficient calculation during the uplink training phase: The system bandwidth W is uniformly divided into K orthogonal sub-channels. A fully synchronized time-slot system is considered, where each time slot t (t≥0) represents a user scheduling period. The wireless time-varying channel fading model used in the training phase is as follows: Where: β m,n This represents the large-scale fading factor, which includes path loss and shadow fading. This represents the small-scale fading of user equipment m (0≤m≤M) at access point n (0≤n≤N) in subchannel k (0≤k≤K) in time slot t, which follows the Z-state Markov model.
[0020] The transition matrix of the Markov model is: Each element q in the matrix zz′ This represents the transition probability from channel state z to z′. During uplink training, UEs simultaneously transmit a predetermined pilot sequence (of length τ) consisting of a set of orthogonal vectors to APs through each sub-channel. p Each UE on a sub-channel has τ p (τ p ≥0) orthogonal pilot sequences can be selected, using This represents user equipment scheduling variables. Specifically, This indicates that the m-th UE will be scheduled to subchannel k in time slot t, and vice versa. Due to the finite coherence time interval τ... c The number of available orthogonal pilot sequences τ p Typically, the number of UEs is less than M, which requires some UEs to reuse the same pilot sequence, leading to pilot pollution. The uplink signal matrix received by access point n on subchannel k is then expressed as: in: It is the transmission power of the pilot sequence of user equipment m on subchannel k. It is the pilot sequence of user equipment m in subchannel k time slot t. It is a noise vector. Based on the received signal vector. The channel coefficient estimate is expressed by the minimum mean square error (LMMSE) as follows: in: Then the estimated channel coefficients Mean square error
[0021] B) Closed-loop uplink rate calculation during the uplink data transmission phase: During the uplink data transmission phase, each AP sends the received signal to the central processing unit (CPU). The CPU combines the received soft estimate to decode the data sent by the user equipment m. in: Data sent to user equipment m It is the uplink transmit power. It is a noise vector. Based on the data decoded value of user equipment m. The closed-loop uplink rate of user equipment m on subchannel k can be obtained by the UatF boundary method as follows:
[0022] Where: The first term in the denominator represents incoherent interference, i.e., the sum of the power of all interfering signals at the access point. The second term in the denominator represents the additional coherent interference caused by pilot pollution. The sum of the first and second terms is the sum of multi-user interference and pilot pollution. The third term represents noise power. Therefore, the total uplink rate of the m-th UE can be obtained by summing the rates of all sub-channels: Clearly, by optimizing user scheduling strategies, multi-user interference and pilot pollution can be effectively mitigated, thereby increasing uplink speed and reducing offload latency.
[0023] The parallel computing offloading model includes local computing latency and edge computing latency, specifically:
[0024] C) Local computation delay calculation: Local computation delay Equals the ratio of local computing power to the user's local computing capacity: Where: D m [t] represents the total amount of computational data for the m-th UE in time slot t, α m [t](0≤α m ≤1) is the unloading ratio. It is the local CPU's computing frequency. It is the computational complexity of the task for the m-th UE.
[0025] D) Edge computing latency: The amount of data offloaded from user device m to the MEC server for parallel computing is α. m [t]D m [t]. Edge latency can be divided into transmission latency for data offloading and computation latency on the edge server. Based on previously obtained R... m [t] represents the delay of offloading data transmission to the AP. for: Computational latency of user equipment m on the MEC server It can be represented as: in: (Total edge computing resources of the MEC server) is the size of edge computing resources allocated by the MEC server to user device m.
[0026] E) Total delay calculation: Due to the parallel computing offloading model adopted, the total delay T of user equipment m is calculated as follows. m [t] represents the larger of the local computing latency and the edge computing latency:
[0027] Step 2: Using Python, build a simulation environment for MEC-assisted decellularized uplink network based on the MEC-assisted decellularized communication model and parallel computing offloading model established in Step 1. Then, construct a deep reinforcement learning model for jointly optimizing computing offloading, user scheduling, and computing resource allocation strategies (JCSCS) and perform interactive training of the near-end policy optimization (PPO) method in the above simulation environment. Finally, obtain the optimal computing offloading strategy, user scheduling strategy, and computing resource allocation strategy through PPO.
[0028] like Figure 3 As shown, the parameters of the JCSCS model include: state space s[t], action space a[t], and reward r[t], where:
[0029] The state space is as follows: Where: D m [t] represents the total amount of computational data for user equipment m in time slot t. It is the sum of multi-user interference and pilot pollution on the sub-channel k of the previous time slot for user equipment m. It is the rate allocation of user equipment m on subchannel k in the previous time slot.
[0030] The action space is: Where: α m [t] is the offloading ratio of user equipment m in time slot t. Indicates whether to schedule user equipment m on subchannel k. It is the size of edge computing resources allocated by the MEC server to user device m.
[0031] The reward is: r[t] = -max{T} m [t]}, where: T m [t] represents the total latency of user equipment m.
[0032] like Figure 3 As shown, the PPO interactive training process specifically includes:
[0033] Step 1: Based on the current state s[t], the old actor network Generate action a[t] to determine user offloading strategy, user scheduling strategy on sub-channels, and edge computing resource allocation strategy; Commentator Network Used to evaluate the reward r[t];
[0034] Step 2, Action space {a[t], u[t], f e The application of [t]} to the environment results in the update of reward r[t] and the state is updated from s[t] to s[t+1]. At the same time, the state space-action space-reward {s[t], a[t], r[t]} under each time slot t is stored in the experience replay area.
[0035] Step 3: Train a new actor network π using samples from the experience replay area. θ and the network of critics The parameters of the deep learning network are updated using gradient descent, where: π is used to train the actor network. θ The loss function is in: ∈ is a hyperparameter, and clip is a cutoff function used to limit the magnitude of policy changes. This is the advantage function used to measure the quality of a movement. λ (0≤λ≤1) is the discount factor. PPO stops training when the reward r[t] converges.
[0036] Step 4: When the reward r[t] converges, PPO stops training and takes the action {α[t], u[t], f} at this time. e The value of [t]} serves as the optimal user offloading strategy, user scheduling strategy on sub-channels, and edge computing resource allocation strategy.
[0037] Through specific practical experiments, PPO was trained using Python 3.9 and TensorFlow 2.10. The network converged after 3500 epochs. Comparisons were then conducted with other methods from different perspectives. Specific environment settings included: epochs: 100 steps, M: 20, K: 10, N: 50, W: ... 5MHz D m : [200,300]KB, c m [500, 1000] cycles / bit GHz, [0.7,1.5]GHz, ρ k 50mW, p k 100mW, τ c 200 samples, τ p 10 samples.
[0038] like Figures 4-6 As shown in the figure, comparison method 1 is the user-free offloading method: all UEs are scheduled to be offloaded simultaneously on the entire channel; comparison method 2 is the random scheduling offloading method: all UEs are randomly scheduled to be offloaded on multiple sub-channels; comparison method 3 is edge-only offloading: all tasks are offloaded and calculated at the edge; comparison method 4 is local-only calculation: all tasks are offloaded and calculated locally.
[0039] like Figure 4 The figure shows the impact of the number of UEs on system performance under different methods. As the number of users increases, the transmission quality of the data link significantly deteriorates, mainly due to increased multi-user interference and pilot pollution, leading to increased transmission delay and system performance degradation. This invention performs best because it dynamically schedules users to allocate bandwidth, mitigating multi-user interference and pilot pollution on each sub-channel while ensuring each UE receives an appropriate bandwidth share. It is worth noting that when M ≤ τ p When the value is 10, the gain of this invention is relatively small because there is no pilot pollution at this time. The gain comes only from relatively small multi-user interference, while pilot pollution has a greater impact on system performance.
[0040] like Figure 5 The figure shows the relationship between maximum latency and total MEC computing power under different methods. It can be observed that the present invention performs best, and when... When it is small, the delay increases with The increase dropped sharply. However, when After reaching a certain value, the reduction in delay tends to slow down. This is because, when... When the latency is small, edge computing latency dominates, while when When the latency is high, offloading latency becomes the dominant factor. Comparing this invention with other methods that do not optimize user scheduling, the lack of suppression of multi-user interference and pilot pollution leads to higher offloading latency, resulting in underutilization of MEC server computing resources and ultimately premature curve convergence.
[0041] like Figure 6 The figure shows the relationship between maximum latency and bandwidth W under different methods. It can be seen that the present invention performs best; as bandwidth W increases, the maximum latency of all offloading methods gradually decreases, but for higher bandwidth values, this reduction tends to saturate. Furthermore, the performance gap between the present invention and other methods gradually narrows. At this point, latency is mainly limited by the computing power of the MEC server, and the gain from user scheduling becomes limited.
[0042] Compared with the prior art, the present invention fully considers the impact of multi-user interference and pilot pollution on offloading delay, and minimizes the delay of all user equipment by dynamically scheduling user equipment for offloading on multiple orthogonal sub-channels.
[0043] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.
Claims
1. A user-schedule-based computing offloading method, characterized in that, After modeling the MEC-assisted decellular communication model and the parallel computing offloading model to construct the uplink network simulation environment, a joint optimization computing offloading, user scheduling, and computing resource allocation strategy (JCSCS) model considering limited orthogonal pilot resources and edge computing resources is constructed and trained in the uplink network simulation environment using deep reinforcement learning based on near-end policy optimization (PPO). Finally, the optimal computing offloading strategy, user scheduling strategy, and computing resource allocation strategy are obtained through PPO.
2. The user-schedule-based computation offloading method of claim 1, wherein, The JCSCS model includes an environment module, a policy generation module, a policy evaluation module, and an experience replay module. The environment encoding module generates feature vectors based on state information reflecting environmental characteristics. The policy generation module outputs a policy consisting of a set of multiple future actions based on the feature vectors. The policy evaluation module evaluates and estimates the current policy and feeds it back to the policy generation module to adjust its policy model. The experience replay module is used to improve the model training efficiency and performance.
3. The user-schedule-based computation offloading method of claim 1, wherein, The aforementioned decellular communication model includes the channel coefficients estimated during the uplink training phase and the closed-loop uplink rate during the uplink data transmission phase, specifically: A) Channel coefficient calculation during the uplink training phase: The system bandwidth W is uniformly divided into K orthogonal sub-channels, and a fully synchronized time-slot system is considered. Each time slot t (t≥0) represents a user scheduling period. The wireless time-varying channel fading model used in the training phase is as follows: Where: β m,n This represents the large-scale fading factor, which includes path loss and shadow fading. This represents the small-scale fading of user equipment m (0≤m≤M) in time slot t of access point n (0≤n≤N) on subchannel k (0≤k≤K), which follows a Z-state Markov model; the uplink signal matrix received by access point n on subchannel k is represented as... in: It is the transmission power of the pilot sequence of user equipment m on subchannel k. It is the pilot sequence of user equipment m in subchannel k time slot t. It is a noise vector; Based on the received signal vector The channel coefficient estimate is expressed by the minimum mean square error (LMMSE) as follows: in: Then the estimated channel coefficients Mean square error B) Closed-loop uplink rate calculation during the uplink data transmission phase: During the uplink data transmission phase, each AP sends the received signal to the central processing unit and decodes the data sent by the user equipment m based on the received soft estimate. in: Data sent to user equipment m It is the uplink transmit power. It is a noise vector, based on the data decoded value of user equipment m. The closed-loop uplink rate of user equipment m on subchannel k can be obtained by the UatF boundary method as follows: Where: the first term in the denominator represents incoherent interference, i.e., the sum of the power of all interfering signals at the access point; the second term in the denominator represents additional coherent interference caused by pilot pollution; the sum of the first and second terms is the sum of multi-user interference and pilot pollution. The third term represents noise power; therefore, the total uplink rate of the m-th user equipment can be obtained by summing the rates of all sub-channels: Clearly, by optimizing the user scheduling strategy, multi-user interference and pilot pollution can be effectively mitigated, thereby improving uplink speed and reducing offload latency.
4. The computational offloading method based on user scheduling according to claim 1, characterized in that, The parallel computing offloading model includes local computing latency and edge computing latency, specifically: C) Local computation delay calculation: Local computation delay Equals the ratio of local computing power to the user's local computing capacity: Where: D m [t] represents the total amount of computational data for the m-th UE in time slot t, α m [t](0≤α m ≤1) is the unloading ratio. It is the local CPU's computing frequency. This is the computational complexity of the task for the m-th UE; D) Edge computing latency: The amount of data offloaded from user device m to the MEC server for parallel computing is α. m [t]D m [t], edge latency is divided into data offloading transmission latency and computation latency on the edge server, based on previously obtained R. m [t] represents the delay in offloading data transmission to the AP. Computational latency of user equipment m on the MEC server in: This refers to the size of edge computing resources allocated by the MEC server to user device m. It represents the total edge computing resources of MEC servers; E) Total delay calculation: due to the parallel computation offloading model adopted, the total delay Tmof the user equipment m is calculated as follows: m [t] is the maximum of the local computation delay and the edge computation delay:
5. The user-schedule-based computation offloading method according to claim 1 or 2, wherein, The parameters of the JCSCS model include: state space s[t], action space a[t], and reward r[t], where: The state space is as follows: Where: D m [t] represents the total amount of computational data for user equipment m in time slot t. It is the sum of multi-user interference and pilot pollution interference of user equipment m on sub-channel k in the previous time slot. It is the rate allocation of user equipment m on subchannel k in the previous time slot; The action space is: Where: a m [t] is the offloading ratio of user equipment m in time slot t. Indicates whether to schedule user equipment m on subchannel k. It is the size of edge computing resources allocated by the MEC server to user device m; The reward is: r[t] = -max{T m [t]}, where: T m [t] is the total latency of the user equipment m.
6. The user-schedule-based computation offloading method of claim 1, wherein, The interactive training in the PPO specifically includes: Step 1, based on the current state s[t], the old actor network generates an action a[t] to decide the user offloading policy, the user scheduling policy on the sub-channel, and the edge computing resource allocation policy; the critic network for evaluating the reward r[t]; Step 2, Action space {a[t], u[t], f e [t]} applied to the environment results in an update of the reward r[t] and an update of the state from s[t] to s[t+1], while storing the state space-action space-reward {s[t], a[t], r[t]} at each time slot t to the experience replay area; Step 3: Train a new actor network π using samples from the experience replay area. θ and the network of critics The parameters of the deep learning network are updated using gradient descent, where: π is used to train the actor network. θ The loss function is in: ∈ is a hyperparameter, and clip is a cutoff function used to limit the magnitude of policy changes. This is the advantage function used to measure the quality of a movement. λ (0≤λ≤1) is the discount factor, and PPO stops training when the reward r[t] converges; Step 4: When the reward r[t] converges, PPO stops training and takes the action {α[t], u[t], f} at this time. e The value of [t]} serves as the optimal user offloading strategy, user scheduling strategy on sub-channels, and edge computing resource allocation strategy.