Mobile edge computing network terminal cooperative scheduling method and device, equipment and medium
By employing a mobile edge computing network terminal collaborative scheduling method in a multi-wireless terminal network, the timeliness of information reception is optimized, the problem of information age optimization is solved, and effective bandwidth and energy management is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2023-11-29
- Publication Date
- 2026-07-28
AI Technical Summary
In multi-wireless terminal networks, the age of information is difficult to optimize effectively, affecting the timeliness of information reception.
A mobile edge computing network terminal collaborative scheduling method is adopted. By dividing time slots, the actions of wireless terminals are determined based on preset loss values and relative post-decision state value functions. The bandwidth and energy constraints are optimized through iterative updates of Lagrange multipliers, which is expressed as a constrained Markov decision process problem.
It improves the timeliness of information reception, optimizes information age, and meets the time average constraints of bandwidth and energy.
Smart Images

Figure CN117693053B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information timeliness optimization, and in particular to a method, apparatus, device and medium for collaborative scheduling of mobile edge computing network terminals. Background Technology
[0002] In 5G (5th Generation Mobile Communication Technology) wireless terminal networks, the timeliness of information is crucial. Traditional indicators for measuring information freshness, such as latency and throughput, are no longer sufficient to fully reflect the timeliness of messages. Age of Information (AoI) is an indicator for measuring the timeliness of data transmission and has important applications in the field of wireless network communication. It represents the time interval between the generation time of the latest data received by the receiver and the current time, describing the freshness of the data currently received by the receiver. In recent years, AoI has become an important indicator for measuring the timeliness of information.
[0003] In related technologies, time slot scheduling is used to plan the information transmission order between communicating parties in a wireless network, and to optimize the average information age of the network through time slot scheduling.
[0004] However, this method still fails to optimize information age in multi-wireless terminal networks. In each time slot of the wireless network, link scheduling is chaotic and irrational, which seriously affects the timeliness of information reception and urgently needs to be addressed. Summary of the Invention
[0005] This application provides a method, apparatus, device, and medium for collaborative scheduling of mobile edge computing network terminals to address the problem that information age is still not well optimized in current multi-wireless terminal networks and improve the timeliness of information reception.
[0006] To achieve the above objectives, the first aspect of this application proposes a mobile edge computing network terminal collaborative scheduling method, comprising the following steps:
[0007] Identify multiple wireless terminals and divide them into multiple time slots;
[0008] Based on preset first-step loss values, preset second-step loss values, and relative post-decision state value functions, the first actions of multiple wireless terminals used for edge transmission and the second actions of the remaining wireless terminals are determined; and
[0009] In the current time slot, the multiple wireless terminals used for edge transmission are controlled to perform a first action, the remaining wireless terminals are controlled to perform a second action to obtain a post-decision state, and the load of each wireless terminal is observed. Based on the post-decision state and the observation results, the relative state value function of each wireless terminal is calculated, and the relative state value function and Lagrange multiplier value of each wireless terminal are updated based on the relative state value function before entering the next time slot.
[0010] According to one embodiment of this application, determining the first action of a plurality of wireless terminals used for edge transmission and the second action of the remaining wireless terminals based on a preset first-step loss value, a preset second-step loss value, and a relative post-decision state value function includes:
[0011] The number of wireless terminals used for edge transmission and the current number of sub-channels are obtained, and the first action of each wireless terminal is determined based on a preset first action strategy.
[0012] When the number of wireless terminals used for edge transmission is greater than the current number of sub-channels, the plurality of wireless terminals used for edge transmission are determined based on the current number of sub-channels, and the other wireless terminals among the plurality of wireless terminals are taken as the remaining wireless terminals.
[0013] Based on a preset second action strategy, the second action of the remaining wireless terminals is determined.
[0014] According to one embodiment of this application, the preset first action strategy is:
[0015]
[0016] Where, α i (t) represents the preset first action strategy. λ is the preset first-step loss value for edge server computation. B For the current bandwidth Lagrange multiplier, For wireless terminal i in the current energy Lagrange multiplier, S i (t) represents the state of wireless terminal i at time t, where t is the current time, and α i For action, Let f be the state-value function after relative decision-making. PDS This represents the state after the decision, where i is the number of the wireless terminal.
[0017] According to one embodiment of this application, the preset second action strategy is:
[0018]
[0019] Among them, a i ′(t) represents the preset second action strategy. This is the second-step loss value preset for local computation.
[0020] According to one embodiment of this application, the relative state value function is:
[0021]
[0022] Among them, V i (S i S is the relative state value function. i For the state of wireless terminal i, v i Let be the value of the Lagrange multiplier.
[0023] According to the mobile edge computing network terminal collaborative scheduling method proposed in this application, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals are determined based on a preset one-step loss value and a relative post-decision state value function. In the current time slot, the multiple wireless terminals used for edge transmission and the remaining wireless terminals are controlled to execute the first and second actions respectively, obtaining the post-decision state. Based on the post-decision state and the observation results of the load of each wireless terminal, the relative state value function of each wireless terminal is calculated. After updating the relative state value function and the Lagrange multiplier value, the system proceeds to the next time slot. Thus, the information timeliness optimization problem is formulated as a constrained Markov decision process problem with bandwidth and energy constraints. The iterative update of the Lagrange multiplier ensures the constraints on bandwidth and energy under time averaging, solving the problem that information age in multi-wireless terminal networks is still not well optimized at present, and improving the timeliness of information reception.
[0024] To achieve the above objectives, a second aspect of this application provides a mobile edge computing network terminal collaborative scheduling device, comprising:
[0025] The first determining module is used to determine multiple wireless terminals and divide multiple time slots;
[0026] The second determining module is used to determine, based on a preset first-step loss value, a preset second-step loss value, and a relative post-decision state value function, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals; and
[0027] The processing module is configured to determine the current time slot from the plurality of time slots, control the plurality of wireless terminals used for edge transmission to perform a first action, control the remaining wireless terminals to perform a second action to obtain a post-decision state, observe the load of each wireless terminal, calculate the relative state value function of each wireless terminal based on the post-decision state and the observation results, and update the relative state value function and Lagrange multiplier value of each wireless terminal based on the relative state value function before entering the next time slot.
[0028] According to one embodiment of this application, the second determining module is specifically used for:
[0029] The number of wireless terminals used for edge transmission and the current number of sub-channels are obtained, and the first action of each wireless terminal is determined based on a preset first action strategy.
[0030] When the number of wireless terminals used for edge transmission is greater than the current number of sub-channels, the plurality of wireless terminals used for edge transmission are determined based on the current number of sub-channels, and the other wireless terminals among the plurality of wireless terminals are taken as the remaining wireless terminals.
[0031] Based on a preset second action strategy, the second action of the remaining wireless terminals is determined.
[0032] According to one embodiment of this application, the preset first action strategy is:
[0033]
[0034] Where, α i (t) represents the preset first action strategy. λ is the preset first-step loss value for edge server computation. B For the current bandwidth Lagrange multiplier, For wireless terminal i in the current energy Lagrange multiplier, S i (t) represents the state of wireless terminal i at time t, where t is the current time, and α i For action, Let f be the state-value function after relative decision-making. PDS This represents the state after the decision, where i is the number of the wireless terminal.
[0035] According to one embodiment of this application, the preset second action strategy is:
[0036]
[0037] Among them, a i ′ (t) represents the preset second action strategy. This is the second-step loss value preset for local computation.
[0038] According to one embodiment of this application, the relative state value function is:
[0039]
[0040] Among them, V i (S i S is the relative state value function. i For the state of wireless terminal i, v i Let be the value of the Lagrange multiplier.
[0041] According to the mobile edge computing network terminal collaborative scheduling device proposed in this application, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals are determined based on a preset one-step loss value and a relative post-decision state value function. In the current time slot, the multiple wireless terminals used for edge transmission and the remaining wireless terminals are controlled to execute the first and second actions respectively to obtain the post-decision state. Based on the post-decision state and the observation results of the load of each wireless terminal, the relative state value function of each wireless terminal is calculated. After updating the relative state value function and the Lagrange multiplier value, the device proceeds to the next time slot. Thus, the information timeliness optimization problem is formulated as a constrained Markov decision process problem with bandwidth and energy constraints. The iterative update of the Lagrange multiplier ensures the constraints on bandwidth and energy under time averaging, solving the problem that information age in current multi-wireless terminal networks is still not well optimized, and improving the timeliness of information reception.
[0042] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the mobile edge computing network terminal collaborative scheduling method as described in the above embodiments.
[0043] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the mobile edge computing network terminal collaborative scheduling method as described in the above embodiments.
[0044] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0045] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0046] Figure 1 This is a flowchart of a mobile edge computing network terminal collaborative scheduling method provided according to an embodiment of this application;
[0047] Figure 2 This is a schematic diagram illustrating the evolution of AoI in a Generate-at-will scenario for a wireless terminal according to an embodiment of this application;
[0048] Figure 3 This is a flowchart of a mobile edge computing network terminal collaborative scheduling method according to another embodiment of this application;
[0049] Figure 4 This is a schematic diagram illustrating the average bandwidth usage according to one embodiment of this application;
[0050] Figure 5 This is a schematic diagram of the average energy consumption according to an embodiment of this application;
[0051] Figure 6 This is a schematic diagram of the average AoI performance according to an embodiment of this application;
[0052] Figure 7 This is a schematic diagram illustrating the impact of M / N on the optimization performance of the AoI algorithm according to an embodiment of this application;
[0053] Figure 8 This is a block diagram of a mobile edge computing network terminal collaborative scheduling device provided according to an embodiment of this application;
[0054] Figure 9 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0055] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0056] The following describes, with reference to the accompanying drawings, a method, apparatus, device, and medium for collaborative scheduling of mobile edge computing network terminals according to embodiments of this application. First, the method for collaborative scheduling of mobile edge computing network terminals according to embodiments of this application will be described with reference to the accompanying drawings.
[0057] Figure 1 This is a flowchart of a mobile edge computing network terminal collaborative scheduling method according to an embodiment of this application.
[0058] Before introducing the mobile edge computing network terminal collaborative scheduling method proposed in the embodiments of this application, let's first introduce the AoI scheduling optimization method in related technologies.
[0059] In related technologies, the researched AoI scheduling optimization methods aim to minimize the average AoI in a randomly generated scenario (i.e., the Generate-at-will scenario) on the one hand, and to minimize the average AoI under sample variation conditions, where state update packets follow a stochastic process (i.e., the Sample-at-change scenario) on the other. Researchers have considered various constraints in real-world networks, such as channel limitations and power limitations, and designed optimization scheduling methods such as the Lyapunov optimization method and the Whittle exponent method to reduce AoI.
[0060] Previous AoI optimization scenarios did not consider the processing time of raw data. In reality, wireless terminals need to perform calculations on the raw data before transmission. However, wireless terminals are typically small in size and have low production costs, which usually limits their computing power in performing fast calculations. Based on this, Multi-access Edge Computing (MEC) systems have been proposed and applied to improve computing performance. MEC systems place edge servers with computing, transmission, and storage capabilities close to the data source. Wireless terminals offload the raw data workload to nearby edge servers for computation via wireless channels, thereby reducing latency, achieving faster network service response, and reducing the amount of data transmitted over long distances, thus reducing energy consumption. Therefore, the application of MEC systems not only reduces the energy consumption of wireless terminals but also improves the freshness of computation results.
[0061] Since the processing time of raw data by wireless terminals needs to be taken into account, the existing definition of AoI needs to be modified. Some researchers have considered the raw data processing time and proposed the Age of Task (AoT) and Age of Process (AoP) as metrics for information freshness, and optimized the information freshness of offline and online scenarios respectively in the Sample-at-change scenario.
[0062] However, the relevant technologies only address the case of a single wireless terminal. It is necessary to explore collaborative scenarios between multiple wireless terminals and between wireless terminals and the network, and to design a method to optimize information freshness in multi-wireless terminal networks under the Generate-at-will scenario.
[0063] In practical MEC networks, the channel state changes rapidly, making it difficult to obtain channel state statistics. The computational workload of the raw data is also variable. Therefore, the system should have the ability to estimate and learn parameters related to the surrounding environment so that the scheduling algorithm can adapt to the constantly changing environment. In this application, the Q-learning algorithm is used to learn the unknown environmental parameters in the Generate-at-will scenario to optimize the AoI in the Generate-at-will scenario.
[0064] For example, such as Figure 1 As shown, the mobile edge computing network terminal collaborative scheduling method includes the following steps:
[0065] In step S101, multiple wireless terminals are identified and multiple time slots are divided.
[0066] Among them, wireless terminal refers to the wireless terminal module used to realize wireless data transmission. It is usually connected to the lower-level machine to achieve the purpose of wireless data transmission. Its transmission principle is basically the same as that of mobile phone data transmission. Typical wireless terminal devices include wireless data transmission, wireless router and wireless modem, etc.; time slot refers to dividing a period of time into several small segments, each of which is a time slot.
[0067] Specifically, the embodiments of this application can model the data processing volume and energy consumption of an MEC system. First, multiple wireless devices (WDs) of an MEC system are identified, denoted as... Wireless terminals can transmit data to edge servers, and the time can be divided into multiple time slots, denoted as... The time interval for each time slot is ΔT.
[0068] Multiple wireless terminals can read raw data in each time slot. The read raw data can form a task. For each wireless terminal i, the size of the task received in each time slot is represented as... Position, among which Follows a specific distribution; for each time slot t, the task In a generate-at-will scenario, the wireless terminal can choose to compute the task for transmission or choose to leave the task idle to conserve energy. Furthermore, each task can be divided into multiple subtasks, which can be computed locally by the wireless terminal's built-in processor or transmitted wirelessly to an edge server for online computation.
[0069] It should be noted that when subtasks are computed locally, the power consumption mainly comes from the operation of the wireless terminal processor. The wireless terminal processor typically uses dynamic voltage and frequency scaling techniques to reduce power consumption. Assume that the processor of wireless terminal i operates at a frequency of f in time slot t. i If processing 1 bit of data requires ω processor cycles, then the amount of data that can be computed in time slot t can be expressed by the following formula:
[0070]
[0071] in, This represents the amount of data that can be computed in time slot t during local computation.
[0072] The energy consumption of a CPU (Central Processing Unit) per cycle is proportional to the square of its frequency, which can be expressed as follows: Therefore, the local energy consumption for this time slot can be expressed by the following formula:
[0073]
[0074] in, γ represents the energy consumption of time slot t during local computation, and γ is a chip structure constant.
[0075] When the subtask is transmitted to the edge server for computation, energy consumption is attributed to data transmission. The channel gain between the wireless terminal i and the edge server in time slot t is denoted as h. i (t), the allocated bandwidth is denoted as W. i (t), the transmission power is denoted as P i (t), the channel noise power is denoted as σ. 2 According to Shannon's formula, the amount of data transmitted in time slot t is expressed by the following formula:
[0076]
[0077] Energy consumption is expressed by the following formula:
[0078]
[0079] Therefore, the total data computation amount for each wireless terminal in each time slot is:
[0080]
[0081] The total power consumption of each wireless terminal in each time slot is:
[0082]
[0083] in, The amount of data transmitted in time slot t during computation on the edge server. Let d be the energy consumption of time slot t when computing is performed on the edge server. i (t) represents the total data computation amount per wireless terminal per time slot, E i (t) represents the total power consumption of each wireless terminal in each time slot.
[0084] Next, this application embodiment incorporates the processing time of the raw data into the definition of AoI under the MEC network. For example... Figure 2 As shown, Figure 2 This illustrates the evolution of the AoI of a wireless terminal in a Generate-at-will scenario. The solid line represents the AoI evolution curve of the wireless terminal, and the dashed line represents the evolution of the reception time of the task being computed in the wireless terminal. If no task is completed in a certain time slot, the AoI and reception time of the task being computed in the next time slot are both increased by 1. The arrow indicates that the task in the wireless terminal has been completed in this time slot, causing the AoI of the wireless terminal to jump in the next time moment to match the reception time of the most recently computed task, and the wireless terminal starts a new task reception and computation.
[0085] At the initial moment (i.e., t=0), the AoI of each wireless terminal is 1, and a new task reception and calculation process begins. Let A... i (t) represents the AoI of wireless terminal i at time slot t, let t i (t) represents the reception time of the most recently completed task by wireless terminal i, then A i The expression for (t) is as follows:
[0086] A i (t)=tt i (t);
[0087] If no new receiving tasks are completed, the AoI will continue to increase. Therefore, a method needs to be designed to optimize the average AoI under the following constraints:
[0088]
[0089] in, This represents the number of remaining bits for the task processed by wireless terminal i in time slot t.
[0090] In step S102, based on the preset first-step loss value, the preset second-step loss value, and the relative decision-after state value function, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals are determined.
[0091] Furthermore, in some embodiments, based on a preset first-step loss value, a preset second-step loss value, and a relative post-decision state value function, determining the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals includes: obtaining the number of wireless terminals used for edge transmission and the current number of sub-channels, and determining the first action of each wireless terminal based on a preset first action strategy; when the number of wireless terminals used for edge transmission is greater than the current number of sub-channels, determining multiple wireless terminals used for edge transmission based on the current number of sub-channels, and designating the other wireless terminals among the multiple wireless terminals as the remaining wireless terminals; and determining the second action of the remaining wireless terminals based on a preset second action strategy.
[0092] Specifically, based on multiple wireless terminals, the number of wireless terminals O(t) used for edge transmission and the current number of sub-channels M are obtained. In each time slot of t = 1, 2, ..., T, the first action of each wireless terminal is determined based on a preset first action strategy.
[0093] In some embodiments, the preset first action strategy is:
[0094]
[0095] Where, α i (t) represents the preset first action strategy. The first-step cost, λ, is preset for calculations on edge servers. B For the current bandwidth Lagrange multiplier, For wireless terminal i in the current energy Lagrange multiplier, S i (t) represents the state of wireless terminal i at time t, where t is the current time, and α i For action, Let f be the state-value function after relative decision-making. PDS This represents the state after the decision, where i is the number of the wireless terminal.
[0096] Furthermore, Specifically, it is expressed as follows:
[0097]
[0098] S i (t) includes four dimensions, that is
[0099] Among them, A i (t) represents the AoI of wireless terminal i at time t, and s i (t) represents the difference between the current time t and the sampling time slot of the task being processed by the wireless terminal i.
[0100] When the number of wireless terminals O(t) used for edge transmission is greater than the current number of sub-channels M, based on the current number of sub-channels, M wireless terminals used for edge transmission are randomly selected from multiple wireless terminals used for edge transmission for transmission, and the other wireless terminals among the multiple wireless terminals are regarded as the remaining wireless terminals. The remaining wireless terminals can only reselect actions from the action space containing local computation, that is, based on the preset second action strategy, the second action of the remaining wireless terminals is determined.
[0101] In some embodiments, the preset second action strategy is:
[0102]
[0103] Where, a′ i (t) represents the preset second action strategy. This is the preset second-step loss value (one-step cost) for local computation.
[0104] in, Specifically, it is expressed as follows:
[0105]
[0106] In step S103, multiple wireless terminals used for edge transmission are controlled to perform a first action in the current time slot, and the remaining wireless terminals are controlled to perform a second action to obtain a post-decision state. The load of each wireless terminal is observed, and the relative state value function of each wireless terminal is calculated based on the post-decision state and the observation results. The relative state value function and Lagrange multiplier value of each wireless terminal are updated based on the relative state value function before entering the next time slot.
[0107] Specifically, for each wireless terminal, multiple wireless terminals used for edge transmission are controlled to perform the first action, while the remaining wireless terminals perform the second action and transition to the post-decision state. The load for each wireless terminal i Observe and convert to S i (t+1), calculate S respectively i (t) and S i (t+1) Relative state value function.
[0108] In some embodiments, the relative state value function is:
[0109]
[0110] Among them, V i (S i S is a relative state value function. i For the state of wireless terminal i, v iIt is the value of the Lagrange multiplier.
[0111] Further updates
[0112]
[0113] Update v i :
[0114]
[0115] in, Let β be the state-value function after relative decision-making, and let β be the Q-learning update weights.
[0116] Furthermore, update each wireless terminal's...
[0117]
[0118] Update λ B :
[0119] λ B (t+1)←ReLU[λ l (t)+θ B (t)(W(t)-W max )];
[0120] Where, θ E (t) represents the energy Lagrange multiplier update rate, E i (t) represents the power consumption of each wireless terminal in each time slot, θ B W(t) is the bandwidth Lagrange multiplier update rate, and W(t) is the allocated bandwidth. max This is a maximum bandwidth constraint.
[0121] After updating the relative state value function and Lagrange multiplier value of each wireless terminal based on the relative state value function, the next time slot is entered. The next time slot is used as the new current time slot. The process of controlling multiple wireless terminals used for edge transmission to perform the first action, controlling the remaining wireless terminals to perform the second action to obtain the post-decision state, and observing the load of each wireless terminal, and calculating the relative state value function of each wireless terminal based on the post-decision state and the observation results, and updating the relative state value function and Lagrange multiplier value of each wireless terminal based on the relative state value function before entering the next time slot.
[0122] It should be noted that the initial settings for each wireless terminal are as follows:
[0123] To facilitate a better understanding of the mobile edge computing network terminal collaborative scheduling method proposed in the embodiments of this application, further explanation is provided below with reference to specific embodiments.
[0124] like Figure 3 As shown, Figure 3 This is a flowchart of a mobile edge computing network terminal collaborative scheduling method according to another embodiment of this application. The method may include the following steps:
[0125] Step S301: Generate the original data for the new time slot and the new communication status.
[0126] Step S302: Select the wireless terminal and perform the action.
[0127] Step S303: Update the relative state value function.
[0128] Step S304: Update the Lagrange multiplier values.
[0129] Jump to the next time slot and continue executing steps S302 to S304. Steps S302 to S304 are the algorithm within time slot t.
[0130] For example, suppose in a certain scenario, the number of time slots T = 3000, the time interval of each time slot ΔT = 10ms, the number of wireless terminals N = 7, the number of sub-channels M = 3, and the channel noise power σ 2 =1×10 -9 W, number of processing cycles ω = 1000, chip structure constant γ = 5 × 10 -28 The maximum local processing frequency constraint is f. i max =1GHz, maximum transmission power constraint is The maximum bandwidth constraint is W max =20MHz, Q-learning updates weights Bandwidth Lagrange multiplier update rate Energy Lagrange multiplier update rate Load of raw data at each time step Following a uniform distribution in the range [10Kb, 50Kb], each task can be divided into subtasks of size 5Kb. For each wireless terminal, h i (t) is modeled as having a channel gain of 3 × 10 -10 6×10 -10 and 9×10 -10 A 3-state Markov chain, where the state transition probability matrix is:
[0131]
[0132] The energy constraints for multiple wireless terminals are set sequentially as follows:
[0133] In summary, (1) the convergence of the embodiments of this application is verified by simulation, and the average bandwidth usage is as follows: Figure 4 As shown, the upper curve represents the theoretically calculated average number of occupied sub-channels when the bandwidth constraint is relaxed to a time-averaged constraint, which converges to M, or 3. The lower curve represents the average number of sub-channels actually occupied by the algorithm to satisfy the bandwidth constraint of each time slot, which converges to a value slightly less than 3, approximately 2.55. Simulation results show that the method proposed in this application embodiment can satisfy the bandwidth constraint; the average power consumption of each wireless terminal converges to the constraint value, such as... Figure 5 As shown, let E i E represents the average power consumption of each wireless terminal. i It eventually converges to the energy constraint. Simulation results show that the method proposed in this application can satisfy energy constraints.
[0134] (2) The AoI performance implemented in this application embodiment is verified by simulation, and compared with the performance of the greedy algorithm in related technologies. The greedy algorithm works as follows: In each time slot, M sub-channels are allocated to the M wireless terminals with the highest AoI. Each wireless terminal selects an action based on the maximum energy it can consume in the current time slot. If the operation selected by the M wireless terminals with the highest AoI does not fully occupy the M sub-channels, the sub-channel allocation will be expanded sequentially according to the AoI size.
[0135] Average AoI performance as Figure 6 As shown, the lower curve represents the average AoI under the algorithm proposed in this application embodiment, and the upper curve represents the average AoI under the greedy algorithm. The Q-learning method in this application embodiment can achieve AoI convergence, and the AoI performance is significantly better than that of the traditional greedy algorithm. This application embodiment can better perceive the surrounding environmental parameters and better allocate bandwidth and energy.
[0136] (3) The impact of M / N on AoI optimization performance was verified through simulation. Assuming N=7, for all wireless terminals, As M varies from 1 to 6, the average AoI convergence value is as follows: Figure 7As shown, the upper curve represents the impact of M / N on AoI optimization performance under the greedy algorithm, while the lower curve represents the impact of M / N on AoI optimization performance under the algorithm proposed in this application embodiment. In terms of AoI optimization, this application embodiment consistently outperforms the greedy algorithm, achieving the minimum AoI when M=3. When M / N is too small, the number of allocated sub-channels is limited, restricting the flexibility of bandwidth resource scheduling and resulting in insufficient utilization of available bandwidth. When M / N is too large, the bandwidth of each sub-channel decreases, leading to increased energy consumption during transmission. In this case, this application embodiment prioritizes local computation, further reducing the effective utilization of bandwidth resources.
[0137] Therefore, (1) the embodiments of this application take into account the computation time of the original data, modify the definition of AoI according to the Generate-at-will scenario of the MEC system, and on this basis, express the AoI optimization problem as a Constraint Markov Decision Process (CMDP) problem with bandwidth and energy constraints, which is composed of multiple wireless terminals coupled together; (2) the embodiments of this application use Lagrange multipliers to relax the constraints, decouple the multi-terminal problem into several sub-problems, each sub-problem containing only one wireless terminal, these sub-problems can be solved by Q-learning, the iterative update of the Lagrange multipliers ensures the constraints on bandwidth and energy under time average, and finally adopts a truncation strategy to ensure that the strict bandwidth constraints are met and improve AoI performance.
[0138] According to the mobile edge computing network terminal collaborative scheduling method proposed in this application, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals are determined based on a preset one-step loss value and a relative post-decision state value function. In the current time slot, the multiple wireless terminals used for edge transmission and the remaining wireless terminals are controlled to execute the first and second actions respectively, obtaining the post-decision state. Based on the post-decision state and the observation results of the load of each wireless terminal, the relative state value function of each wireless terminal is calculated. After updating the relative state value function and the Lagrange multiplier value, the system proceeds to the next time slot. Thus, the information timeliness optimization problem is formulated as a constrained Markov decision process problem with bandwidth and energy constraints. The iterative update of the Lagrange multiplier ensures the constraints on bandwidth and energy under time averaging, solving the problem that information age in current multi-wireless terminal networks is still not well optimized, and improving the timeliness of information reception.
[0139] Next, with reference to the accompanying drawings, a mobile edge computing network terminal collaborative scheduling device proposed according to an embodiment of this application is described.
[0140] Figure 8This is a block diagram of a mobile edge computing network terminal collaborative scheduling device according to an embodiment of this application.
[0141] like Figure 8 As shown, the mobile edge computing network terminal collaborative scheduling device 10 includes: a first determining module 100, a second determining module 200, and a processing module 300.
[0142] The first determining module 100 is used to determine multiple wireless terminals and divide multiple time slots;
[0143] The second determining module 200 is used to determine the first actions of multiple wireless terminals used for edge transmission and the second actions of the remaining wireless terminals by using a preset first-step loss value, a preset second-step loss value, and a relative post-decision state value function; and
[0144] The processing module 300 is used to control multiple wireless terminals used for edge transmission to perform a first action in the current time slot, control the remaining wireless terminals to perform a second action to obtain the post-decision state, observe the load of each wireless terminal, calculate the relative state value function of each wireless terminal based on the post-decision state and the observation results, and update the relative state value function and Lagrange multiplier value of each wireless terminal based on the relative state value function before entering the next time slot.
[0145] Furthermore, in some embodiments, the second determining module 200 is specifically used for:
[0146] The number of wireless terminals used for edge transmission and the current number of sub-channels are obtained, and the first action of each wireless terminal is determined based on a preset first action strategy.
[0147] When the number of wireless terminals used for edge transmission is greater than the current number of sub-channels, multiple wireless terminals used for edge transmission are determined based on the current number of sub-channels, and the other wireless terminals among the multiple wireless terminals are regarded as the remaining wireless terminals.
[0148] Based on the preset second action strategy, determine the second action of the remaining wireless terminals.
[0149] Furthermore, in some embodiments, the preset first action strategy is:
[0150]
[0151] Where, α i (t) represents the preset first action strategy. λ is the preset first-step loss value for edge server computation. B For the current bandwidth Lagrange multiplier, For wireless terminal i in the current energy Lagrange multiplier, S i(t) represents the state of wireless terminal i at time t, where t is the current time, and α i For action, Let f be the state-value function after relative decision-making. PDS This represents the state after the decision, where i is the number of the wireless terminal.
[0152] Furthermore, in some embodiments, the preset second action strategy is:
[0153]
[0154] Among them, a i i (t) represents the preset second action strategy. This is the second-step loss value preset for local computation.
[0155] Furthermore, in some embodiments, the relative state value function is:
[0156]
[0157] Among them, V i (S i S is a relative state value function. i For the state of wireless terminal i, v i It is the value of the Lagrange multiplier.
[0158] It should be noted that the foregoing explanation of the mobile edge computing network terminal collaborative scheduling method embodiment also applies to the mobile edge computing network terminal collaborative scheduling device of this embodiment, and will not be repeated here.
[0159] According to the mobile edge computing network terminal collaborative scheduling device proposed in this application, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals are determined based on a preset one-step loss value and a relative post-decision state value function. In the current time slot, the multiple wireless terminals used for edge transmission and the remaining wireless terminals are controlled to execute the first and second actions respectively to obtain the post-decision state. Based on the post-decision state and the observation results of the load of each wireless terminal, the relative state value function of each wireless terminal is calculated. After updating the relative state value function and the Lagrange multiplier value, the device proceeds to the next time slot. Thus, the information timeliness optimization problem is formulated as a constrained Markov decision process problem with bandwidth and energy constraints. The iterative update of the Lagrange multiplier ensures the constraints on bandwidth and energy under time averaging, solving the problem that information age in current multi-wireless terminal networks is still not well optimized, and improving the timeliness of information reception.
[0160] Figure 9 A schematic diagram of the structure of a vehicle provided in an embodiment of this application. The vehicle may include:
[0161] The memory 901, the processor 902, and the computer program stored on the memory 901 and capable of running on the processor 902.
[0162] When the processor 902 executes the program, it implements the mobile edge computing network terminal collaborative scheduling method provided in the above embodiments.
[0163] Furthermore, the vehicle also includes:
[0164] Communication interface 903 is used for communication between memory 901 and processor 902.
[0165] The memory 901 is used to store computer programs that can run on the processor 902.
[0166] The memory 901 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0167] If the memory 901, processor 902, and communication interface 903 are implemented independently, then the communication interface 903, memory 901, and processor 902 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0168] Optionally, in a specific implementation, if the memory 901, processor 902, and communication interface 903 are integrated on a single chip, then the memory 901, processor 902, and communication interface 903 can communicate with each other through an internal interface.
[0169] The processor 902 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0170] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described mobile edge computing network terminal collaborative scheduling method.
[0171] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0172] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0173] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for collaborative scheduling of mobile edge computing network terminals, characterized in that, Includes the following steps: Identify multiple wireless terminals and divide them into multiple time slots; Based on the preset first-step loss value, the preset second-step loss value, and the relative decision-after state value function, the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals are determined. as well as In the current time slot, the multiple wireless terminals used for edge transmission are controlled to perform a first action, the remaining wireless terminals are controlled to perform a second action to obtain a post-decision state, and the load of each wireless terminal is observed. Based on the post-decision state and the observation results, the relative state value function of each wireless terminal is calculated, and the relative state value function and Lagrange multiplier value of each wireless terminal are updated based on the relative state value function before entering the next time slot. The step of determining the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals based on a preset first-step loss value, a preset second-step loss value, and a relative post-decision state value function includes: obtaining the number of wireless terminals used for edge transmission and the current number of sub-channels, and determining the first action of each wireless terminal based on a preset first action strategy; when the number of wireless terminals used for edge transmission is greater than the current number of sub-channels, determining the multiple wireless terminals used for edge transmission based on the current number of sub-channels, and designating the other wireless terminals among the multiple wireless terminals as the remaining wireless terminals; and determining the second action of the remaining wireless terminals based on a preset second action strategy. The preset first action strategy is: ; in, The preset first action strategy, The first step loss value preset for edge server calculation. For the current bandwidth Lagrange multiplier, For wireless terminals i In the current energy Lagrange multiplier, for t Time Wireless Terminal i state, For the current moment, For action, The state-value function is relative to the decision. This is the state after the decision. i This is the serial number of the wireless terminal; ; in, for t Time Wireless Terminal i Information age AoI, For the current moment t With wireless terminals i The difference between sampling time slots of the task being processed. For wireless terminals i In the time slot The number of bits remaining for the task being processed. For the allocated bandwidth, Total power consumption of each wireless terminal per time slot; The preset second action strategy is: ; in, This is the preset second action strategy. This is the second-step loss value preset during local computation; The relative state value function is: ; in, The relative state value function, For wireless terminals i state, The value of the Lagrange multiplier; renew : ; renew : ; in, The state-value function is relative to the decision. Update the weights for Q-learning; Update each wireless terminal : ; renew : ; in, The energy Lagrange multiplier update rate, The power consumption of each wireless terminal in each time slot. For the bandwidth Lagrange multiplier update rate, For the allocated bandwidth, Maximum bandwidth constraint; For each wireless terminal, the initial settings are: .
2. A mobile edge computing network terminal collaborative scheduling device, characterized in that, include: The first determining module is used to determine multiple wireless terminals and divide multiple time slots; The second determining module is used to determine the first action of multiple wireless terminals used for edge transmission and the second action of the remaining wireless terminals based on the preset first-step loss value, the preset second-step loss value and the relative decision-after state value function. as well as The processing module is configured to control the plurality of wireless terminals used for edge transmission to perform a first action in the current time slot, control the remaining wireless terminals to perform a second action to obtain a post-decision state, observe the load of each wireless terminal, calculate the relative state value function of each wireless terminal based on the post-decision state and the observation results, and update the relative state value function and Lagrange multiplier value of each wireless terminal based on the relative state value function before entering the next time slot. The second determining module is specifically configured to: acquire the number of wireless terminals used for edge transmission and the current number of sub-channels, and determine the first action of each wireless terminal based on a preset first action strategy; when the number of wireless terminals used for edge transmission is greater than the current number of sub-channels, determine the plurality of wireless terminals used for edge transmission based on the current number of sub-channels, and designate the other wireless terminals among the plurality of wireless terminals as the remaining wireless terminals; and determine the second action of the remaining wireless terminals based on a preset second action strategy. The preset first action strategy is: ; in, The preset first action strategy, The first step loss value preset for edge server calculation. For the current bandwidth Lagrange multiplier, For wireless terminals i In the current energy Lagrange multiplier, for t Time Wireless Terminal i state, For the current moment, For action, The state-value function is relative to the decision. This is the state after the decision. i This is the serial number of the wireless terminal; ; in, for t Time Wireless Terminal i Information age AoI, For the current moment t With wireless terminals i The difference between sampling time slots of the task being processed. For wireless terminals i In the time slot The number of bits remaining for the task being processed. For the allocated bandwidth, Total power consumption of each wireless terminal per time slot; The preset second action strategy is: ; in, This is the preset second action strategy. This is the second-step loss value preset during local computation; The relative state value function is: ; in, The relative state value function, For wireless terminals i state, The value of the Lagrange multiplier; renew : ; renew : ; in, The state-value function is relative to the decision. Update the weights for Q-learning; Update each wireless terminal : ; renew : ; in, The energy Lagrange multiplier update rate, The power consumption of each wireless terminal in each time slot. For the bandwidth Lagrange multiplier update rate, For the allocated bandwidth, Maximum bandwidth constraint; For each wireless terminal, the initial settings are: .
3. An electronic device, characterized in that, include: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the mobile edge computing network terminal collaborative scheduling method as described in claim 1.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the mobile edge computing network terminal collaborative scheduling method as described in claim 1.