Mobile edge computing network joint scheduling and unloading method for optimizing information age

Through the near-end strategy optimization algorithm of parameterized action space, the problem of insufficient adaptability of traditional edge computing architecture in dynamic scenarios is solved, efficient resource scheduling of hybrid mobile edge computing networks is realized, and data freshness and reliability of computing services are improved.

CN120456124APending Publication Date: 2025-08-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510604092.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the dynamic scenarios such as smart mine trucks and smart logistics trucks, traditional edge computing architectures are difficult to adapt to the dynamic changes in the location of on-board servers, resulting in fluctuations in link quality, increased task transmission failure rate, deterioration of information age, and traditional reinforcement learning algorithms are difficult to achieve efficient exploration and dynamic adaptation of hybrid action space.

Method used

The near-end strategy optimization algorithm with parameterized action space is adopted, and discrete server selection and continuous offload rate allocation are coordinated through a hierarchical training mechanism, dynamically perceive channel status and server load, and joint decision-making task scheduling strategies to achieve accurate resource coordination of hybrid mobile edge computing networks.

Benefits of technology

It significantly improves the flexibility and decision-making efficiency of task scheduling, reduces link instability, reduces task failure rate, maintains data timeliness, shortens processing delays, ensures data freshness, and provides high-reliability and low-latency computing services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456124A_ABST
    Figure CN120456124A_ABST
Patent Text Reader

Abstract

The invention relates to a mobile edge computing network joint scheduling and unloading method for optimizing information age, and belongs to the field of mobile edge computing. In order to solve the problems of network topology dynamic change, complex hybrid action space decision and data freshness guarantee caused by mobility of a vehicle-mounted server, a near-end strategy optimization algorithm based on a parameterized action space is provided, discrete server selection and continuous unloading rate allocation are collaboratively optimized through a hierarchical training mechanism, and the network topology dynamic change and the hybrid action space decision are optimized. And dynamically sensing a channel state and a server load, and jointly deciding a task scheduling strategy. According to the method, dynamic network state modeling and hybrid action space optimization are fused, precise coordination of local and edge computing resources is realized, task processing time delay is reduced, the adaptability of a system to topology change and load fluctuation is improved, and long-term data freshness and scheduling stability in a mobile edge computing scene are effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile edge computing and relates to a mobile edge computing network joint scheduling and offloading method for optimizing information age. Background Art

[0002] With the rapid development of mobile communication technology, mobile edge computing (MEC) provides low-latency, highly reliable computing services to terminal devices by sinking computing power to the edge of the network. However, in dynamic scenarios such as smart mining vehicles and smart logistics vehicles, traditional edge computing architectures face multiple challenges: mobility of onboard servers, time-varying network topologies, and ensuring data freshness. Information age, a key metric for measuring data timeliness, requires comprehensive consideration of multiple factors, including transmission latency, task processing time, and update intervals. This places higher demands on resource scheduling in hybrid MEC environments.

[0003] In existing technologies, scenarios where fixed servers and vehicle-mounted edge servers work together usually adopt static task allocation strategies, which fail to effectively adapt to the dynamic changes in the location of the vehicle-mounted servers. Since the coverage of the vehicle-mounted server continuously changes with movement, the transmission distance and channel conditions between the device node and the server are time-varying, making it difficult for traditional scheduling algorithms based on fixed topology to maintain the optimal information age. In addition, the significant differences in computing power and coverage between fixed servers and vehicle-mounted servers have increased the design complexity of partial offloading strategies: on the one hand, the task offloading rate needs to dynamically match the real-time load and computing resources of the two types of servers; on the other hand, the parallel execution characteristics of local computing and edge computing require precise coordination of task division to minimize overall processing latency.

[0004] Current approaches to optimizing the long-term average information age in hybrid server environments have the following limitations: First, they fail to fully account for link quality fluctuations caused by the mobility of onboard servers, resulting in increased task transmission failure rates and degraded information age. Second, traditional reinforcement learning algorithms struggle to efficiently explore the hybrid action space for the joint decision-making process of discrete server selection and continuous offloading rate allocation in partial offloading. Third, fixed strategies cannot dynamically adapt to changes in server load and the randomness of task arrival, resulting in uneven utilization of computing resources and reduced data freshness. Therefore, an intelligent scheduling method that integrates dynamic network state awareness, hybrid action space optimization, and long-term information age modeling is urgently needed to improve system performance in complex mobile edge computing scenarios. Summary of the Invention

[0005] In view of this, the object of the present invention is to provide a mobile edge computing network joint scheduling and offloading method for optimizing information age.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] A joint scheduling and offloading method for mobile edge computing networks that optimizes information age. In a hybrid mobile edge computing network system model where fixed servers and vehicle-mounted edge servers coexist, a portion of computing tasks is offloaded to the edge servers through a partial offloading mode, fully utilizing the computing resources of the edge servers and terminal devices to meet the information freshness of the hybrid mobile edge computing network system. This method performs hierarchical training of discrete and continuous actions based on a proximal policy optimization algorithm in a parameterized action space, minimizing the long-term average information age of the system while satisfying the resource constraints of the edge servers, link conflict constraints, and the location constraints of the vehicle-mounted mobile servers.

[0008] The method specifically comprises the following steps:

[0009] S1: Obtain relevant parameter information of the hybrid mobile edge computing network system and build a system model by comprehensively considering server computing resource constraints, link conflict constraints, and vehicle-mounted edge server location constraints;

[0010] S2: Based on the constructed system model and the optimization goal of minimizing the long-term average information age of the system, establish the state space, action space and reward function;

[0011] S3: Updates the source information age, destination information age, server status flag, and remaining time slots of each terminal device node. Calculates the total time slots required for each terminal device node to be offloaded to the two types of edge servers in the current time slot. Updates the action space of the parameterized action space proximal policy optimization algorithm and outputs the optimal "discrete-continuous" action set based on the current state.

[0012] S4: Train the network and calculate the loss function of each network. Then update the network parameters through the gradient descent algorithm until the network converges, and finally obtain the optimal offloading scheduling strategy under the hybrid mobile edge computing partial offloading mode.

[0013] Furthermore, S1 specifically includes the following steps:

[0014] S11: In a hybrid mobile edge computing system consisting of M-1 vehicle-mounted edge servers, a fixed server, and N terminal device nodes, the computing tasks generated by the device terminal nodes are transmitted to the server nodes through wireless channels for processing; the server index is recorded as Where m = 0 represents a fixed server; in addition, a separate set of vehicle-mounted edge servers is defined as The index of the device node is defined as And its position is fixed; suppose the system is based on time slots, Represents the index of the time slot; at the beginning of each time slot, the system can schedule at most min{M,N} device nodes' data packets for conflict-free transmission through the orthogonal channel. Considering the distortion and thermal noise of the transmission channel, data transmission is not always successful. After the data packet is successfully transmitted, it will be calculated and processed on the corresponding server, and the calculation result will be returned to the corresponding terminal device node; the uplink transmission rate of the terminal device calculation task in time slot t is:

[0015]

[0016] in, is the Gaussian noise power, B is the bandwidth, is the transmission power of device node n, h n,m (t) is the wireless channel gain from device node n to server node m, and its calculation expression is:

[0017] h n,m (t)=h0(d n,m (t)) -2

[0018] Where h0 is the unit power gain of the channel; d n,m (t) is the Euclidean distance from the terminal device node to the target edge server;

[0019] S12: When a device node generates a computing task, the scheduling system needs to make a decision by comprehensively considering multiple factors such as the available computing resources of the local device, the current load status of each edge server, the quality of the wireless channel, and the real-time requirements of the computing task. The scheduling system dynamically determines the selection of the offloading target server and the offloading rate ρ n (t)∈[0,1], the task offloading rate is defined as:

[0020]

[0021] Among them, D n (t) represents the total amount of data generated by terminal device node n in time slot t, It represents the amount of data unloaded from the terminal device node n to the edge server in time slot t.

[0022] S13: In the partial offloading mode, the data of the computing task is divided into two parts: one part is retained in the local device for execution, and the other part is offloaded to the edge server for processing; specifically, when the device node generates a computing task, ρ n (t) data is offloaded to the edge server for calculation, 1-ρ n The data portion of (t) will remain in local execution; define To express the time consumed by the computing task in local calculation, its expression is:

[0023]

[0024] in, Indicates the calculation frequency of device node n, It represents the amount of data unloaded from the terminal device node n to the edge server in time slot t.

[0025] S14: Another part of the data volume of the computing task is offloaded to the edge server for calculation. There are two types of edge servers: VEC server and base station server. Due to the differences in computing power, transmission characteristics and coverage of the two types of servers, it is necessary to analyze the total time consumed for transmission and calculation respectively. n (t) = {0, 1, 2} represents the task offloading decision of device node n in time slot t, specifically: n (t) = 0 means that the server node does not schedule the device node n in time slot t; n (t) = 1 means that the device node n is scheduled to the vehicle-mounted edge server for computing and processing in time slot t; n (t) = 2 means that the device node n is scheduled to the fixed server for computing processing in time slot t;

[0026] When uninstalling decision O n When (t) = 1, the device node transfers part of the offloaded computing task to the nearby VEC server for calculation; since the transmission rate remains unchanged during the offload process, the total delay required for the offloaded part to complete the computing task is The calculation expression is:

[0027]

[0028] in, represents the total time required for device node n to offload the computing task to the vehicle-mounted mobile server m in time slot t, D n (t) represents the computing task of device node n in time slot t, R n,m (t) represents the uplink transmission rate from device node n to server node m at time slot t, C represents the number of CPU cycles required to process 1 bit of data, represents the calculation frequency of the vehicle-mounted mobile server m;

[0029] Since local computing and edge offloading are performed in parallel, the total time taken by the device node to complete the computing task is It can be determined by the larger value of the local computing time and the edge computing offloading time:

[0030]

[0031] in, It represents the total time required to complete the calculation by offloading to the vehicle edge server. To represent the time taken to compute the task locally.

[0032] S15: When the uninstall decision O n (t) = 2, the device node n transmits the offloaded computing task to the fixed server for computing. Since the device node and the fixed server are fixed in position, the uplink transmission rate does not change, so the total time required to complete the computing task is Calculated as:

[0033]

[0034] in, Indicates the calculation frequency of a fixed server.

[0035] The total time it takes to complete the computational task under this decision for:

[0036]

[0037] in, Indicates the total time required to complete the calculation by offloading to a fixed server. To represent the time taken to compute the task locally.

[0038] Furthermore, S2 specifically includes the following steps:

[0039] S21: Construct the optimization objective of minimizing the long-term average information age of the system, expressed as:

[0040]

[0041] Where t represents the time slot, n represents the terminal device node index, Indicates the initialization state information age of the system destination;

[0042] The constraints are as follows:

[0043] The fixed server can cover all nodes with a coverage radius of R, so the vehicle-mounted edge server is within the coverage radius of the fixed server:

[0044]

[0045] in, Indicates the coordinates of the vehicle-mounted mobile server, represents the coordinates of the fixed edge server;

[0046] When the terminal device node's offloading decision is 0 nWhen (t) = 1, the terminal device node must be within the coverage of the vehicle-mounted edge server, and its coverage radius is r. The constraint is expressed as:

[0047]

[0048] in, Indicates the coordinates of the vehicle-mounted mobile server, Indicates the coordinates of the terminal device node;

[0049] At time slot t, the edge server can only schedule the computing task of one terminal device node, and the computing task of a terminal device node can only be scheduled by one edge server:

[0050]

[0051] Among them, y n,m (t)∈{0,1} indicates the establishment of a link between the terminal device n and the edge server m;

[0052] At time slot t, the offloading rate ρ when the computing task of the terminal device is offloaded to the edge server n (t) should meet the following requirements:

[0053] 0≤ρ n (t)≤1

[0054] The three scheduling strategies can be expressed as:

[0055] O n (t)∈{0,1,2}

[0056] S22: Establishing the system state space S in the mobile edge computing partial offloading mode t for:

[0057]

[0058] in, represents the information age of the computing task data packet of the terminal device node n in the buffer, represents the information age of the data packet of the terminal device node n at the destination, v(t)=(v0(t),...,v N-1 (t)) represents the optimal unloading flag of the device node at the time slot, i(t)=(i0(t),…,i M-1 (t)) represents the state of server m at time slot t, l(t)=(l0(t),…,l N-1 (t)) represents the number of time slots required to complete the computation task of the terminal device node n in the current time slot;

[0059] S23: The action space is expanded into two dimensions. First, the edge server selects the node index for scheduling, and then determines the node's computing task offloading rate. The scheduling action of the server node is defined as selecting a node from the device node set for task scheduling, which is recorded as The action space of each server node in each time slot t is:

[0060]

[0061] Among them, 0 to N-1 represent the indexes of N device nodes. The server node does not perform any scheduling actions in the current time slot. This design takes into account the actual working status of the server node. When the server node's computing resources are occupied or the channel conditions do not meet the transmission requirements, the -1 action can be selected to avoid wasting computing resources.

[0062] The decision of the system also needs to determine the offloading rate of the corresponding node computing task, and define it as the offloading rate determination action Obviously, this action is a continuous action, and its value range is consistent with the definition range of the uninstall rate:

[0063]

[0064] By combining these two actions, the overall action form of the scheduling system can be further determined as:

[0065]

[0066] S24: Establish the reward function r(t):

[0067]

[0068] Among them, r b Represents the benchmark value of the system's long-term average information age, which is given by Calculated.

[0069] Furthermore, S3 updates the information age and action space of each terminal device node, specifically including the following steps:

[0070] S31: Data packets will age over time while waiting to be transmitted in the buffer, which will affect the evolution of the information age at the destination. For device nodes, the system model needs to consider the information age of its data packets when they are in the buffer and the information age when the computing task is completed. The information age of the buffer will affect the information age at the destination. Therefore, define The information age of the data packet of device node n in the buffer at time slot t is expressed as follows:

[0071]

[0072] Among them, g n (t)∈{0,1} indicates whether the device node n generates a computing task at time slot t. If g n (t) = 1, device node n generates a computing task in time slot t, otherwise g n (t) = 0;

[0073] S32: Definition l n (t)∈{0,1,…,c n,m (t)} represents the number of remaining time slots required for device n to receive the calculation result at the beginning of time slot t; let i m (t)∈{0,1} represents the state of the server in time slot t, where i m (t) = 1 means that server m is in computing state, otherwise i m (t) = 0; l n The updating process of (t) is:

[0074]

[0075] Among them, y n,m (t)∈{0,1} represents the state of device node n being scheduled by server m at the beginning of time slot t, y n,m (t) = 1 means that the computing task of device node n is scheduled to server node m, otherwise y n,m (t) = 0; k n,m (t)∈{0,1} indicates whether the device node n successfully sends the data packet to the server node m;

[0076] S33: When the server node completes the computing task, it transmits the result back to the corresponding device node; represents the information age of the destination device node n, and The update process is expressed as:

[0077]

[0078] Among them, l n (t) represents the number of data packets remaining from the device node n task calculated by the server node in time slot t, i m (t) represents the state of server m in the time slot.

[0079] S34: By rounding up the function Discretize the total time into the number of time slots, using a finite positive integer c n,m (t), so the total time slot to complete the computing task under different strategies is:

[0080]

[0081] in, Represents the uninstall decision O n The total time when (t) = 1, Represents the uninstall decision O n (t) = total time of 2 hours;

[0082] S35: Design a proximal policy optimization algorithm based on parameterized action space (Hybrid Dynamic Scheduling Proximal Policy Optimization, HDS-PPO), which uniformly encodes discrete server selection actions and continuous offload rate actions into high-dimensional continuous vectors, achieving a unified representation of the hybrid action space; specifically, HDS-PPO uses two parallel actor networks. The discrete actor network selects discrete actions, that is, selects the index of the scheduling node; the continuous actor network is responsible for selecting continuous parameters, in this case, assigning appropriate offload rates to discrete actions; the global critic network is used to evaluate the value function of the state to guide the training of all child actor networks.

[0083] Furthermore, in S4, the following steps are specifically included:

[0084] S41: After S3, the DHS-PPO algorithm is obtained. First, it interacts with the environment according to the current strategy and collects the state s(t), action a(t), and reward r(t).

[0085] S42: Calculate the advantage value of the current action a(t) relative to the average performance of the strategy in the state s(t). The advantage function of the two-level action can be calculated as:

[0086] A t =Q((s(t),a(t))-V(s(t))

[0087] Among them, s(t) represents the state value, a(t) represents the action, and V(s(t)) represents the value function of the current state;

[0088] S43: The critic network is able to accurately estimate the state value by minimizing the mean square error loss function, and then update the parameters by gradient descent. The loss function can be expressed as:

[0089] L(θ t )=E t [(V(s(t);θ t )-R t ) 2 ]

[0090] Among them, V(s(t);θ t ) represents the value function prediction of the current state, Rt Indicates cumulative rewards;

[0091] S45: After multiple rounds of iteration, the strategy will be continuously optimized until the network converges.

[0092] The beneficial effects of the present invention are:

[0093] The present invention unifies the encoding of discrete server selection and continuous offloading rate decision-making through parameterized action space design, realizes the collaborative optimization of hybrid action space, and significantly improves the flexibility and decision-making efficiency of task scheduling in dynamic environments. In view of the mobility characteristics of vehicle-mounted edge servers, this method dynamically perceives network topology changes and channel quality fluctuations, adaptively adjusts the offloading strategy, effectively reduces the link instability problem caused by changes in transmission distance, thereby reducing task failure rate and maintaining data timeliness. Through the deep reinforcement learning framework, local computing and edge computing resources are jointly optimized, the task division ratio is accurately coordinated, and computing resources are maximized under the parallel processing mechanism, the overall task processing delay is shortened, and end-to-end data freshness is guaranteed. In addition, the hierarchical training mechanism based on proximal strategy optimization solves the defects of low exploration efficiency and slow strategy convergence of traditional algorithms in hybrid action space. Through the coordinated update of global state value evaluation and local action strategy, the system's adaptability to server load fluctuations and random arrival of tasks is enhanced, and finally the long-term average information age is minimized in complex mobile edge computing scenarios, providing highly reliable and low-latency computing service support for real-time sensitive applications such as smart vehicles and industrial Internet of Things.

[0094] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0096] Figure 1 This is a training diagram of the hybrid mobile edge computing scheduling algorithm DHS-PPO based on deep reinforcement learning in the present invention;

[0097] Figure 2 Flowchart of the joint scheduling and offloading method for mobile edge computing networks that optimizes information age according to the present invention. DETAILED DESCRIPTION

[0098] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0099] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0100] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0101] See also Figures 1 and 2 This paper targets hybrid mobile edge computing scenarios, comprehensively considering constraints such as the computing resources of terminal devices and edge servers, link conflicts, and the location of onboard edge servers to minimize the system's long-term average information age. Using a deep reinforcement learning method with a parameterized action space, the discrete server action selection and the continuous offload rate action are uniformly encoded into a high-dimensional continuous vector. Through iterative optimization of the strategy, a scheduling method suitable for mobile edge computing systems in partial offload mode is obtained.

[0102] Figure 1 The following is a training diagram of the hybrid mobile edge computing scheduling algorithm DHS-PPO based on deep reinforcement learning. Figure 1As shown in Figure 1, the HDS-PPO algorithm achieves flexible control of complex partial offloading problems by collaboratively training discrete and continuous action policies. The algorithm's training process begins with the agent interacting with the environment, extracting environmental features using a shared state encoding network. Action distributions are then generated using discrete and continuous action heads, respectively. The discrete action head outputs indexed probabilities based on a Softmax strategy, while the continuous action head uses a Gaussian strategy to generate real-valued actions. Together, they form a hybrid action space. In each iteration, the algorithm collects the state, action, and reward trajectories generated under the current policy and uses generalized advantage estimation to calculate the advantage function value at each time step to evaluate the long-term benefits of the actions. The algorithm simultaneously optimizes the loss functions of both discrete and continuous policies and reuses sampled data across multiple rounds of mini-batch updates to improve sample efficiency. Ultimately, a dynamic scheduling method suitable for partial offloading is obtained.

[0103] Figure 2 The flowchart of the computing network scheduling method assisted by vehicle-mounted edge computing for optimizing information age of the present invention is as follows: Figure 2 As shown, the specific steps include:

[0104] V1~V3: Obtain parameter information of the hybrid mobile edge computing system in partial offloading mode, parameters of each network in the initial algorithm, and parameters of the shared module.

[0105] V4~V5: Establish the system's state space, action space, and reward function, obtain the location of the on-board edge server, calculate the theoretical time slot for each device node to complete the calculation under different offloading objects, and update the system's state space.

[0106] V6-V12: Continuously interacts with the current environment to collect new state values, action values, and rewards, copies the parameters of the actor network to the old actor network, calculates the advantage function of each network's action, outputs the current action value and the action's log probability density, calculates the actor network's loss function, and then uses gradient descent to update the actor network's parameters.

[0107] V13~V16: Input the state into the Critic network, calculate the state value, and use gradient descent to update the network parameters to guide the Actor network to select the optimal action.

[0108] V17~V18: After determining whether the loss function is stable and the training termination conditions are met, if the termination conditions are not met, execute V6. If they are met, output the hybrid dynamic scheduling method based on the fully trained scheduling network.

[0109] Example 1: Dynamic Scheduling Process of an In-Vehicle Edge Server

[0110] Workflow:

[0111] 1. Environmental perception: The vehicle-mounted edge server reports location coordinates and coverage radius in real time, while the fixed base station collects device node task queue length, local computing resource status, and channel quality data.

[0112] 2. State encoding: The source information age, destination information age, server load status, and remaining task time slots of the device node are encoded into a multidimensional state vector and input into the shared feature extraction network of the HDS-PPO algorithm;

[0113] 3. Action generation: The discrete Actor network outputs the server selection action and selects the vehicle-mounted or fixed server for execution scheduling; the continuous Actor network generates the offload rate parameters for the corresponding devices and determines the task division ratio;

[0114] 4. Execution and feedback: Execute task offloading based on joint actions, calculate actual transmission delay and processing delay, update information age indicator and generate reward signal;

[0115] 5. Model update: The critic network evaluates the state value, calculates the policy gradient and updates the discrete and continuous actor network parameters, iteratively optimizing the scheduling strategy.

[0116] Effect: Dynamically adapt to coverage changes caused by the movement of the vehicle-mounted server, reduce the probability of task interruption through hybrid action collaboration, and improve the timeliness of data updates.

[0117] Example 2: Joint Optimization Process of Hybrid Action Space

[0118] Workflow:

[0119] 1. Task triggering: When a device node generates a computing task, the HDS-PPO algorithm obtains the current channel gain, server computing queue length, and device local resource occupancy rate;

[0120] 2. Discrete action generation: The discrete Actor network filters the available server set based on server coverage status and load level, and outputs the optimal scheduling node index;

[0121] 3. Continuous Action Generation: The continuous actor network generates task offload rate parameters based on the remaining computing capacity of the selected server and channel conditions, dynamically allocating local and edge computing loads.

[0122] 4. Parallel processing: Local computing and edge offloading are performed simultaneously. The system monitors the progress of both processes and takes the longest processing time as the task completion time.

[0123] 5. Strategy Optimization: The critic network calculates the advantage function based on the change in information age, constrains the strategy update range through the proximal strategy optimization algorithm, and balances exploration and utilization efficiency.

[0124] Effect: Achieve precise matching of server selection and offload rate, maximize parallel computing resource utilization, and shorten end-to-end task processing latency.

[0125] Example 3: Closed-loop control process for long-term information age optimization

[0126] Workflow:

[0127] 1. Periodic sampling: The system collects information age status, server location, and task queue data of all network devices at fixed time slots to build a historical status trajectory data set;

[0128] 2. Offline training: Use historical data to pre-train the HDS-PPO model, the critic network learns the state-value function, and the discrete and continuous actor networks initially establish a hybrid action strategy;

[0129] 3. Online deployment: Load the pre-trained model to the edge controller, which receives real-time environmental status and outputs scheduling decisions while also recording online interaction data.

[0130] 4. Incremental Update: Periodically fine-tune actor network parameters based on online data, and adjust action distribution using a weighted advantage function to adapt to server load fluctuations and changes in task arrival patterns.

[0131] 5. Closed-loop feedback: Dynamically adjust the training cycle and exploration rate parameters based on the information age optimization effect to maintain the long-term stability and environmental adaptability of the strategy.

[0132] Effect: Build a learning-execution-feedback closed loop, continuously optimize the system's adaptability to dynamic environments, and ensure the long-term stability of data freshness.

[0133] Example 4: Multi-device scheduling process under conflict constraints

[0134] Workflow:

[0135] 1. Conflict detection: The system detects the scheduling actions of each server in the current time slot and verifies whether the constraints of "a single server can schedule at most one device, and a single device can connect to at most one server" are met;

[0136] 2. Action correction: For scheduling actions that violate constraints, the scheduling requests of devices with high information age are retained first, and the remaining actions are invalidated and the re-decision mechanism is triggered;

[0137] 3. Resource reallocation: Recalculate available server-device pairings and generate new offload rate parameters based on the revised state to ensure conflict-free utilization of channel resources.

[0138] 4. Penalty mechanism: imposes negative rewards on conflicting actions to guide the policy network to avoid resource competition scenarios in subsequent decisions;

[0139] 5. Conflict history records: Conflict events and handling results are stored in the experience pool to enhance the model's ability to learn complex constraints.

[0140] Effect: Effectively avoids transmission failures caused by channel resource competition, and improves scheduling success rate and system robustness in multi-device concurrent scenarios.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A joint scheduling and offloading method for mobile edge computing networks that optimizes information age, characterized by: In a hybrid mobile edge computing system consisting of fixed servers and vehicle-mounted edge servers, computing tasks are dynamically allocated to edge servers and local devices through partial offloading mode. A proximal policy optimization algorithm with a parameterized action space is used to jointly train discrete and continuous actions. The system includes the following steps: S1: Obtain parameter information of the hybrid mobile edge computing network system and build a system model that includes server computing resource constraints, link conflict constraints, and vehicle server location constraints; S2: Based on the system model, an optimization objective of minimizing the long-term average information age is established, and a state space, an action space, and a reward function are defined; wherein the state space includes the source information age of the terminal device, the destination information age, the server status flag, and the number of remaining time slots of the task; and the action space includes a discrete server selection action and a continuous task offloading rate allocation action. S3: Based on the communication parameters of each terminal device node in the current time slot, the time slots required for offloading to different servers are calculated. The source information age, destination information age, and number of remaining time slots of each node are updated. The optimal "discrete-continuous" action set is generated through the proximal policy optimization algorithm of the parameterized action space. S4: Calculate the policy gradient through the reinforcement learning algorithm and optimize the parameters of the discrete and continuous action policy networks until the network converges to obtain the optimal unloading scheduling strategy.

2. The method for joint scheduling and offloading of mobile edge computing networks with optimized information age according to claim 1, characterized in that: In S1, the system model includes M-1 vehicle-mounted edge servers, 1 fixed server, and N terminal device nodes. The coverage range of the vehicle-mounted edge server changes dynamically, and the coverage radius of the fixed server is R. At time slot t, the uplink transmission rate from device node n to server m is: Where B is the bandwidth, is the transmission power of device node n, is the Gaussian noise power, h n,m (t)=h0(d n,m (t)) -2 is the channel gain, d n,m (t) is the Euclidean distance between the device node and the server, and h0 is the unit power gain coefficient.

3. The method for joint scheduling and offloading of mobile edge computing networks with optimized information age according to claim 2, characterized in that: In the partial unloading mode, the unloading rate ρ of the device node n in time slot t n (t) is defined as: Among them, D n (t) is the total task data volume, The amount of data offloaded to the edge server; local computing time for: Where C is the number of CPU cycles required to process unit data, Calculates the frequency locally.

4. The method for joint scheduling and offloading of mobile edge computing networks with optimized information age according to claim 3 is characterized in that: When the device node selects the vehicle-mounted edge server to unload, the total time Including transmission time and server calculation time: in, is the calculation frequency of the vehicle server; represents the total time required for device node n to offload the computing task to the vehicle-mounted mobile server m in time slot t; the total delay for task completion is the larger value of local computing and edge computing:

5. The method for joint scheduling and offloading of mobile edge computing networks with optimized information age according to claim 3, characterized in that: When the device node selects a fixed server to uninstall, the total time consumed for: in, is the calculation frequency of the fixed server; the total delay for task completion is:

6. The mobile edge computing network joint scheduling and offloading method for optimizing information age according to claim 1 is characterized by: In S2, the optimization goal is to minimize the long-term average information age: in, is the destination information age of device node n in time slot t; the constraints include: The distance between the vehicle-mounted server and the fixed server shall not exceed R; The distance between the device node and the vehicle server does not exceed its coverage radius r; Each server schedules at most one device node in time slot t; Uninstall rate ρ n (t)∈[0,1].

7. The method for joint scheduling and offloading of mobile edge computing networks with optimized information age according to claim 1, characterized in that: In S3, the proximal policy optimization algorithm of the parameterized action space adopts a discrete Actor network to select the server to schedule the action. Continuous Actor Network Spawn Offload Rate in Indicates that scheduling is not performed; the joint actions are:

8. The method for joint scheduling and offloading of mobile edge computing networks with optimized information age according to claim 7, characterized in that: The reward function is: where r b It is the benchmark value of the system's long-term average information age.

9. The mobile edge computing network joint scheduling and offloading method for optimizing information age according to claim 1 is characterized in that: In S3, updating the source information age of each node includes: Source information age Reset to 1 when a new task is generated, otherwise incremented; Destination Information Age Updated to when the task is completed Otherwise increment; Remaining calculation time slot l n (t+1) Dynamically updated according to server status.

10. The mobile edge computing network joint scheduling and offloading method for optimizing information age according to claim 1 is characterized in that: In S4, by calculating the advantage function A t =Q((s(t),a(t))-V(s(t))Optimize the policy network, and the state value loss function of the critic network is: L(θ t )=E t [(V(s(t);θ t )-R t ) 2 ] Among them, R t is the cumulative reward, θ t is the network parameter.