Satellite communication network optimization method and system based on MEC

By applying the MADDPG algorithm in a network combined with MEC and satellite communication, we optimize task placement, access control, service instance selection and bandwidth allocation, solving the problem of system performance imbalance, achieving more efficient resource utilization and lower economic costs.

CN120075060APending Publication Date: 2025-05-30TIANYUAN RUIXIN COMM TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510074075.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When using edge computing (MEC) with satellite communications, there are problems such as unreasonable task offloading strategies, inefficient service placement strategies, unreasonable computing resource allocation and high environmental dynamics, resulting in unbalanced system performance.

Method used

A joint optimization algorithm is designed to optimize task placement and removal, access control, business instance selection and bandwidth allocation through the multiagent depth deterministic policy gradient (MADDPG) algorithm to ensure optimal operation results in a dynamic environment.

Benefits of technology

Through joint optimization algorithms, the system performance is significantly improved, the long-term average weighted sum of task failure rate and the economic costs of IoT devices within a given time period is reduced, and the total expenditure efficiency of the system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075060A_ABST
    Figure CN120075060A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of wireless communication, discloses a satellite communication network optimization method and system based on MEC, and aims to meet the time delay requirement of quality of service (QoS) and improve the service performance of the satellite communication network through joint optimization of task scheduling, access control, service instance selection and bandwidth allocation. And the long-term average weighted sum of the task failure rate and the equipment economic cost is minimized. According to the method, the mobile edge computing technology is utilized, the performance and reliability of a communication network are improved, the problems of wireless resource shortage and insufficient computing resources of a user side are solved, and energy consumption and time delay in the computing unloading process of equipment are reduced. Aiming at task dynamics, network condition volatility and complexity of an optimization problem in a satellite ground integrated system, the problem is converted into a deep reinforcement learning (DRL) problem, a joint optimization algorithm based on a multi-agent depth deterministic policy gradient (MADDPG) is provided, and the task failure rate and the economic cost are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a joint optimization algorithm, system, medium, device and application for task placement and removal, access control, service instance selection and bandwidth allocation. Background Art

[0002] Edge computing technology (MEC) significantly improves the speed and efficiency of data processing by deploying computing and storage resources at the network edge. As a communication method with wide coverage, large capacity and high reliability, satellite communication has become a research hotspot when integrated with the terrestrial communication network. Combining MEC with satellite communication, especially in areas where the terrestrial network is unreachable or unstable, can provide critical low-latency and high-bandwidth communication services. By deploying MEC near satellite ground stations or user terminals, the need to transmit data to remote data centers can be reduced, thereby reducing communication latency and improving service response speed. This is crucial for applications that require real-time data processing. Therefore, the framework of combining MEC with satellite communication not only expands the network coverage, but also improves the network performance and reliability, providing an efficient and reliable communication solution for modern society to meet the communication needs of users in various environments.

[0003] When using MEC in combination with satellites, several challenges need to be addressed.

[0004] (1) Existing research mainly focuses on the optimization of task offloading. However, the key step is to pre-deploy task applications on satellite edge servers. Therefore, in order to successfully offload tasks to Internet of Things (IoT) devices, a reasonable task offloading strategy needs to be designed.

[0005] (2) The caching capacity of satellites is limited, and it is impossible to deploy all task applications on each satellite. Therefore, it is crucial to design an efficient service placement strategy to reduce the communication distance between IoT devices and service-providing satellites.

[0006] (3) When an IoT device connects to a satellite, the satellite has two options: one is to use existing idle service instances to serve the device; the other is to start a new service instance to specifically manage the device. Therefore, in the case of multiple terminals entrusting tasks to a single satellite, the satellite needs to reasonably allocate computing resources among these terminals to ensure service efficiency.

[0007] (4) In satellite-based systems, the environmental dynamics are high, and the goal setting often focuses on the long term. Therefore, when developing algorithms, it is necessary to consider the algorithm complexity and performance while maximizing the goal to achieve optimal operating results.

[0008] Through the above analysis, based on the existing research and the problems found, the present invention: (1) designs an integrated satellite-ground network using MEC, which consists of a group of LEO satellites with MEC capabilities and a series of IoT devices. Each satellite and its deployed edge server are considered as a unified entity. (2) By jointly optimizing task placement and removal, access control, service instance selection, and bandwidth allocation between satellites and IoT devices, the total system expenditure is reduced over a long period of time. Among them, the optimization is constrained by factors such as the number of service instances, the remaining space of satellites, the bandwidth allocation between devices, and the number of deployed tasks. (3) Considering the dynamics of task arrival and the volatility of network conditions. We use the enhanced MADDPG algorithm to comprehensively understand the environmental dynamics and derive the most effective collaborative decision-making strategies for task allocation, removal, access control, task instance selection, and bandwidth allocation. Summary of the Invention

[0009] In view of the problems existing in the prior art, the present invention provides a joint optimization algorithm, system, medium, device, and application for task placement and removal, access control, service instance selection, and bandwidth allocation.

[0010] The present invention is implemented as follows. An optimization method for a satellite communication network based on MEC, the optimization method for the satellite communication network based on MEC includes the following steps:

[0011] First step: The current network and the target network of each agent jointly form the model of the training network, where the current network includes the current policy network actor and the current value network critic, and the target network includes the target policy network actor and the target value network critic; establish all these network models.

[0012] Second step: Initialize the relevant parameters in the network.

[0013] Third step: Initialize the state o k (t) of each agent, and at the same time obtain the global initial state s(t).

[0014] Fourth step: Set the time t = t0.

[0015] Fifth step: The agent k obtains the action a k = μ k (s k |θ k u ) of the agent through the policy network a k at time t. k (t).

[0016] Sixth step: Reconstruct the action a k (t).

[0017] Step 7: Determine whether k is equal to K. If not, return to Step 5; otherwise, execute Step 8.

[0018] Step 8: Obtain the global action a(t).

[0019] Step 9: Execute the action a(t) to obtain the reward and the state s(t + 1) at the next moment.

[0020] Step 10: Store the experience tuple <s(t), a(t), r(t), s(t + 1)> in the buffer.

[0021] Step 11: Determine whether the buffer is full. If so, take out mini-batch samples from the buffer to train the network; if not, determine whether t is equal to T max , if so, jump back to execute Step 3; otherwise, execute Step 12.

[0022] Step 12: Calculate the gradient of the current value network corresponding to the agent and update the current policy network through gradient ascent.

[0023] Step 13: Calculate the value of the action of the agent at time t using the current value network, and calculate the value of the action at time t + 1 using the target network, and then update the current value network through gradient descent.

[0024] Step 14: Update the policy network and the value network in the target network using soft update.

[0025] Step 15: Determine whether convergence occurs. If so, obtain the final optimal solution; otherwise, jump back to execute Step 3.

[0026] Furthermore, the parameters initialized in the second step include:

[0027] Initialize the weight parameters θ j and w j of each actor and critic network, and initialize the parameters of each target actor and critic network and Initialize the learning rates α and β, the discount factor γ, the number of episodes EP, and the maximum number of training steps T for each episode corresponding to the critic and actor networks. max . Initialize the replay buffer size D, the size M of the mini-batch, and the random process Ψ for action exploration. Initialize the network layout parameters, such as the number N of IoT devices, the number K of satellites, and task parameters.

[0028] Furthermore, in the third step, the state of the agent and the global state are respectively represented as:

[0029]

[0030] where represents the task requests of IoT devices covered by satellite k at the beginning of time step t; represents the number of occurrences of task q on satellite k at the end of time slot t - 1, i.e., at the beginning of time slot t; is the remaining storage space on the satellite at the beginning of time slot t; is the path loss between the IoT device and the satellite; is to re - allocate the satellite computing power to the IoT device.

[0031] Furthermore, in the fifth step, the obtained action a j (t) is represented as follows:

[0032]

[0033] where, is the task placement decision of satellite k; represents access control; represents the instance selection strategy; represents the bandwidth allocation between satellite k and IoT device n.

[0034] Furthermore, in the ninth step, the calculation formula of the reward is as follows:

[0035]

[0036] Furthermore, in the twelfth step, the current policy network is updated by gradient ascent as follows:

[0037]

[0038] Furthermore, in the thirteenth step, the update process of the current value network by the gradient descent method is as follows:

[0039]

[0040] where and respectively represent the TD target and the TD error.

[0041] Furthermore, in the fourteenth step, the soft update formula of the target network is as follows:

[0042]

[0043] where θj represent the parameters of the current policy network, represent the parameters of the target policy network, w j represent the parameters of the current value network, represent the parameters of the target value network,

[0044] Another object of the present invention is to provide a wireless communication information data processing terminal, and the wireless communication information data processing terminal is used to implement the method for jointly optimizing task placement and removal, access control, service instance selection, and bandwidth allocation.

[0045] Another object of the present invention is to provide a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method for jointly optimizing task placement and removal, access control, service instance selection, and bandwidth allocation.

[0046] In summary, the advantages and positive effects of the present invention are as follows: By jointly optimizing task placement and removal, access control, service instance selection, and bandwidth allocation, the present invention minimizes the long-term average weighted sum of the task failure rate and the economic cost of IoT devices within a given time period while ensuring task processing constraints. Specifically, we pre-install new task instances into the satellite and remove idle service instances from the satellite to free up space for new service deployment. In addition, we also make task offloading decisions, optimize access control policies, select the optimal service instance for each offloaded service request, and optimize resource allocation. Given the complexity of the problem and the dynamic characteristics of the network, we formulate the problem as a Markov Decision Process (MDP). Based on this, a joint optimization algorithm based on Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is proposed to achieve low complexity and real-time adaptive decision-making. Since MADDPG is designed to solve problems with continuous action spaces, and our problem contains integer, binary, and continuous variables, MADDPG cannot be directly applied to our scenario. To solve this problem, we take the following measures: First, we normalize the continuous variables; Second, we reformulate the continuous output of MADDPG as discrete and binary variables; and at the same time ensure that the coupling constraints between different variables are satisfied. Through these methods, we can effectively apply MADDPG to problems with mixed variables, thereby realizing the optimization of satellite service deployment and offloading.

[0047] The present invention proposes a satellite communication network architecture based on MEC, aiming to achieve task processing services with QoS guarantee. Different from existing research that mainly focuses on the task offloading process, the present invention comprehensively optimizes the success rate of the entire task offloading, covering multiple aspects such as service deployment, access control, service instance selection, and bandwidth allocation. The present invention considers the service deployment process, including the initialization and deletion of service instances, and optimizes these processes. In addition, in-depth research is carried out on access control, service instance selection, and bandwidth allocation during the task offloading process, and corresponding optimization strategies are proposed. Based on the actual dynamic environment, the present invention proposes an optimization method for dynamic cooperative task placement and removal, access control, service instance selection, and bandwidth allocation for the network combining MEC and satellite communication. Through comparative analysis with experimental data, it is confirmed that this dynamic resource allocation method has higher accuracy compared with the traditional static resource allocation method and can more realistically simulate the actual environment. Therefore, the method of cooperative task placement and removal, access control, service instance selection, and bandwidth allocation proposed by the present invention can effectively solve the problem of unbalanced performance between the MEC system and the satellite communication system caused by unreasonable computing resource allocation, thereby significantly improving the system performance.

[0048] The present invention also provides a joint optimization system for task placement and removal, access control, service instance selection, and bandwidth allocation based on deep reinforcement learning, including:

[0049] A training network module for establishing a current network and a target network for each agent, where the current network includes a current policy network (actor) and a current value network (critic), and the target network includes a target policy network (actor) and a target value network (critic);

[0050] A parameter initialization module for initializing the parameters in the network, setting the initial state of the agent, and obtaining the global initial state at the same time;

[0051] An action generation and reconstruction module for obtaining the action of the agent through the policy network and performing action reconstruction;

[0052] A reward calculation and state update module for calculating the reward value after executing the global action, updating to the state at the next moment, and storing the experience tuple in the replay buffer pool at the same time;

[0053] A gradient update module for calculating the values of the current and next moment actions using the current value network and the target network respectively, updating the policy network through gradient ascent, and updating the value network through gradient descent;

[0054] A target network update module for updating the target policy network and the target value network through a soft update method;

[0055] The convergence judgment module determines whether the convergence condition is satisfied. If it is satisfied, the final optimized solution is output. If not, it returns to execute the action generation and reconstruction module to continue the optimization iteration.

[0056] The system further includes:

[0057] The replay buffer module is used to store the experience tuples of the agents, and samples are drawn from the buffer pool using the mini-batch sampling method for optimizing network parameters;

[0058] The global state management module is used to manage the local states and actions of each agent, and generate the global state and global action to achieve the collaborative optimization of multiple agents;

[0059] The time step control module is used to dynamically adjust the time step t in the optimization process to ensure that action execution, reward calculation, and network parameter update are completed within each time step.

[0060] Compared with the existing MEC and satellite communication integrated system solutions, the present invention has achieved significant technological progress.

[0061] The method proposed by the present invention has significant advantages in terms of operational simplicity, real-time performance, and closeness to real scenarios, which is beneficial to network optimization and system performance improvement. By comprehensively considering the dynamic characteristics of MEC and satellite communication networks, the overall optimization of the task offloading process is realized, providing a new solution for improving network efficiency and user experience.

[0062] The present invention is a joint optimization of a mobile edge computing and satellite communication network integrated system, specifically related to a collaborative satellite communication network optimization method based on MEC, belonging to the field of wireless communication technology.

[0063] The present invention relates to a resource optimization method in a satellite-ground integrated Internet of Things (IoT) system supported by edge computing. By jointly optimizing task scheduling, access control, service instance selection, and bandwidth resource allocation, this method aims to minimize the long-term average weighted sum of the task failure rate and the economic cost of IoT devices within a given time period while meeting the quality of service (QoS) latency requirements of user tasks. By integrating mobile edge computing technology, the present invention aims to improve the performance and reliability of communication networks, alleviate the high system economic expenditure problem caused by wireless resource tension and lack of computing resources on the user side, and at the same time reduce the energy consumption and latency of mobile devices during the computing offloading process, thereby reducing the task failure rate and the economic cost of IoT devices. Given the task dynamics, network condition volatility, user real-time dynamics, and the mixed-integer non-linear characteristics of the problem in the satellite-ground integrated system, this optimization problem is extremely complex, and traditional convex optimization and dynamic programming methods are difficult to effectively solve. For the formulated optimization problem, the present invention transforms this joint optimization problem into a deep reinforcement learning (DRL) problem and proposes an efficient joint optimization algorithm for task placement and removal, access control, service instance selection, and bandwidth allocation based on multi-agent deep deterministic policy gradient (MADDPG). BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. Obviously, the following described drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0065] Figure 1 It is a flowchart of an MEC-based satellite communication network optimization method provided by an embodiment of the present invention.

[0066] Figure 2 It is a scenario diagram that can be applied provided by an embodiment of the present invention.

[0067] Figure 3 It is a flowchart of the implementation of the joint optimization method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the present invention in detail with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0069] In view of the problems existing in the prior art, the present invention provides an optimization method for a satellite communication network based on MEC. The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0070] As Figure 1 shown, the optimization method for a satellite communication network based on MEC provided by the present invention includes the following steps:

[0071] S101: For the training network of each agent, envelope the current network and the target network, where the current network includes the current policy network actor and the current value network critic, and the target network includes the target policy network actor and the target value network critic; establish all these network models.

[0072] S102: Initialize the parameters in the network.

[0073] S103: Initialize the state of each agent, and at the same time obtain the global initial state, and set the time t = t0.

[0074] S104: The agent obtains the action of the agent through the policy network and reconstructs the action.

[0075] S105: Obtain the global action, execute the global action, obtain the reward and the state at the next moment; store the experience tuple in the replay buffer pool.

[0076] S106: Judge whether the replay buffer is full. If so, take out mini-batch samples from the buffer, otherwise judge whether t is equal to T max , if so, step S104 is executed again.

[0077] S107: Calculate the gradient of the current value network, update the current policy network through gradient ascent; calculate the value of the action at time t using the current value network, and at the same time calculate the value of the action at time t + 1 using the target network, and then update the current value network through gradient descent.

[0078] S108: Use soft update to update the policy network and the value network in the target network.

[0079] S109: Judge whether it converges. If so, obtain the final optimal solution, otherwise start from step S104 and execute again.

[0080] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0081] Figure 2This is a scenario diagram to which the method of the present invention can be applied. The system includes a group of LEO satellites with MEC capabilities and a series of Internet of Things devices. The set of satellites and Internet of Things devices is denoted as K = {1, 2,..., k,..., K} and N = {1, 2,..., i,..., N}, where K and N represent the total amounts of satellites and Internet of Things devices respectively. Each satellite gets a value k from the set K, and each Internet of Things device gets a position n in the set N. In our system, the satellite and the edge server deployed on it are regarded as a unified entity, which can be interchangeably called satellite k or edge server k. Satellites can directly establish relevant communication connections through wireless links. The system operates using a time-slot method, where each time slot represents a different time unit for task assignment, deletion, and access management. The present invention uses T = {1, 2,..., t,..., T} to represent the set of time intervals, where T represents the total number of time periods within the system. The duration of each time interval is t. At the beginning of each time period, each Internet of Things device is responsible for processing a task that requires a large amount of computing resources. These tasks can be processed by the Internet of Things device itself or offloaded to the satellite at a certain cost to enhance the task execution ability. To minimize the task processing delay, a successful strategy includes installing service instances on the satellite before the user requests. Existing idle instances can also be deleted to free up space for deploying other instances of frequently requested services. In particular, the present invention refers to the first satellite as the main satellite, which has the strongest processing ability, and other satellites as general satellites, with relatively limited computing power.

[0082] In a related scenario, there are Q different service categories, where the task set is represented by Q = {1, 2,..., q,..., Q}, and the parameter is used to describe task Q. Among them, the variable D q (bits) represents the size of the input data related to the task; Ξ q (CPUcycles / bit) represents the task processing density, that is, the number of CPU cycles required to process a single bit of the task; According to D q and Ξ q , the amount of computation C q can be represented by C q = D q Ξ q ; O q (bits) represents the size of the result obtained by task processing; w q (bits) represents the size of the service instance of a specific service q; is the maximum allowed time to complete the task, setting the deadline for completing the task.

[0083] As Figure 3 shown, the collaborative MEC-based satellite communication network optimization method of the present invention includes the following steps:

[0084] Step 1: For each agent, the training network envelopes the current network and the target network, where the current network includes the current policy network actor and the current value network critic, and the target network includes the target policy network actor and the target value network critic; establish all these network models.

[0085] Step 2: Initialize the parameters in the network.

[0086] Step 3: Initialize the state o k (t) of each agent, and at the same time obtain the global initial state s(t).

[0087] Step 4: Set the time t = t0.

[0088] Step 5: The agent k obtains the action a of the agent k through the policy network k (t).

[0089] Step 6: Reconstruct the action a k (t).

[0090] Step 7: Judge whether k is equal to K. If not, return to Step 5; otherwise, execute Step 8.

[0091] Step 8: Obtain the global action a(t).

[0092] Step 9: Execute the action a(t) to obtain the reward and the state s(t + 1) at the next moment.

[0093] Step 10: Store the experience tuple <s(t), a(t), r(t), s(t + 1)> into the replay buffer pool.

[0094] Step 11: Judge whether the replay buffer is full. If so, take out mini - batch samples from the buffer to train the network; otherwise, judge whether t is equal to T max , if so, jump back to execute Step 3; otherwise, execute Step 12.

[0095] Step 12: Calculate the gradient of the current value network corresponding to the agent, and update the current policy network through gradient ascent.

[0096] Step 13: Calculate the value of the action at time t of the agent using the current value network, and at the same time calculate the value of the action at time t + 1 using the target network, and then update the current value network through gradient descent.

[0097] Step 14: Use soft update to update the policy network and value network in the target network.

[0098] Step 15: Determine whether convergence has occurred. If so, obtain the final optimal solution; otherwise, jump back to Step 3.

[0099] In a preferred embodiment of the present invention, the parameters initialized in the second step include:

[0100] Initialize the weight parameters θ j and w j of each actor and critic network, and initialize the parameters and of each target actor and critic network through and Initialize the learning rates α and β corresponding to the critic and actor networks, the discount factor γ, the number of episodes EP, and the maximum number of training steps T for each episode. max Initialize the replay buffer size D, the size M of the mini-batch, and the random process Ψ for action exploration. Initialize the network layout parameters, such as the number N of IoT devices, the number K of satellites, and task parameters.

[0101] In a preferred embodiment of the present invention, the state of the agent and the global state in the third step are respectively represented as:

[0102]

[0103] where represents the task requests of the IoT devices covered by satellite k at the start of time step t; represents the number of occurrences of task q on satellite k at the end of time slot t - 1, i.e., at the start of time slot t; is the remaining storage space on the satellite at the start of time slot t; is the path loss between the IoT device and the satellite; is the reallocation of satellite computing power to the IoT device.

[0104] In a preferred embodiment of the present invention, in the fifth step, the obtained action a k (t) is represented as follows:

[0105]

[0106] where is the task placement decision of satellite k; represents access control; represents the instance selection strategy; Represents the bandwidth allocation between satellite k and IoT device n.

[0107] In a preferred embodiment of the present invention, in the ninth step, the calculation formula of the reward reward is as follows:

[0108]

[0109] In a preferred embodiment of the present invention, in the twelfth step, the current policy network is updated by gradient ascent as follows:

[0110]

[0111] In a preferred embodiment of the present invention, in the thirteenth step, the update process of the current value network by gradient descent is as follows:

[0112]

[0113] Where and represent the TD target and TD error respectively.

[0114] In a preferred embodiment of the present invention, in the fourteenth step, the soft update formula of the target network is as follows:

[0115]

[0116] Where θ j represents the parameters of the current policy network, represents the parameters of the target policy network, w j represents the parameters of the current value network, represents the parameters of the target value network,

[0117] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, such as provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.

[0118] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A satellite communication network optimization method based on MEC, characterized in that: The following steps are involved: Step 1: The current network and target network of each agent together form the model of the training network, where the current network includes the current policy network actor and the current value network critic, and the target network includes the target policy network actor and the target value network critic; establish all these network models; Step 2: Initialize relevant parameters in the network; Step 3: Initialize the state of each agent k (t), and at the same time obtain the global initial state s(t); Step 4: Set time t=t0; Step 5: Agent k Through the strategic network Get agent k Action k (t); Step 6: Refactor action a k (t); Step 7: Determine whether k is equal to K. If not, return to step 5, otherwise execute step 8; Step 8: Get the global action a(t); Step 9: Execute action a(t) to get reward and the next state s(t+1); Step 10: Store the experience tuple <s(t), a(t), r(t), s(t+1)> into the buffer area; Step 11: Determine whether the buffer is full. If so, take out a mini-batch of samples from the buffer to train the network. If not, determine whether t is equal to T. max If yes, jump back to step 3, otherwise go to step 12; Step 12: Calculate the gradient of the current value network corresponding to the agent, and update the current policy network through gradient ascent; Step 13: Use the current value network to calculate the value of the agent's action at time t, and use the target network to calculate the value of the action at time t+1, and then update the current value network through gradient descent; Step 14: Use soft update to update the policy network and value network in the target network; Step 15: Determine whether it converges. If so, get the final optimal solution. Otherwise, jump back to step 3.

2. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: The parameters initialized in the second step include: Initialize the weight parameters θ of each actor and critic network j and w j ,pass and Initialize the parameters of each target actor and critic network and Initialize the learning rates α and β corresponding to the critic and actor networks, the discount factor γ, the number of episodes EP, and the maximum number of training steps T for each episode max ; Initialize the replay buffer size D, the mini-batch size M, and the random process Ψ for action exploration; Initialize network layout parameters, such as the number of IoT devices N, the number of satellites K, mission parameters, etc.

3. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: The agent state and global state in the third step are respectively expressed as: in represents the mission request of IoT devices covered by satellite k at the beginning of time step t; represents the number of occurrences of task q on satellite k at the end of time slot t-1, i.e., at the beginning of time slot t; is the remaining storage space on the satellite at the beginning of time slot t; is the path loss between the IoT device and the satellite; The goal is to redistribute satellite computing power to IoT devices.

4. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: In the fifth step, the action a is obtained k (t) is expressed as follows: in, is the mission placement decision of satellite k; Indicates access control; Indicates the instance selection strategy; represents the bandwidth allocation between satellite k and IoT device n.

5. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: In the ninth step, the calculation formula of reward is as follows:

6. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: In the twelfth step, the current policy network is updated by gradient ascent as follows:

7. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: In the thirteenth step, the updating process of the current value network by gradient descent is as follows: in and denote TD target and TD error respectively.

8. The method for optimizing a satellite communication network based on MEC according to claim 1, characterized in that: In the fourteenth step, the soft update formula of the target network is as follows: where θ j Represents the parameters of the current policy network, represents the parameters of the target policy network, w j Represents the parameters of the current value network, represents the parameters of the target value network, 9. A joint optimization system for task placement and removal, access control, service instance selection, and bandwidth allocation based on deep reinforcement learning, characterized in that: include: The training network module is used to establish a current network and a target network for each agent, where the current network includes the current policy network (actor) and the current value network (critic), and the target network includes the target policy network (actor) and the target value network (critic); Parameter initialization module, used to initialize the parameters in the network and set the initial state of the agent, and obtain the global initial state; The action generation and reconstruction module obtains the action of the intelligent agent through the policy network and reconstructs the action; The reward calculation and state update module calculates the reward value after executing the global action, updates the state to the next moment, and stores the experience tuple in the replay buffer pool; The gradient update module uses the current value network and the target network to calculate the value of the current and next actions respectively, updates the policy network through gradient ascent, and updates the value network through gradient descent; The target network update module updates the target policy network and the target value network through a soft update method; The convergence judgment module determines whether the convergence conditions are met. If so, the final optimization solution is output. If not, it returns to the execution action generation and reconstruction module to continue the optimization iteration.

10. The system for joint optimization of task placement and removal, access control, service instance selection and bandwidth allocation based on deep reinforcement learning according to claim 9, characterized in that: The system further comprises: The replay cache module is used to store the agent's experience tuples and extract samples from the cache pool using mini-batch sampling to optimize network parameters. The global state management module is used to manage the local state and action of each agent and generate the global state and global action to achieve collaborative optimization of multiple agents; The time step control module is used to dynamically adjust the time step t in the optimization process to ensure that action execution, reward calculation and network parameter update are completed within each time step.

Citation Information

Patent Citations

  • MEC-based air-space-ground integrated network task segmentation and resource allocation method

    CN118075772A

  • Joint optimization method based on multi-agent depth deterministic strategy gradient

    CN118741605A

  • Edge computing enabled air-to-ground network service deployment and resource allocation method and system

    CN119255292A