Communication resource configuration method for assisting mobile edge joint optimization

By using non-orthogonal multiple access technology to assist mobile edge computing in 5G IoT networks, combining the computing power of static and mobile base stations, channel selection and resource allocation are optimized, and resource allocation is solved, and the resource allocation problem of computing-intensive delay-sensitive applications in heterogeneous multi-user scenarios is improved, achieving the improvement of spectrum utilization and network efficiency.

CN120343632APending Publication Date: 2025-07-18SHENYANG INSTITUTE OF CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510274004.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In 5G IoT networks, the prior art is difficult to effectively solve the application needs of computing-intensive but latency-sensitive, especially in heterogeneous multi-user scenarios. The allocation of non-orthogonal multiple access resources leads to insufficient utilization of frequency resources and cannot meet the service quality requirements of different users.

Method used

Non-orthogonal multiple access technology assists mobile edge computing, by constructing multi-objective optimization problems, using Markov decision-making process and deep reinforcement learning algorithm, combining the computing capabilities of static and mobile base stations, optimizing channel selection, offloading strategies and computing resource allocation, and proposing the MADDPG-CCRC algorithm for joint optimization.

Benefits of technology

It improves spectrum utilization, reduces network latency and energy consumption, meets users' service quality requirements, and optimizes network resource utilization and transmission rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343632A_ABST
    Figure CN120343632A_ABST
Patent Text Reader

Abstract

The invention discloses a communication resource configuration method for assisting mobile edge joint optimization, and relates to a communication resource configuration method, which comprises the following steps of: S1, building a system model of heterogeneous multi-edge multi-user non-orthogonal multiple access assisted mobile edge calculation; s2, establishing a multi-objective optimization problem by taking a weighted sum of minimum delay, energy consumption and service quality as an objective while complying with a service quality requirement under the constraint of scheduling and resource allocation; and S3, due to the non-convexity of the problem, a multi-objective optimization problem is converted into three sub-problems to be solved, and an MADDPG-CCRC algorithm is proposed to perform joint optimization on communication and computing resource configuration. According to the method, the non-orthogonal multiple access technology is applied to resource allocation of edge computing, the requirement of a large number of devices on computing-intensive and delay-sensitive services of the Internet is met, and a simulation result shows that the framework method improves spectrum efficiency and reduces energy consumption and delay of unloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a communication resource allocation method, specifically a communication resource allocation method for assisting mobile edge joint optimization. Background Art

[0002] With the rapid development of communication technologies, more and more people, data, and things are intelligently connected to the Internet, moving from the Internet of Things era to the era of everything connected. This transformation has spawned various computationally intensive but latency-sensitive applications. However, due to hardware limitations, the user side often faces limitations in computing resources and cannot meet task requirements. Integrating mobile edge computing in 5G Internet of Things networks is to alleviate latency issues and optimize network bandwidth utilization, which helps to process data closer to the data source, thereby reducing the distance and time required to transmit data to cloud-based servers, thus reducing latency. In addition, in the field of intelligent transportation, mobile edge computing plays a crucial role in optimizing traffic flow, enhancing security protocols, and alleviating congestion problems, thereby improving the overall operational efficiency of urban transportation systems.

[0003] Generally, the computing tasks generated in the same scenario are usually heterogeneous, resulting in different quality of service requirements. For example, VR / AR videos in a factory environment or entertainment applications in a vehicle require high data rates and broadband communication; on the contrary, tasks such as transmitting factory control signals or performing vehicle avoidance tasks require deterministic communication with ultra-low latency. With the continuous progress of technology, the proliferation of access devices in industrial and vehicle networks has highlighted the need to address the challenges brought by large-scale communication. Generally, these solutions involve multiple terminal devices and multiple edge servers. In order to successfully offload computing tasks to edge servers, network access must first be successful. Therefore, the paradigm of multi-access mobile edge computing is considered a promising approach.

[0004] In the context of 5G systems, fixed orthogonal resource allocation in non-orthogonal multiple access may lead to insufficient utilization of frequency resources, especially in scenarios with different user demands and traffic conditions. This limitation has stimulated the implementation of non-orthogonal multiple access technology, which accommodates multiple user transmissions within a single time-frequency resource block by using user-specific signatures, thus promoting more efficient sharing of frequency resources. Non-orthogonal multiple access improves spectrum utilization, enhances system capacity, and increases overall efficiency in high-density and dynamic network environments to meet the changing communication requirements of next-generation communication systems. Combining non-orthogonal multiple access with mobile edge computing technology has several key advantages. First, it improves the utilization of network resources and spectrum efficiency, thereby increasing the transmission rate and meeting the quality-of-service requirements of users. Second, by quickly scheduling and allocating resources for user requests, it reduces network latency and enhances the performance of real-time applications. Third, by directly deploying user requests and services on edge computing nodes, it reduces the data transmission distance and energy consumption, ultimately improving energy efficiency. Summary of the Invention

[0005] The object of the present invention is to propose a communication resource configuration method for assisted mobile edge joint optimization. This method is a resource allocation method in mobile edge computing assisted by non-orthogonal multiple access technology, which is decomposed into a series of tractable sub-problems. For the sub-problems of communication resource allocation and offloading strategies, they are transformed into Markov decision processes, and deep reinforcement learning algorithms are used to find optimal solutions. A multi-agent Actor-Critic (AC) network is implemented, which interacts with the environment to obtain rewards, thereby adjusting actions to continuously optimize the solutions.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A communication resource configuration method for assisted mobile edge joint optimization includes the following steps:

[0008] S1: Build a system model for non-orthogonal multiple access-assisted mobile edge computing, which includes a time-slot communication model and a computing model. The computing model includes a local computing model and an edge computing model; the framework of the present invention considers a system composed of multiple edge nodes and multiple users. Among them, the base stations are divided into two types: static base stations (SBS) and mobile base stations (MBS). The computing capabilities of these two types of base stations are different, and each base station is equipped with a successive interference cancellation (SIC) receiver and a multi-access edge computing (MEC) server. By dividing the entire system's spectrum into several orthogonal channels for wireless transmission connection, the resource allocation for computing tasks is completed. The model specifically includes channel parameters, base station parameters, the number of users, the number of orthogonal channels, work task data packets, local and edge offloading parameters, etc.;

[0009] S2: In the system model constructed according to S1, we developed a multi - edge and multi - terminal non - orthogonal multiple access (NOMA) assisted mobile edge computing (MEC) system framework. In the present invention, we considered the time delay and energy consumption in task completion, and defined the task cost as a weighted combination of time delay, energy consumption, and quality of service (QoS). While complying with the QoS requirements under scheduling and resource allocation constraints, the goal was to minimize the weighted sum of delay, energy consumption, and QoS. Specifically, it was optimized by searching for four parameters: a feasible channel selection strategy, an edge offloading strategy, a transmission power set, and an edge computing resource allocation strategy. The multi - objective optimization problem was established through a Markov decision process to minimize the system's cost function;

[0010] S3: Due to the non - convexity of the problem, the multi - objective optimization problem was transformed into three sub - problems for solution, and a MADDPG - CCRC algorithm was proposed to jointly optimize the communication and computing resource allocation, and finally the best strategy was found. The interaction between the agent and the environment was carried out by establishing the state space, action space, and reward function in the Markov decision process that adapted to the built system environment. Deep reinforcement learning technology was used to determine the joint channel and offloading strategy to solve the first sub - problem; the power allocation was optimized according to the characteristics of non - orthogonal multiple access to solve the second sub - problem; and a convex optimization algorithm was applied to obtain the optimal solution of the computing resource allocation to solve the third sub - problem. The main innovation points of the method of the present invention are as follows: This system combines static and mobile base stations, which have different computing capabilities according to their types; at the same time, the tasks generated by users show heterogeneity and time dynamics. Within this framework, we formulated an optimization problem aimed at minimizing the system cost, involving the joint optimization of channel selection, offloading decision, transmission power, and computing resource allocation.

[0011] Furthermore, the hierarchical architecture 5G network slice system model described in S1 is built through the following steps:

[0012] S11: We consider a non - orthogonal multiple access assisted mobile edge computing system with multiple edges and multiple terminals. N user equipments (UEs) are randomly distributed in a circular area with a radius of R and communicate with M base stations, where M1 is the static base station and M2 is the mobile base station, M = M1+M2. Each base station is equipped with a receiver and an edge server. The computing capabilities of the edge servers associated with the static base station and the mobile base station are represented as F1 and F2 respectively. The entire system's spectrum is divided into K orthogonal channels. We use UE n , ED m and C k to represent the nth user equipment, the mth edge (i.e., base station), and the kth channel respectively;

[0013] S12: In each time slot, each user equipment generates a computing task. We represent the tasks of several user equipments as Γ n= d n , w n , T max,n 1 ≤ n ≤ N, where d n represents the size of the data packet, w n represents the number of CPU cycles T required for the calculation max,n is the deadline of the task, limited within the duration of one time slot;

[0014] S13: Due to limited computing resources, the user equipment can choose to offload its task to the edge server for help. We assume that within one time slot, all base stations and user equipment remain stationary, but in different time slots, the mobile base station changes its position along with the user equipment;

[0015] S14: If the user equipment chooses to offload its task, it first selects a channel to transmit its data packet. In the invention, we apply non - orthogonal multiple access technology, where each base station is equipped with an L - SIC receiver, which means that each time - frequency resource block can transmit signals for L users simultaneously. Since there are K channels in the system, at most L * K data packets can be transmitted within one time slot. At the base station, the received signals of each channel are decoded in descending order of received power. That is to say, by applying the receiver, the base station first decodes the user equipment with the strongest received power, and regards the other signals from users in the same channel as interference, so as to distinguish the offloading priority of user equipment. We define the channel selection strategy as where u n,k is expressed as:

[0016]

[0017] When k = 0, the user equipment chooses not to transmit the data packet, that is, local calculation;

[0018] The received power of the user equipment signal at ED m through C k is:

[0019]

[0020] where, is the transmission power of the user, h n,m,k and α represent the small - scale Rayleigh fading coefficient and the large - scale path loss, follows an exponential distribution with unit parameter, r n,m is the Euclidean distance between the user and the base station.

[0021] According to Shannon's theorem, the signal - to - noise ratio of the user from ED m to C k is expressed as:

[0022]

[0023] where is the indicator function, and σ 2 represents the noise power. The condition for the base station to successfully decode is that the SINR is greater than the threshold θ. After successful decoding, the strongest signal is subtracted from the received signal, and then the process is repeated; if the decoding fails, the decoding process will terminate immediately, and the remaining undecoded signals are regarded as failures;

[0024] S15: Local computing model: We define the edge offloading strategy as:

[0025]

[0026] where when m = 0, the user performs local computing, denotes the local computing ability, and the computing delay and energy consumption are respectively:

[0027]

[0028] where κ is the energy consumption constant, usually taking the value 10 -28 ; We consider the time delay and energy consumption at the completion of the task, and define the task cost as a weighted combination of these factors, expressed as:

[0029]

[0030] where β is the weight of energy consumption;

[0031] S16: Edge computing model: The users who complete the computing tasks through offloading need to be divided into two stages: the stage of transmitting tasks to the edge through the channel and the stage of completing task computing at the edge. Therefore, the offloading delay includes the delay involved in transmitting data packets and the delay generated during edge computing; the transmission delay depends on the size of the data packet denoted by d n and the transmission rate is denoted by R n,m,k , and the transmission rate of the user equipment from the channel to the base station edge is obtained as:

[0032] R n,m,k =v n,m Blog21 + SINR n,m,k ;

[0033] where B is the channel bandwidth. Then, for v n,m = 1, the transmission delay Γ n of the task is expressed as:

[0034]

[0035] For v n,m= 1, we represent the computing resources allocated by the base station to the user as Therefore, for task Γ n the offloading computing delay can be expressed as:

[0036]

[0037] Regarding the energy consumption, since the user side is more restricted by battery limitations, this paper only considers the energy required for data packet transmission at the user side. Then, in the edge computing model, the energy consumption required to complete the task can be expressed as:

[0038]

[0039] Similar to the local computing model, the task cost function of edge computing is defined as:

[0040]

[0041] Furthermore, the optimization problem of establishing the corresponding hierarchical structure according to the system model described in S2 includes the following steps:

[0042] S21: The present invention takes into account the time delay and energy consumption in task completion. We define the task cost as a weighted combination of these factors, which can be expressed as:

[0043]

[0044] where x n,m,k = u n,k ∩ v n,m , considering the different levels of time sensitivity between different tasks, we enhance Z n by incorporating whether the task can be completed within the deadline:

[0045] Z n + η n δT n - T max,n ;

[0046] where T n represents the delay required to complete the task and can be expressed as:

[0047]

[0048] where η n represents the importance of the task, pays special attention to the sensitivity to delay and is related to T max,n , and δ(x) is defined as: The total cost function of the system model is the sum of the costs of all tasks and can be expressed as:

[0049]

[0050] S22: and respectively represent the set of user transmission capabilities and the base station computing resource allocation strategy. Our goal is to minimize the cost function of the system by searching for a feasible channel selection strategy edge offloading strategy transmission power set and edge computing resource allocation strategy Therefore, the optimization problem is formulated as follows:

[0051]

[0052] The specific constraint conditions of the optimization problem described in S22 are further explained as follows:

[0053] S23: The constraint conditions of the above S22 optimization problem are respectively represented as follows:

[0054] S231: F m represents the computing resources owned by the base station, and P max,n represents the maximum transmission power of the user, and impose constraints on the users, allowing them to select at most one channel to send their data packets;

[0055] S232: and restrict the users from offloading their tasks to at most one edge node;

[0056] S233: x n,m,k = u n,k ∩ v n,m represents the joint scheme of channel selection and edge offloading; only when the two are combined will the user transmit their tasks to the base station through the channel;

[0057] S234: and represent that the computing resources allocated to tasks by each edge server cannot exceed their own computing resources;

[0058] S235: Ensure that the transmission power of the user does not exceed its maximum allowable value.

[0059] Furthermore, the optimization problem described in S3 is transformed into a Markov decision process for solution, which is divided into three sub-problems for separate solution, including the following steps:

[0060] S31: Obviously, the optimization problem is a mixed - integer and non - convex problem, which is usually classified as a difficult problem, making it challenging to solve using traditional optimization methods. To address this issue, our approach involves decomposing the problem into multiple sub - problems and optimizing the variables step by step. More specifically, we initially use deep reinforcement learning techniques to determine the joint channel and offloading strategy Next, we optimize the power allocation according to the characteristics of non - orthogonal multiple access Finally, we apply a convex optimization algorithm to obtain the optimal solution for the computing resource allocation ;

[0061] S32: In consecutive time slots, the optimization problem can be regarded as a partially observable multi - agent Markov decision process. Each user is regarded as an agent. Partial observability means that each agent only knows its own current information, including but not limited to channel state, packet size, distance to each base station, task deadline, etc., and does not know the information of other agents. To solve the above problem, it is divided into two sub - problems. First, each agent makes an independent decision to solve the joint offloading scheme of the channel and the edge node ; then, the agent calculates the corresponding and to further reduce the cost function. Next, we use a multi - agent Markov decision process to reconstruct the problem. We use to represent the state, action, reward of the Markov decision process and their respective sets

[0062] S33: Power allocation sub - problem: After determining , the users specify their computing models. In the case of local computing, the cost is completely determined by the computing resources required for the task and the local computing power, resulting in a fixed value. On the contrary, if offloading computing is selected, the transmission power needs to be further optimized because the transmission power affects the transmission rate and thus the transmission time. Our goal is to optimize the transmission power to minimize the transmission time. Therefore, the sub - problem is formulated as follows:

[0063]

[0064] where for the edge time, since the transmission power only affects the transmission delay and does not affect the offloading computing delay, we omit the term

[0065] S34: Computing resource allocation sub - problem: Allocate computing resources at the edge according to the received computing tasks. Similar to the transmission power problem, the configuration of computing resources only affects the offloading computing delay. Our goal is to optimize the offloading computing delay. The sub - problem is formulated as follows:

[0066]

[0067] Furthermore, listing the space described in S32 respectively includes the following steps:

[0068] S321: List the state space in the Markov decision process of the optimization problem. At time slot t, we use s nt to represent the state of agent N, including the action in the previous time slot, the current state of task N, local and edge computing resources, and channel state information, which is expressed as:

[0069]

[0070] where o nt = a n(t-1) represents the action output in the previous time slot, and represents the matrix of channel gains and the distance to the edge at the current time slot, represents the computing resources available at the edge, and the state set of all tasks is represented as s t = s nt , 1 ≤ n ≤ N;

[0071] S322: List the action space in the Markov decision process of the optimization problem. Similar to the state, we represent the action of agent n at time slot t as a nt which includes the index of the accessed channel and the offloaded edge server, and is expressed as:

[0072] a nt = C nt , ED nt ;

[0073] where C nt ∈ {0, 1, 2,..., K} and ED nt ∈ {0, 1, 2,..., M}. When C nt or ED nt is 0, it means that the task will perform local computing. According to the action of the agent, we obtain the corresponding and and derive the required channel selection strategy and offloading scheme All action sets can be represented as: a t = a nt , 1 ≤ n ≤ N;

[0074] S323: We divide the reward into two parts: one part is the offloading reward, and the other part is the quality-of-service reward. The offloading reward is defined as the cost reduction achieved by performing offloading computing compared with local computing, and is expressed as:

[0075]

[0076] The reward for service quality is the reward obtained for completing tasks within the deadline, defined as:

[0077]

[0078] Therefore, the reward of agent n and the total reward of the system at time slot t are:

[0079]

[0080] S324: Therefore, regarding the determination of the optimal channel and offloading joint scheme and can be rephrased as:

[0081]

[0082]

[0083] x n,m,k = u n,k ∩ v n,m ;

[0084] where γ ∈ 0,1 represents the discount factor reflecting the influence of future time slots on the current time slot. The higher γ is, the greater the influence of future rewards on the current reward. Our goal is to maximize the long-term reward.

[0085] Furthermore, solving sub-problem two described in S33 includes the following steps:

[0086] S331: The condition for successful decoding is that the SINR of the received signal is greater than or equal to the preset threshold, i.e., SINR n,m,k ≥ θ, subject to the constraint of the limited battery capacity of the user equipment to extend the battery life and reduce the transmission power while maintaining successful decoding; therefore, the minimum received power for successful decoding in the first power level is:

[0087]

[0088] S332: The channel gain of the nth agent arriving at the edge m through channel k is defined as Agents with higher channel gains are assigned higher power levels because this assignment aims to minimize the total transmission power; subsequently, the assigned level and will be passed back to each agent; when the nth agent receives the signal returned by the environment, the corresponding transmit power is calculated as:

[0089]

[0090] where l ntdenotes the power level allocated to the n-th agent at time slot t; if it is above the allowed maximum power, the n-th agent performs local computation. In this way, we solve this sub-problem.

[0091] Furthermore, the solution to sub-problem three described in S34 includes the following steps:

[0092] S341: After receiving the computing task, the edge base station will optimize the allocation of computing resources to further reduce the system cost; since the computations at the edge are independent of each other, the optimization objective is adjusted to minimize the computing delay of each edge. Therefore, for the specified edge base station m, we define Its problem can be reformulated as:

[0093]

[0094] After determining v n,m the problem effectively becomes a convex optimization problem, which can be solved using Lagrange multipliers and KKT conditions to find the global optimal solution. Its Lagrangian is established as follows:

[0095]

[0096] where ρ and τ are non-negative Lagrange multipliers associated with the constraints;

[0097] S342: Let denote the optimal solution, then the KKT conditions can be derived as:

[0098]

[0099] Obviously, for this sub-problem, we only need to focus on the case of, so we can derive:

[0100]

[0101] Since ρ * > 0, we obtain The computing resources we allocate to the n-th agent at time slot t are derived as:

[0102]

[0103] Finally, we solve this sub-problem in this way.

[0104] Furthermore, for a communication resource configuration method to assist mobile edge joint optimization, to solve the three sub-problems described in S3, we propose an algorithm called MADDPG-CCRC, which includes the following steps:

[0105] S35: The invention is named as a MADDPG-CCRC algorithm, and its steps are as follows:

[0106] S351: The algorithm starts to execute, and relevant system parameters are input and initialized, including user equipment, number of base stations, number of orthogonal channels, signal-to-noise ratio, channel parameters, etc.;

[0107] S352: Initialize the policy set of the four parameters to be optimized, and complete the setting of algorithm hyperparameters, etc.;

[0108] S353: Reset the environment variables, enter the iterative main loop, and put the state space, action space, reward function, and the next moment state space minibatch into the replay pool;

[0109] S354: Observe the current state to generate a random policy; and judge whether it is less than the greedy parameter according to the condition. If the condition holds, randomly select an action and generate an action set, and then return to execute S353; otherwise, obtain the average maximum value of the action set according to the policy, output the channel selection and offloading policy sets, and execute S359;

[0110] S355: After solving sub-problem one, set both its user base station and channel selection to 1, calculate relevant parameters to determine the priority level of the agents that select the same base station, and send the result back to each agent to facilitate the solution of subsequent problems;

[0111] S356: Allocate computing resources for each user. If offloading is performed, make a judgment. If the transmission power is greater than the maximum limit power, set the relevant parameters to zero and return to execute S353; otherwise, use the transmission power set to solve sub-problem two and execute S359;

[0112] S357: Calculate the current reward value, observe and record the next moment state, and put the state-reward combination into the experience replay pool;

[0113] S358: Calculate the Q value and the loss function, update the internal network parameters of the algorithm; and obtain the optimal resource allocation scheme according to the convex optimization algorithm and the KKT condition to solve sub-problem three;

[0114] S359: Output the channel selection policy, offloading policy, transmission power policy, and computing resource allocation policy. After finding the optimal approximate policy value, the algorithm execution is completed.

[0115] The invention has the following advantages and effects:

[0116] 1. We have developed a system framework for scenarios involving hybrid multi-base stations and heterogeneous tasks. In this case, hybrid base stations refer to the combination of static and mobile base stations, while heterogeneous tasks involve variations in task size, computing resource requirements, and deadlines. By combining non-orthogonal multiple access and mobile edge computing, and considering system latency, energy consumption, and quality-of-service constraints, we formulate a joint communication and computing optimization problem. To solve the above problem, we first decompose it into a series of tractable sub-problems. For the sub-problems of communication resource allocation and offloading strategies, we transform them into Markov decision processes and use deep reinforcement learning algorithms to find optimal solutions. Specifically, we implement a multi-agent Actor-Critic (AC) network that interacts with the environment to obtain rewards, thereby adjusting actions to continuously optimize the solution.

[0117] 2. When solving the power allocation sub-problem, we combine the characteristics of non-orthogonal multiple access to minimize the transmission power while ensuring successful decoding, thereby reducing the energy consumption at the user end and extending the battery life. For the computing resource allocation sub-problem, the optimal solution is derived using Lagrange multipliers and KKT conditions. Finally, by integrating these sub-problems, we propose the MADDPG-CCRC algorithm framework, which proves the effectiveness and superiority of the algorithm proposed in the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0118] Figure 1 is the flowchart of the implementation process of the method of the present invention;

[0119] Figure 2 is the system model diagram built by the method of the present invention;

[0120] Figure 3 is the network framework diagram of the algorithm proposed in the present invention;

[0121] Figure 4 is the algorithm execution flowchart. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0122] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0123] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0124] As Figure 1 and Figure 2As shown in the figure, a communication resource allocation method for assisting mobile edge joint optimization according to an embodiment of the present invention includes the following steps:

[0125] S1: Build a system model for non-orthogonal multiple access-assisted mobile edge computing, including a time-slot communication model and a computing model. The computing model includes a local computing model and an edge computing model; the framework of the present invention considers multiple edge nodes and multiple users. Among them, the base stations are divided into two types: static base stations (SBS) and mobile base stations (MBS). The computing capabilities of these two types of base stations are different, and each base station is equipped with a receiver (SIC) and an edge server (MEC). By dividing the spectrum of the entire system into several orthogonal channels for wireless transmission connection, the computing task resource allocation is completed. The model specifically includes channel parameters, base station parameters, the number of users, the number of orthogonal channels, work task data packets, local and edge offloading parameters, etc.;

[0126] Specifically, in the S1, the system model of non-orthogonal multiple access technology-assisted mobile edge computing is as follows:

[0127] S11: We consider a non-orthogonal multiple access-assisted mobile edge computing system with multiple edges and multiple terminals. N user equipment (UE) are randomly distributed in a circular area with a radius of R and communicate with M base stations, where M1 is a static base station and M2 is a mobile base station, and M = M1 + M2. Each base station is equipped with a receiver and an edge server. The computing capabilities of the edge servers associated with the static base station and the mobile base station are represented as F1 and F2 respectively. The spectrum of the entire system is divided into K orthogonal channels. We use UE n , ED m and C k to represent the nth user equipment, the mth edge node (i.e., the base station), and the kth channel respectively;

[0128] S12: In each time slot, each user equipment generates a computing task. We represent the tasks of several user equipment as Γ n = d n , w n , T max,n 1 ≤ n ≤ N, where d n represents the size of the data packet, w n represents the number of CPU cycles required for computing, and T max,n is the deadline of the task, which is limited within the duration of one time slot;

[0129] S13: Due to limited computing resources, the user equipment can choose to offload its tasks to the edge server for help. We assume that within one time slot, all base stations and user equipment remain stationary, but in different time slots, the mobile base station will change its position along with the user equipment;

[0130] S14: If the user equipment chooses to offload its task, it first selects a channel to transmit its data packets. In the invention, we apply the non-orthogonal multiple access technology, where each base station is equipped with an L-SIC receiver, which means that each time-frequency resource block can transmit signals for L users simultaneously. Since there are K channels in the system, at most L*K data packets can be transmitted in one time slot. At the base station, the received signals of each channel are decoded in descending order of received power. That is to say, by applying the receiver, the base station first decodes the user equipment with the strongest received power, and regards the other signals from the users on the same channel as interference, so as to distinguish the offloading priorities of user equipment. We define the channel selection strategy as where u n,k is expressed as:

[0131]

[0132] When k = 0, the user equipment chooses not to transmit data packets, that is, local computing;

[0133] The received power of the user equipment signal at ED m through C k is:

[0134]

[0135] where is the transmission power of the user, h n,m,k and α represent the small-scale Rayleigh fading coefficient and the large-scale path loss. |h n,m,k | 2 follows an exponential distribution with unit parameter, and r n,m is the Euclidean distance between the user and the base station.

[0136] According to the Shannon theorem, the signal-to-noise ratio of the user from ED m to C k is expressed as:

[0137]

[0138] where is the indicator function, and σ 2 represents the noise power. The condition for the base station to successfully decode is that the SINR is greater than the threshold θ. After successful decoding, the strongest signal is subtracted from the received signal, and then the process is repeated; if the decoding fails, the decoding process will terminate immediately, and the remaining undecoded signals are regarded as failures;

[0139] S15: Local computing model: We define the edge offloading strategy as:

[0140]

[0141] When m = 0, the user performs local computation, which is expressed as local computing power. Then the computing delay and energy consumption are respectively:

[0142]

[0143]

[0144] where κ is the energy consumption constant, usually taking the value of 10 -28 ; We consider the time delay and energy consumption at the completion of the task and define the task cost as a weighted combination of these factors, expressed as:

[0145]

[0146] where β is the weight of energy consumption;

[0147] S16: Edge computing model: The users who complete the computing tasks through offloading need to be divided into two stages: the stage of transmitting tasks to the edge through the channel and the stage of completing task computing at the edge. Therefore, the offloading delay includes the delay involved in transmitting data packets and the delay generated during edge computing; The transmission delay depends on the size of the data packet denoted by d n and the transmission rate is denoted as R n,m,k , and the transmission rate of the user device from the channel to the base station edge is obtained as:

[0148] R n,m,k = v n,m Blog21 + SINR n,m,k ;

[0149] where B is the channel bandwidth. Then for v n,m = 1, the transmission delay Γ n of the task is expressed as:

[0150]

[0151] For v n,m = 1, we denote the computing resources allocated by the base station to the user as Therefore, the offloading computing delay of task Γ n can be expressed as:

[0152]

[0153] Regarding energy consumption, since the user side is more restricted by battery limitations, this paper only considers the energy required for data packet transmission at the user side. Then, in the edge computing model, the energy consumption required to complete the task can be expressed as:

[0154]

[0155] Similar to the local computing model, the task cost function of edge computing is defined as:

[0156]

[0157] S2: In the system model constructed according to S1, we developed a multi - edge multi - terminal non - orthogonal multiple access assisted mobile edge computing system framework. In the present invention, we considered the time delay and energy consumption in task completion, defined the task cost as a weighted combination of time delay, energy consumption, and quality of service, and aimed to minimize the weighted sum of delay, energy consumption, and quality of service while complying with the quality - of - service requirements under scheduling and resource allocation constraints. Specifically, it was optimized by searching for four parameters: a feasible channel selection strategy, an edge offloading strategy, a transmission power set, and an edge computing resource allocation strategy, and a multi - objective optimization problem was established through a Markov decision process to minimize the cost function of the system;

[0158] Specifically, in S2, an optimization problem was established based on the system model, and the specific optimization problem is as follows:

[0159] S21: The present invention considered the time delay and energy consumption in task completion. We defined the task cost as a weighted combination of these factors, which can be expressed as:

[0160]

[0161] where x n,m,k = u n,k ∩v n,m , considering the different levels of time sensitivity between different tasks, we enhanced Z n by incorporating whether the task can be completed within the deadline:

[0162] Z n +η n δT n -T max,n ;

[0163] where T n represents the delay required to complete the task and can be expressed as:

[0164]

[0165] where η n represents the importance of the task, pays particular attention to the sensitivity to delay and is related to T max,n , and δ(x) is defined as: The total cost function of the system model is the sum of the costs of all tasks and can be expressed as:

[0166]

[0167] S22: and represent the set of user transmission capabilities and the base station computing resource allocation strategy respectively. Our goal is to minimize the cost function of the system by searching for a feasible channel selection strategy edge offloading strategy transmission power set and edge computing resource allocation strategy Thus, the optimization problem is formulated as follows:

[0168]

[0169] Specifically, in the aforementioned S22, an optimization problem is established according to the system model, and its specific constraint conditions are as follows:

[0170] S23: The constraint conditions of the above S22 optimization problem are represented as follows:

[0171] S231: F m represents the computing resources owned by the base station, and P max,n represents the maximum transmission power of the user, and impose constraints on the users, allowing them to select at most one channel to send their data packets;

[0172] S232: and restrict the users from offloading their tasks to at most one edge node;

[0173] S233: x n,m,k = u n,k ∩ v n,m represents the joint scheme of channel selection and edge offloading; only when the two are combined will the user transmit their tasks to the base station through the channel;

[0174] S234: and indicate that the computing resources allocated to tasks by each edge server cannot exceed their own computing resources;

[0175] S235: Ensure that the transmission power of the user does not exceed its maximum allowable value.

[0176] S3: Due to the non-convexity of the problem, the multi-objective optimization problem is transformed into three sub-problems for solution, and a MADDPG-CCRC algorithm is proposed to jointly optimize the communication and computing resource allocation, and finally the optimal strategy is found. The interaction between the agent and the environment is carried out by establishing the state space, action space and reward function in the Markov decision process that adapt to the built system environment. Deep reinforcement learning technology is used to determine the joint channel and offloading strategy to solve sub-problem one; according to the characteristics of non-orthogonal multiple access, the power allocation is optimized to solve sub-problem two; the convex optimization algorithm is applied to obtain the optimal solution of the computing resource allocation to solve sub-problem three. The main innovation of the method of the present invention is reflected in that this system combines static and mobile base stations and has different computing capabilities according to their types; at the same time, the tasks generated by users show heterogeneity and time dynamics. Within this framework, we formulate an optimization problem aiming to minimize the system cost, which involves the joint optimization of channel selection, offloading decision, transmission power and computing resource allocation.

[0177] Specifically, in the said S3, according to the conversion of the problem into three sub-problems, sub-problem one is listed as follows:

[0178] S31: Obviously, the optimization problem is a mixed integer and non-convex problem, which is usually classified as a difficult problem, making it challenging to solve with traditional optimization methods. To solve this problem, our method includes decomposing the problem into multiple sub-problems and optimizing the variables step by step; more specifically, we initially use deep reinforcement learning technology to determine the joint channel and offloading strategy Next, we optimize the power allocation according to the characteristics of non-orthogonal multiple access Finally, we apply the convex optimization algorithm to obtain the computing resource allocation of the optimal solution;

[0179] S32: In consecutive time slots, the optimization problem can be regarded as a partially observable multi-agent Markov decision process. Each user is regarded as an agent. Partial observability means that each agent only knows its own current information, including but not limited to channel state, packet size, distance to each base station, task deadline, etc., and does not know the information of other agents. To solve the above problem, it is divided into two sub-problems. First, each agent makes an independent decision to solve the joint offloading scheme of the channel and the edge node ; then, the agent calculates the corresponding and to further reduce the cost function. Next, we use the multi-agent Markov decision process to reconstruct the problem. We use to represent the state, action, reward of the Markov decision process and their respective sets;

[0180] Specifically, in S32, to solve its sub-problem 1, the deep reinforcement learning method is used, and the steps are as follows:

[0181] S321: List the state space in the Markov decision process of the optimization problem. At time slot t, we use s nt to represent the state of agent N, including the action in the previous time slot, the current state of task N, local and edge computing resources, and channel state information, expressed as:

[0182]

[0183] where o nt = a n(t-1) represents the action output in the previous time slot, and represent the matrix of channel gains and the distance to the edge at the current time slot, represents the computing resources available at the edge, and the state set of all tasks is expressed as s t = s nt , 1 ≤ n ≤ N;

[0184] S322: List the action space in the Markov decision process of the optimization problem. Similar to the state, we represent the action of agent n at time slot t as a nt which includes the index of the accessed channel and the offloaded edge server, expressed as:

[0185] a nt = C nt , ED nt ;

[0186] where C nt ∈ {0, 1, 2,..., K} and ED nt ∈ {0, 1, 2,..., M}, when C nt or ED nt is 0, it means that the task will perform local computing; according to the action of the agent, we obtain the corresponding and to derive the required channel selection policy and offloading scheme All action sets can be expressed as: a t = a nt , 1 ≤ n ≤ N;

[0187] S323: We divide the reward into two parts: one part is the offloading reward, and the other part is the quality of service reward. The offloading reward is defined as the cost reduction achieved by performing offloading computing compared with local computing, expressed as:

[0188]

[0189] The reward for service quality is the reward obtained for completing tasks within the deadline, defined as:

[0190]

[0191] Therefore, the reward of agent n and the total reward of the system at time slot t are:

[0192]

[0193] S324: Therefore, regarding the determination of the optimal channel and offloading joint scheme and can be rephrased as:

[0194]

[0195] x n,m,k = u n,k ∩ v n,m ;

[0196] where γ ∈ 0, 1 represents the discount factor reflecting the influence of future time slots on the current time slot. The higher γ is, the greater the influence of future rewards on the current reward. Our goal is to maximize the long-term reward.

[0197] Specifically, in S3, sub-problem two includes the following steps:

[0198] S33: Power allocation sub-problem: After determining , users specify their computing models. In the case of local computing, the cost is completely determined by the computing resources required for the task and the local computing power, resulting in a fixed value. On the contrary, if offloading computing is selected, the transmission power needs to be further optimized because the transmission power affects the transmission rate, which in turn affects the transmission time. Our goal is to optimize the transmission power to minimize the transmission time. Therefore, this sub-problem is formulated as follows:

[0199]

[0200] s.t. x n,m,k = u n,k ∩ v n,m ;

[0201]

[0202] where for the edge time, since the transmission power only affects the transmission delay and does not affect the offloading computing delay, we omit the term

[0203] Specifically, in S33, to solve sub-problem two, it includes the following steps:

[0204] S331: The condition for successful decoding is that the SINR of the received signal is greater than or equal to a preset threshold, i.e., SINR n,m,k ≥θ. There is a constraint on the limited battery capacity of the user equipment to extend the battery life and reduce the transmission power while maintaining successful decoding; thus, the minimum received power for successful decoding in the first power level is:

[0205]

[0206] S332: The channel gain of the nth agent arriving at the edge m through channel k is defined as Agents with higher channel gains are assigned higher power levels because this assignment aims to minimize the total transmission power; subsequently, the assigned level, as well as will be sent back to each agent; when the nth agent receives the signal returned by the environment, the corresponding transmit power is calculated as:

[0207]

[0208] where l nt represents the power level assigned to the nth agent at time slot t; if exceeds the allowed maximum power, the nth agent performs local calculations. In this way, we solve this sub-problem.

[0209] Specifically, in the aforementioned S3, sub-problem three includes the following steps:

[0210] S34: Calculate the resource allocation sub-problem: Allocate computing resources at the edge according to the received computing tasks. Similar to the transmission power problem, the configuration of computing resources only affects the offloading computing delay. Our goal is to optimize the offloading computing delay. The sub-problem is formulated as follows:

[0211]

[0212] s.t. x n,m,k = u n,k ∩v n,m ;

[0213]

[0214] Specifically, in the aforementioned S34, to solve sub-problem three, it includes the following steps:

[0215] S341: After receiving the computing task, the edge base station will optimize the computing resource allocation to further reduce the system cost; since the computing at the edge is independent of each other, the optimization goal is adjusted to minimize the computing delay of each edge. Therefore, for the specified edge base station m, we define Its problem can be reformulated as:

[0216]

[0217] After determining v n,m To effectively make the problem a convex optimization problem, the Lagrange multipliers and KKT conditions can be used to solve it to find the global optimal solution. Its Lagrange rule is established as follows:

[0218]

[0219] where ρ and τ are non - negative Lagrange multipliers related to the constraints;

[0220] S342: Let represent the optimal solution, then the KKT conditions can be derived as:

[0221]

[0222] Obviously, for this sub - problem, we only need to focus on the case of Therefore, we can derive:

[0223]

[0224] Since ρ * > 0, we obtain The computing resources we allocate to the n - th agent at time slot t are derived as:

[0225]

[0226] Finally, we solved this sub - problem in this way.

[0227] The above is only the preferred embodiment of the present invention. It should be pointed out that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can also be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.

[0228] Specifically, as Figure 3 and Figure 4 shown, in the execution process of a MADDPG - CCRC algorithm proposed in S35 of S3, the steps are as follows:

[0229] S35: The algorithm proposed by the present invention is named a MADDPG - CCRC algorithm, and its steps are as follows:

[0230] S351: The algorithm starts to execute, input and initialize relevant system parameters, including user equipment, number of base stations, number of orthogonal channels, signal-to-noise ratio, channel parameters, etc.;

[0231] S352: Initialize the policy set of the four parameters to be optimized, and complete the setting of algorithm hyperparameters, etc.;

[0232] S353: Reset the environment variables, enter the iterative main loop, and put the state space, action space, reward function, and the next moment state space mini-batch into the replay pool;

[0233] S354: Observe the current state to generate a random policy; and judge whether it is less than the greedy parameter according to the condition. If the condition holds, randomly select an action and generate an action set, and then return to execute S353; otherwise, obtain the average maximum value of the action set according to the policy, output the channel selection and offloading policy sets, and execute S359;

[0234] S355: After solving sub-problem one, set both its user base station and channel selection to 1, calculate relevant parameters to determine the priority level of the agents that select the same base station, and send the result back to each agent to facilitate solving subsequent problems;

[0235] S356: Allocate computing resources for each user. If offloading is performed, make a judgment. If the transmission power is greater than the maximum limit power, set the relevant parameters to zero and return to execute S353; otherwise, use the transmission power set to solve sub-problem two and execute S359;

[0236] S357: Calculate the current reward value, observe and record the next moment state, and put the state-reward combination into the experience replay pool;

[0237] S358: Calculate the Q value and the loss function, update the internal network parameters of the algorithm; and obtain the optimal resource allocation scheme according to the convex optimization algorithm and the KKT condition to solve sub-problem three;

[0238] S359: Output the channel selection policy, offloading policy, transmission power policy, and computing resource allocation policy. After finding the best approximate policy value, the algorithm execution is completed.

[0239] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention. Without departing from the concept of the present invention, several deformations and improvements can also be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.

Claims

1. A communication resource allocation method for assisting mobile edge joint optimization, characterized in that, The method includes the following steps: S1: Build a system model for non-orthogonal multiple access (NOMA)-assisted mobile edge computing, which includes a time-slot communication model and a computing model. The computing model includes a local computing model and an edge computing model. This framework consists of multiple edge nodes and multiple users. The base stations are divided into two types: static base stations (SBS) and mobile base stations (MBS). These two types of base stations have different computing capabilities, and each base station is equipped with a successive interference cancellation (SIC) receiver and a multi-access edge computing (MEC) server. The spectrum of the entire system is divided into several orthogonal channels for wireless transmission connection to complete the allocation of computing task resources. The model specifically includes channel parameters, base station parameters, the number of users, the number of orthogonal channels, work task data packets, local and edge offloading parameters. S2: Based on the system model constructed in S1, a multi-party multi-node NOMA-assisted mobile edge computing system framework is developed. Considering the time delay and energy consumption in task completion, the task cost is defined as a weighted combination of time delay, energy consumption, and quality of service (QoS). While complying with the QoS requirements under scheduling and resource allocation constraints, the goal is to minimize the weighted sum of delay, energy consumption, and QoS. Specifically, it is optimized by searching for four parameters: a feasible channel selection strategy, an edge offloading strategy, a transmission power set, and an edge computing resource allocation strategy. The multi-objective optimization problem is established through a Markov decision process to minimize the system's cost function. S3: Due to the non-convexity of the problem, the multi-objective optimization problem is transformed into three sub-problems for solution, and a MADDPG-CCRC algorithm is proposed to jointly optimize the communication and computing resource allocation, and finally find the optimal strategy. The interaction between the agent and the environment is carried out by establishing a state space, an action space, and a reward function in the Markov decision process that adapt to the constructed system environment. Deep reinforcement learning technology is used to determine the joint channel and offloading strategy to solve the first sub-problem; the power allocation is optimized according to the characteristics of non-orthogonal multiple access to solve the second sub-problem; a convex optimization algorithm is applied to obtain the optimal solution of the computing resource allocation to solve the third sub-problem. This system combines static and mobile base stations, which have different computing capabilities according to their types. At the same time, the tasks generated by users show heterogeneity and time dynamics. Within this framework, an optimization problem aiming to minimize the system cost is formulated, involving the joint optimization of channel selection, offloading decision, transmission power, and computing resource allocation.

2. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, wherein The steps for building the system model for NOMA-assisted mobile edge computing described in S1 are as follows: S11: Consider a non-orthogonal multiple access assisted mobile edge computing system with multiple edges and multiple terminals; N user equipments (UEs) are randomly distributed in a circular area with a radius of R and communicate with M base stations, where M1 are static base stations and M2 are mobile base stations, M = M1 + M2, each base station is equipped with a receiver and an edge server, the computing capabilities of the edge servers associated with the static base stations and the mobile base stations are represented as F1 and F2 respectively, the spectrum of the whole system is divided into K orthogonal channels, and we use UE n , ED m and C k to represent the nth user equipment, the mth edge (i.e., base station), and the kth channel respectively; S12: In each time slot, each user equipment generates a computing task, and the tasks of several user equipments are represented as Γ n =(d n , w n , T max,n ) 1≤n≤N, where d n represents the size of the data packet, w n represents the number of CPU cycles required for computing, and T max,n is the deadline of the task, limited within the duration of one time slot; S13: Due to limited computing resources, user equipment can choose to offload its tasks to the edge server for help. Assume that within a time slot, all base stations and user equipment remain stationary, but in different time slots, the mobile base station changes its position along with the user equipment.

3. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, wherein The relevant parameters of the communication model described in S1 are as follows: S14: If the user equipment chooses to offload its task, it first selects a channel to transmit its data packets; non-orthogonal multiple access technology is applied, where each base station is equipped with an L-SIC receiver, which means that each time-frequency resource block can transmit signals for L users simultaneously. Since the system has K channels, at most L*K data packets can be transmitted in one time slot. At the base station, the received signals of each channel are decoded in descending order of received power. That is to say, by applying the receiver, the base station first decodes the user equipment with the strongest received power, treats other signals from users on the same channel as interference, and thus distinguishes the offloading priorities of user equipment. The channel selection strategy is defined as , where u n,k is expressed as: When k = 0, the user equipment chooses not to transmit data packets, i.e., local computing. The user equipment signal at ED m Through C k The received power is: Among them, is the transmission power of the user, h n,m,k and α represent the small-scale Rayleigh fading coefficient and the large-scale path loss. |h n,m,k | 2 follows an exponential distribution with unit parameter, r n,m is the Euclidean distance between the user and the base station; According to Shannon's theorem, the signal-to-noise ratio of the user from ED m to C k is expressed as: where Ⅱ{x} ∈ {0,1} is the indicator function, and σ 2 represents the noise power. The condition for the base station to successfully decode is that the SINR is greater than the threshold θ. After successful decoding, the strongest signal is subtracted from the received signal, and then the process is repeated; if decoding fails, the decoding process will terminate immediately, and the remaining undecoded signals are regarded as failures. S15: Local computing model: The edge offloading strategy is defined as: When m = 0, the user performs local computing, which is expressed as local computing power. Then the computing latency and energy consumption are respectively: where k is the energy consumption constant, usually taking the value of 10 -28 ; We considered time delay and energy consumption when the task was completed, and defined the task cost as a weighted combination of these factors, expressed as: where β is the weight of energy consumption. S16: Edge Computing Model: Users who complete computing tasks through offloading need to be divided into two stages: the stage of transmitting tasks to the edge through the channel and the stage of completing task computing at the edge. Therefore, the offloading delay includes the delay involved in transmitting data packets and the delay generated during edge computing; the transmission delay depends on the size of the data packet represented by d n and the transmission rate is expressed as R n,m,k , and the transmission rate of the user device from the channel to the base station edge is obtained as: R n,m,k = v n,m B log2(1 + SINR n,m,k ) where B is the channel bandwidth, then for v n,m = 1, the transmission delay Γ of the task n is expressed as: For v n,m = 1, we represent the computing resources allocated by the base station to the user as Therefore, for task Γ n the offloading computing delay can be expressed as: Regarding energy consumption, since the user side is more restricted by battery limitations, this paper only considers the energy required for data packet transmission on the user side. Then, in the edge computing model, the energy consumption required to complete a task can be expressed as: Similar to the local computing model, the task cost function of edge computing is defined as:

4. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, wherein As described in S2, the goal is to minimize the weighted sum of latency, energy consumption, and quality of service. Specifically, it is optimized by searching for four parameters: a feasible channel selection strategy, an edge offloading strategy, a transmission power set, and an edge computing resource allocation strategy. The multi-objective optimization problem of minimizing the system cost function is established through a Markov decision process, including the following steps: S21: Consider the time delay and energy consumption in task completion; define the task cost as a weighted combination of these factors, which can be expressed as: where x n,m,k = u n,k ∩ v n,m , considering different levels of time sensitivity among different tasks, Z is enhanced n by incorporating whether the task can be completed within the deadline: Z n +η n δ(T n -T max,n ); where T n represents the latency required to complete the task and can be expressed as: where η n represents the importance of the task, with particular attention paid to the sensitivity to latency and related to T max,n , and δ(x) is defined as: The total cost function of the system model is the sum of the costs of all tasks and can be expressed as: S22: and respectively represent the set of user transmission capabilities and the base station computing resource allocation strategy. Our goal is to search for a feasible channel selection strategy edge offloading strategy transmission power set and edge computing resource allocation strategy to minimize the cost function of the system. Therefore, the optimization problem is formulated as follows: x n,m,k = u n,k ∩ v n,m ; S23: The constraint conditions of the above S22 optimization problem are respectively expressed as follows: S231: F m represents the computing resources owned by the base station, P max,n represents the maximum transmission power of the user, and imposes a constraint on the users, allowing them to select at most one channel to send their data packets; S232: and restrict the user from offloading their tasks to at most one edge node; S233: x n,m,k = u n,k ∩ v n,m represents a joint solution for channel selection and edge offloading; only when the two are combined will the user transfer their tasks to the base station through the channel; S234: and indicates that the computing resources allocated to tasks by each edge server cannot exceed their own computing resources; S235: Ensure that the user's transmission power does not exceed its maximum allowable value.

5. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, wherein S3 transforms the optimization problem into three sub-problems for separate solution, including the following steps: S31: Obviously, the optimization problem is a mixed integer and non-convex problem, which is usually classified as a difficult problem, making it challenging to solve using traditional optimization methods. To address this issue, it involves decomposing the problem into multiple sub-problems and gradually optimizing the variables; more specifically, initially, deep reinforcement learning techniques are used to determine the joint channel and offloading strategy. Next, the power allocation is optimized according to the characteristics of non-orthogonal multiple access. Finally, a convex optimization algorithm is applied to obtain the optimal solution for the computing resource allocation. ​ S32: In consecutive time slots, the optimization problem is regarded as a partially observable multi-agent Markov decision process; each user is regarded as an agent, and partial observability means that each agent only knows its own current information, including but not limited to channel state, packet size, distance to each base station, task deadline, etc., without knowing the information of other agents. To solve the above problems, it is divided into two sub-problems; first, each agent makes independent decisions to solve the joint offloading scheme of the channel and the edge node ; then, the agent calculates the corresponding and to further reduce the cost function. Next, we use the multi-agent Markov decision process to reconstruct the problem, using to represent the state, action, reward of the Markov decision process and their respective sets, and provide specific definitions in the following sections: S321: List the state space in the Markov decision process of the optimization problem. At time slot t, use s nt to represent the state of agent N, including the action in the previous time slot, the current state of task N, local and edge computing resources, and channel state information, which is expressed as: where o nt = a n(t-1) represents the action output of the previous time slot, and represents the matrix of channel gains and the distance to the edge at the current time slot, represents the computing resources available at the edge side, and the state set of all tasks is denoted as s t = {s nt , 1 ≤ n ≤ N}; S322: List the action space in the Markov decision process of the optimization problem. Similar to the state, represent the action of agent n at time slot t as a nt which includes the index of the accessed channel and the offloaded edge server, and is expressed as: a nt =(C nt ,ED nt ); where C nt ∈ {0, 1, 2, ..., K} and ED nt ∈ {0, 1, 2, ..., M}, when C nt or ED nt is 0, it means that the task will perform local computation; according to the actions of the agent, we obtain the corresponding and to derive the required channel selection strategy and offloading scheme , and all the action sets can be expressed as: a t = {a nt , 1 ≤ n ≤ N}; S323: Divide the reward into two parts: one is the offloading reward, and the other is the quality-of-service reward. The offloading reward is defined as the cost reduction achieved by performing offloading calculations compared to local computing, which is expressed as: The quality-of-service reward is the reward obtained for completing a task within the deadline, which is defined as: Therefore, the reward of agent n and the total reward of the system at time slot t are: S324: Therefore, regarding the determination of the optimal channel and offloading joint scheme and can be rephrased as: x n,m,k = u n,k ∩ v n,m ; Where γ ∈ [0,1) represents the discount factor reflecting the influence of future time slots on the current time slot. The higher γ is, the greater the influence of future rewards on the current reward. The goal is to maximize the long-term reward.

6. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, characterized in that The above optimization problem is divided into three sub-problems for separate solution. Sub-problem one is solved as in S31 and S32, and sub-problems two and three are solved including the following steps: S33: Power allocation sub-problem: After determining the user specifies the computing model. In the case of local computing, the cost is completely determined by the computing resources required by the task and the local computing power, resulting in a fixed value. On the contrary, if offloading computing is selected, the transmission power needs to be further optimized because the transmission power affects the transmission rate and thus the transmission time. Our goal is to optimize the transmission power to minimize the transmission time. Therefore, this sub-problem is formulated as follows: s.t.x n,m,k = u n,k ∩ v n,m ; For the edge time, since the transmission power only affects the transmission delay and does not affect the offloading calculation delay, the term is omitted S34: Computational resource allocation sub-problem: Allocate computational resources at the edge according to the received computational tasks. Similar to the transmission power problem, the configuration of computational resources only affects the offloading calculation delay. Our goal is to optimize the offloading calculation delay. The sub-problem is formulated as follows: s.t.x n,m,k = u n,k ∩ v n,m ; 7. The communication resource allocation method for assisting mobile edge joint optimization according to claim 1, wherein A key aspect of specifically solving sub-problem two is to introduce non-orthogonal multiple access technology and jointly optimize communication and computing in combination with mobile edge computing: S331: The condition for successful decoding is that the SINR of the received signal is greater than or equal to a preset threshold, i.e., SINR n,m,k ≥θ. Constrained by the limited battery capacity of the user equipment to extend battery life and reduce transmission power while maintaining successful decoding; thus, the minimum received power for successful decoding in the first power level is: S332: The channel gain of the n-th agent reaching edge m through channel k is defined as Agents with higher channel gains are assigned higher power levels because this assignment aims to minimize the total transmission power; subsequently, the assigned levels, as well as will be sent back to each agent; when the n-th agent receives the signal returned by the environment, the corresponding transmit power is calculated as: where l nt represents the power level allocated to the n-th agent at time slot t; if it exceeds the allowed maximum power, the n-th agent performs local computation, and in this way, the sub-problem is solved.

8. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, characterized in that The steps to specifically solve sub-problem three are as follows: S341: After receiving the computing tasks, the edge base stations will optimize the computing resource allocation to further reduce the system cost. Since the computations at the edge are independent of each other, the optimization objective is adjusted to minimize the computing latency of each edge. Therefore, for a specified edge base station m, we define Its problem can be reformulated as: After determining v n,m To effectively transform the problem into a convex optimization problem, the Lagrange multipliers and KKT conditions can be used to solve it in order to find the global optimal solution. Its Lagrange rule is established as follows: Where ρ and τ are non-negative Lagrange multipliers related to the constraints; S342: Let denote the optimal solution, then the KKT conditions can be derived as follows: Obviously, for this sub - problem, only the situation of needs to be concerned about. Therefore, it can be deduced that: Since ρ * > 0, it follows that The computational resources we allocate to the nth agent at time slot t are derived as: Finally, we solve this sub-problem in this way.

9. The communication resource allocation method for auxiliary mobile edge joint optimization according to claim 1, wherein The specific MADDPG-CCRC algorithm proposed to solve the problem is executed as follows: S35: The proposed algorithm is called a MADDPG-CCRC algorithm, and its steps are as follows: S351: The algorithm starts to execute, input and initialize relevant system parameters, including user equipment, number of base stations, number of orthogonal channels, signal-to-noise ratio, and channel parameters, etc.; S352: Initialize the policy sets of the four parameters to be optimized and complete the setting of algorithm hyperparameters, etc.; S353: Reset the environmental variables, enter the iterative main loop, and put the state space, action space, reward function, and next moment state space in small batches into the replay pool; S354: Observe the current state to generate a random policy; and determine whether it is less than the greedy parameter according to the condition. If the condition holds, randomly select an action and generate an action set, and then return to execute S353; otherwise, obtain the average maximum value of the action set according to the policy, output the channel selection and offloading policy sets, and execute S359; S355: After solving sub-problem one, set both its user base station and channel selection to 1, calculate relevant parameters to determine the priority levels of the agents that select the same base station, and send the results back to each agent for solving subsequent problems; S356: Allocate computing resources to each user. If offloading is performed, make a judgment. If the transmission power is greater than the maximum limit power, set the relevant parameters to zero and return to execute S353; otherwise, use the transmission power set to solve sub-problem two and execute S359; S357: Calculate the current reward value, observe and record the state at the next moment, and put the state-reward combination into the experience replay pool; S358: Calculate the Q value and the loss function, update the internal network parameters of the algorithm; and obtain the optimal resource allocation scheme according to the convex optimization algorithm and the KKT condition to solve sub-problem three; S359: Output the channel selection policy, offloading policy, transmission power policy, and computing resource allocation policy. After finding the optimal approximate policy value, the algorithm is completed.