A Resource Allocation Method for Semantic Communication Systems Oriented to Image Restoration Tasks
The LSTM-NOISE MAPPO algorithm optimizes resource allocation of semantic communication system, solves the problem of scarcity of wireless resources in 6G network, realizes efficient semantic service transmission, and improves system performance and stability.
Patent Information
- Application Number
- CN202411616037.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-13
AI Technical Summary
In 6G networks, there is scarcity of wireless resources in the semantic communication system, which leads to difficulty in managing communication resources and is difficult to achieve high-quality semantic service transmission.
The semantic communication resource allocation method based on LSTM-NOISE MAPPO is adopted, and the power, semantic compression rate and calculation capacity allocation of the semantic communication system are optimized through channel estimation, signal-to-interference noise ratio analysis, modulation method calculation and deep reinforcement learning algorithm.
It improves the system's data rate, system capacity and user experience, reduces energy consumption stability and delay stability, and improves symbol error performance.
Smart Images

Figure CN119545525B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies, and particularly relates to a resource allocation method for a semantic communication system for image restoration tasks. Background Art
[0002] With the acceleration of 5G commercialization, 6G has gradually become the direction for countries around the world to focus on deployment. In recent years, with the strong promotion and support of the Chinese government and driven by the strategic needs of the country's medium- and long-term development, in order to maintain and expand China's leading edge in 5G, China is actively exploring the technical route of 6G independent innovation. The 6G network is expected to connect people and machines with different levels of intelligence to each other, and semantic communication is one of the key technologies currently being considered for 6G. Semantic communication systems mainly focus on the semantic representation, transmission, and reconstruction of the source content, as well as semantic-based wireless transmission. Further considering that when there are multiple users at the sending end or the receiving end of a semantic communication system simultaneously, a semantic communication network is formed, and the scarcity of wireless resources in the network makes communication resource management very important. Therefore, there is an urgent need to study a resource allocation strategy that can not only improve the overall performance of the network but also support high-quality semantic services. Through reasonable communication resource allocation, better communication services can be provided to users, such as higher data rates, larger system capacities, and better user experiences. Therefore, reasonable allocation of communication resources is the most direct and effective means to solve the shortage of communication resources in a complex communication environment. Summary of the Invention
[0003] The purpose of the present invention is to provide a new resource allocation method for a semantic communication system based on MAPPO. On the basis of considering the dynamic changes of user locations and the physical channels of semantic communication, the proposed LSTM-NOISE MAPPO semantic communication resource algorithm is used to solve the allocation problems of power, semantic compression ratio, and computing capacity in a cell semantic communication system.
[0004] The technical solution for achieving the purpose of the present invention is as follows: On the one hand, a resource allocation method for a semantic communication system for image restoration tasks is provided, and the method includes:
[0005] Step 1, establish a semantic communication system, including a base station and several users randomly distributed within the signal coverage range of the base station; a neural network for extracting image semantic information is stored on each user, and an edge computing node is installed at the base station, and this node stores a neural network for restoring images using semantic information;
[0006] Step 2, the users send pilot signals to the base station, and the different pilot signals are orthogonal to each other;
[0007] Step 3, the base station performs channel estimation on the uplink channel of each user based on the pilot signals sent by all users;
[0008] Step 4: Conduct theoretical analysis on the signal-to-interference-plus-noise ratio (SINR) of the transmitted signal at the receiving end and the statistical characteristics of the SINR based on the uplink channel estimation results.
[0009] Step 5: Analyze the transmission energy consumption of the semantic information sent by the user using the SINR of each user, and calculate the theoretical average symbol error performance of the signal sent by each user according to the modulation method.
[0010] Step 6: Obtain the mapping curve of the number of CPU cycles required for the user to extract semantic information and the base station to restore the image information under different semantic extraction rates, perform numerical fitting on the mapping curve, and calculate the computing energy consumption for extracting semantic information and restoring image information based on the fitting results.
[0011] Step 7: Construct an optimization problem, and simplify the problem to obtain the final semantic communication resource allocation problem that needs to be optimized.
[0012] Step 8: Design and build a new deep reinforcement learning algorithm (LSTM-NOISE MAPPO semantic communication resource allocation algorithm) based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm to solve the semantic communication resource allocation problem. During the training process, an offline training of the user is carried out by constructing a dynamically changing training environment by updating the position and channel state of the user in each interaction round until the training result converges.
[0013] On the other hand, a semantic communication system resource allocation system for an image restoration task is provided. The system includes:
[0014] The first module is used to establish a semantic communication system, including a base station and several users randomly distributed within the signal coverage area of the base station; a neural network for extracting image semantic information is stored on each user, and an edge computing node is installed at the base station, and a neural network for restoring the image using semantic information is stored on this node.
[0015] The second module is used for the user to send pilot signals to the base station, and the different pilot signals are orthogonal to each other.
[0016] The third module is used for the base station to perform channel estimation on the uplink channel of each user based on the pilot signals sent by all users.
[0017] The fourth module is used to conduct theoretical analysis on the signal-to-interference-plus-noise ratio (SINR) of the transmitted signal at the receiving end and the statistical characteristics of the SINR based on the uplink channel estimation results.
[0018] The fifth module is used to analyze the transmission energy consumption of users sending semantic information by using the SINR of each user, and calculate the theoretical average symbol error performance of the signals sent by each user according to the modulation mode;
[0019] The sixth module is used to obtain the mapping curve of the number of CPU cycles required for users to extract semantic information and the base station to restore image information under different semantic extraction rates, perform numerical fitting on the mapping curve, and calculate the computational energy consumption for extracting semantic information and restoring image information based on the fitting results;
[0020] The seventh module is used to construct an optimization problem and simplify the problem to obtain the final semantic communication resource allocation problem that needs to be optimized;
[0021] The eighth module is used to design and build a new deep reinforcement learning algorithm based on the multi-agent proximal policy optimization algorithm to solve the semantic communication resource allocation problem. During the training process, an offline training of users is carried out by constructing a dynamically changing training environment by updating the positions and channel states of users in each interaction round until the training results converge.
[0022] Compared with the prior art, the remarkable advantages of the present invention are as follows:
[0023] (1) In the construction of the communication system, the present invention conducts channel estimation and analyzes to obtain a closed-form expression of the average symbol error performance of the user communication link, and constructs a semantic communication theory model closer to the actual communication process, further laying a foundation for users to adaptively adjust the resource allocation strategy during the offline training process.
[0024] (2) The proposed LSTM-NOISE MAPPO semantic communication resource allocation algorithm can not only solve discrete resource allocation strategies but also has the ability to solve continuous resource allocation strategies. At the same time, the algorithm of the present invention can not only solve the single-modal semantic communication resource allocation problem but also be applied to the solution of the multi-modal multi-task semantic communication resource allocation problem.
[0025] (3) The proposed LSTM-NOISE MAPPO semantic communication resource allocation algorithm solves the problem of poor training effect brought by the dynamic change of the environment to the MAPPO algorithm. The advantages of the proposed algorithm compared with the MAPPO algorithm are specifically manifested as follows:
[0026] This algorithm can bring better data utilization rate to the system than the MAPPO algorithm, specifically manifested as having fewer training convergence steps;
[0027] This algorithm can bring better energy consumption stability to the system than the MAPPO algorithm, specifically manifested as that after training convergence, it can make the system have a smaller energy consumption variance on the premise of achieving the same training effect as the MAPPO algorithm;
[0028] This algorithm can bring better latency stability to the system. Specifically, after training convergence, it can make the system have a smaller latency variance on the premise of achieving the same training effect as the MAPPO algorithm.
[0029] This algorithm can bring better symbol error performance to the system. Specifically, after training convergence, the system can obtain a smaller average maximum symbol error rate and variance.
[0030] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings
[0031] Figure 1 It is a specific flowchart of the resource allocation method for the semantic communication system of the present invention for the image restoration task.
[0032] Figure 2 It is a schematic diagram of the cell semantic communication system related to the present invention.
[0033] Figure 3 It is a schematic diagram of the image sending and restoration process.
[0034] Figure 4 It is a comparison chart of the system energy consumption of the algorithm of the present invention and other different algorithms, where Figure 4 (a) to (d) in it are the system energy consumption of the algorithm of the present invention, the LSTM MAPPO algorithm, the NOISE MAPPO algorithm, and the MAPPO algorithm respectively.
[0035] Figure 5 It is a comparison chart of the system latency of the algorithm of the present invention and other different algorithms, where Figure 5 (a) to (d) in it are the system latency of the algorithm of the present invention, the LSTM MAPPO algorithm, the NOISE MAPPO algorithm, and the MAPPO algorithm respectively.
[0036] Figure 6 It is a comparison chart of the system link rate performance of the algorithm of the present invention and other different algorithms, Figure 6 and (a) to (d) in it are the system link rate performance of the algorithm of the present invention, the LSTM MAPPO algorithm, the NOISE MAPPO algorithm, and the MAPPO algorithm respectively. Detailed Embodiments
[0037] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0038] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0039] In one embodiment, in combination with Figure 1 , a method for resource allocation of a semantic communication system for an image restoration task is provided. The method includes:
[0040] Step 1, establish a semantic communication system, including a base station and several users randomly distributed within the signal coverage area of the base station; a neural network for extracting image semantic information is stored on each user, and an edge computing node (Edge Computing Point, ECP) is installed at the base station, and a neural network for restoring an image using semantic information is stored at this node; as Figure 2 shown, in the center of the cell, there is a base station (BS) equipped with N r antennas and an edge computing node (Edge Computing Point, ECP), providing semantic communication services for M single-antenna users randomly scattered in the cell (N r ≥ M). The device set is defined as
[0041] Step 2, the user sends a pilot signal to the base station, and the different pilot signals are orthogonal to each other;
[0042] Step 3, the base station estimates the uplink channel of each user based on the pilot signals sent by all users;
[0043] Step 4, based on the uplink channel estimation result, conduct a theoretical analysis on the signal-to-interference-plus-noise ratio (SINR) of the transmitted signal at the receiving end and the statistical characteristics of the SINR;
[0044] Step 5, use the SINR of each user to analyze the transmission energy consumption of the user sending semantic information, and calculate the theoretical average symbol error performance of the signal sent by each user according to the modulation method;
[0045] Step 6: Obtain the mapping curves of the CPU cycle numbers required for the user to extract semantic information and the base station to restore image information under different semantic extraction rates, perform numerical fitting on the mapping curves, and calculate the computational energy consumption for extracting semantic information and restoring image information based on the fitting results.
[0046] Step 7: Construct an optimization problem and simplify it to obtain the final semantic communication resource allocation problem to be optimized.
[0047] Step 8: Design and build a new deep reinforcement learning algorithm (Multi-Agent Proximal Policy Optimization, MAPPO) based on the multi-agent proximal policy optimization algorithm to solve the semantic communication resource allocation problem. During the training process, an offline training of the user is carried out by constructing a dynamically changing training environment by updating the user's location and channel state in each interaction round until the training result converges.
[0048] Furthermore, in one embodiment, step 1 further includes:
[0049] Before performing the semantic communication task, train the semantic transceiver networks of both the semantic communication transceiver parties for the image restoration task. After the training is completed, the transceiver network parameters corresponding to specific semantic background knowledge are obtained. The semantic transceiver network includes a neural network for the user side to extract image semantic information and a neural network for the edge computing node to restore the image using semantic information. After the semantic transceiver network training is completed, the user and the base station perform the image restoration task under specific semantic background knowledge through semantic communication. The schematic diagrams of the above training process and the communication task execution process are as Figure 3 shown.
[0050] Consider using an autoencoder as the structure of the transceiver network. The semantic background knowledge base used for training is associated with the specific semantic task to be performed, and can include image background knowledge in specific application fields such as medicine, transportation, and chemical engineering. After the training is completed, the transceiver network parameters corresponding to specific semantic background knowledge are obtained.
[0051] Furthermore, in one embodiment, in step 2, at the beginning of each channel coherence interval, M single-antenna users in the cell simultaneously send mutually orthogonal pilot signals to the base station. The number of antennas of the base station is denoted as N r , the number of pilot signals is M, the pilot dimension size is 1×τ p , the pilot vector sent by the kth user is denoted as Φ k , and its elements are denoted as satisfying τ p is greater than or equal to M, and the transmission power of the pilot signal is the maximum transmission power of the user, denoted as p k. To save pilot overhead, let τ p = M.
[0052] Further, in one embodiment, step 3 specifically includes:
[0053] Let denote the uplink fading between the N r antennas of the base station in the cell and the k-th user, expressed as:
[0054]
[0055] where h k is an N r ×1 vector, and each element follows a complex Gaussian distribution with a mean of 0 and a variance of 1, representing the small-scale fading between the user and each antenna of the base station; β k is a scalar representing the large-scale fading between the k-th user and the base station;
[0056] After all users send pilot signals, the received signal at the base station side is expressed as Y:
[0057]
[0058] where the dimension of Y is N r ×M, M is the number of pilot signals, and n is a complex additive Gaussian noise matrix of dimension N r ×M, and its elements satisfy independent and identically distributed CN(0,σ 2 ), and CN(0,σ 2 ) represents a complex Gaussian distribution with a mean of 0 and a covariance of σ 2 ; obtained from the orthogonality between different pilot vectors, Φ j represents the pilot sent by the j-th user, and H represents the conjugate transpose;
[0059] Let right multiply both sides of formula (2) to obtain:
[0060]
[0061] where has a dimension of N r ×1;
[0062] Based on formula (3), use the MMSE method for channel estimation to obtain the uplink channel of the k-th user as
[0063]
[0064] Further, in one embodiment, step 4 specifically includes:
[0065] (1) After channel estimation, the received signal at the base station during the communication process is expressed as y:
[0066]
[0067] Where x k represents the transmission signal of the k-th user, and p′ k represents the signal transmission power of the k-th user,
[0068] For any k-th user, at the base station side, by calculating the inner product of the weight vector v r with dimension 1×N k and the received signal y to perform linear combining processing on the received signal, which is expressed as:
[0069]
[0070] Where p′ k' represents the signal transmission power of the k'-th user, and x k' represents the transmission signal of the k'-th user, represents the channel estimation value of the k'-th user,
[0071] Based on formula (6), the maximum SINR of the k-th user is expressed as γ k :
[0072]
[0073] Where is the identity matrix with dimension N r ×N r , and C k′ is the mean square value of the channel estimation error of the k'-th user, which is expressed as:
[0074]
[0075] Where β k' is a scalar representing the large-scale fading between the k'-th user and the base station;
[0076] For the k-th user, the general closed-form expression of the SINR distribution of its uplink channel has high complexity, and it is difficult to obtain an accurate closed-form solution when the number of antennas configured in the system is large. Therefore, an asymptotic method is adopted to analyze the probability density function of the uplink channel SINR of any user in this system.
[0077] (3) Analyze the probability density function of the uplink channel SINR of any user by the asymptotic method
[0078] Let the matrix have its eigenvalues converge to λ1. According to the Marchenko-Pastur distribution criterion, the mean of λ1 is calculated as:
[0079]
[0080] Therefore, Equation (7) is simplified to:
[0081]
[0082] where represents the j-th element of the channel estimation vector of the k-th user, and the expression of φ is as follows:
[0083]
[0084] The right-hand side of Equation (10) consists of two terms. Each term can be regarded as the sum of exponential random variables with the same mean, which follows a chi-square distribution. Therefore, the statistical distribution of SINR in Equation (10) is characterized as the convolution of the distributions of the two terms on the right-hand side of the equation, expressed as:
[0085]
[0086] where
[0087]
[0088] N1 = M - 1
[0089]
[0090] N2 = N r -M + 1
[0091]
[0092] Furthermore, in one of the embodiments, step 5 specifically includes:
[0093] For various digital modulation systems, such as PAM, PSK, and QAM, their average symbol error rate is expressed by the Gaussian Q function:
[0094]
[0095] where the Gaussian Q function is expressed as:
[0096]
[0097] The parameters a and b are constants determined by the specific modulation scheme;
[0098] Substituting Equation (14) into Equation (13) gives:
[0099]
[0100] Among them, the closed expression of C(x) is as follows:
[0101]
[0102] Among them,
[0103] Substituting C(x) into formula (15) gives the expression of.
[0104] Furthermore, in one of the embodiments, step 6 specifically includes:
[0105] For each user, its energy consumption includes two parts: the energy consumption E 1k for extracting semantic information and the energy consumption E 2k for uplink transmission;
[0106] (1) The energy consumption E 1k for extracting semantic information is:
[0107] E 1k = κy 1k (D k , S k )f 2 (17)
[0108] In the formula, y 1k (D k , S k ) represents the number of CPU cycles for the k-th user in the cell to extract semantic information. D k and S k respectively represent the data volume sizes before and after the k-th user extracts semantics. This size value is related to the semantic extraction rate ρ k . f represents the computing capacity of the k-th user. It is assumed that the local computing capacities of each user in the system are equal; κ represents the energy consumption coefficient;
[0109] The relationship between y 1k (D k , S k ) and ρ k is obtained by numerical fitting and is expressed as follows:
[0110] y 1k (D k , S k ) = η1ρ k + η2 (18)
[0111] Wherein, η1>0 and η2>0 are constant parameters respectively; it should be noted that the numerical fitting results obtained from formula (18) are achieved under the semantic coding model selected in the present invention and the simulation software and hardware conditions. Specific analysis should be carried out for the numerical fitting research under different semantic coding models and software and hardware conditions.
[0112] The time for the user to extract semantic information is:
[0113]
[0114] (2) The energy consumption E of the uplink transmission 2k is:
[0115] E 2k =t 2k p k ′(20)
[0116] Wherein, the time required for the uplink transmission is t 2k :
[0117]
[0118] In the formula, Cap k represents the uplink channel capacity of the k-th user, and its expression is Cap k =Wlog2(1 + γ k ), where W is the uplink bandwidth allocated to each user, and S k is the data volume size after the k-th user extracts semantic information;
[0119] (3) For the edge computing node, the energy consumption is the computing energy consumption when performing image restoration on the semantic information sent by the users in the cell, which is expressed as:
[0120]
[0121] In the formula, E 3k represents the energy consumption for decoding the information sent by the k-th user, represents the number of CPU cycles required for decoding the information sent by the k-th user, and respectively represent the data volume size of the semantic information after wireless channel transmission and the data volume size after semantic decoding. Since the neural network structures in the semantic coding and semantic decoding processes are symmetric, it can be considered that f 3k represents the computing capacity allocated to the k-th user;
[0122] The corresponding time consumption is:
[0123]
[0124] In summary, the total energy consumption \(E\) of the \(k\) -th user for performing its semantic communication task k is expressed as: \(E\) k = \(E\) 1k + \(E\) 2k + \(E\) 3k .
[0125] Furthermore, in one embodiment, the system objective is to minimize the total energy consumption of all users in the semantic communication system by adjusting the power of each user, the semantic compression ratio, and the computational capacity allocation of the edge computing nodes under the semantic communication task. Therefore, the optimization problem constructed in step 7 is as follows:
[0126]
[0127] This optimization problem is a mixed - integer non - linear programming problem. Among them, the constraint condition \(C1\) is the delay constraint for each user to perform the semantic communication task in the cell, \(C2\) represents the accuracy constraint for each user to perform the semantic task, \(t\) th represents the upper limit value of the delay constraint for each user, \(u\) k (\(D\) k , \(S\) k ) represents the semantic accuracy of the \(k\) -th user for semantic communication, \(A\) k represents the lower limit value of the semantic accuracy constraint for the \(k\) -th user, \(C3\) and \(C4\) are the computational capacity constraints of the edge computing nodes, \(F\) max represents the maximum computational capacity of the edge computing nodes, \(C5\) is the signal transmission power constraint for each user, \(p\) level1 , \(p\) level2 , \(p\) level3 , \(p\) level4 respectively represent the optional powers, \(C6\) is the link symbol error performance constraint for each user, is the upper limit value of the symbol error performance constraint for each user, \(C7\) represents the link rate sum constraint of all users in the cell, \(\Gamma\) th represents the lower limit value of the link rate sum constraint,
[0128] During the semantic communication process, since the higher the semantic extraction rate will make the transmitted signal carry more information, it can be considered that there is a positive - correlation characteristic between the execution accuracy of the semantic task and the semantic extraction rate. Therefore, the optimization problem \(P1\) can be equivalent to:
[0129]
[0130]
[0131] In the formula, the \(C2\) constraint is the constraint condition for the semantic extraction rate of each user, \(\rho\)th Represents the lower limit value of the semantic extraction rate constraint for each user.
[0132] Further, in one of the embodiments, step 8 specifically includes:
[0133] For the optimization problem P2, the overall idea of the solution is: decompose the problem into two sub-problems, and by alternately solving the two sub-problems in turn, make P2 converge to the optimal value;
[0134] First, fix the variables ρ k and p′ k When, P2 is transformed into a sub-problem with only a single variable f 3k which is expressed as follows:
[0135]
[0136] P3 is a convex optimization problem, which is solved using the convex optimization tool CVXPY;
[0137] Regarding each user as an agent in deep reinforcement learning, under the condition that information can be shared among all users, define the observation state of the kth user in the tth training round as o k (t)@{Task k (t), Energy(t), Max_Delay(t), Max_SER(t), Action - (t)}, where Task k (t) represents the size of the original data that the kth user needs to transmit in the tth training round, Energy(t) represents the total energy consumption of all users in the (t - 1)th training round, Max_Delay(t) represents the maximum value of the time delay experienced by all users in the (t - 1)th training round, Max_SER(t) represents the maximum value of the average symbol error rate of all user links in the (t - 1)th training round, and define the action of the kth user as a k (t)@{p k (t), ρ k (t)}, where p k (t) represents the power selection scheme of the user in the tth training round, ρ k (t) represents the semantic extraction rate selection scheme of the user in the tth training round, Action - (t) represents the set of actions taken by other users except the user himself in the (t - 1)th training round, that is, Action - (t) = {a1(t) ∪ a2(t)... a k-1 (t) ∪ a k+1 (t)... ∪ a M (t)};
[0138] The training adopts a centralized training method, defining the global state as the set of all user states, that is
[0139] The reward value of each user in the training round is expressed as: The second, third, and fourth terms on the right side of the r(t) equation are penalty terms, and λ1, λ2, λ3, and λ4 are the weight coefficients of the corresponding terms respectively. All users share the same reward value. In each round, each user uses its own policy network to output actions based on its respective observed state and uploads its own state to the base station. At the base station, the edge computing node calculates r(t) and uses the evaluation network to output the global state value function V ζ (s(t)), where ζ represents the parameters of the evaluation network.
[0140] Furthermore, in one of the embodiments, the Long short-term memory-NoiseMAPPO (LSTM-NOISEMAPPO) algorithm proposed in step 8 is designed based on the MAPPO algorithm. The algorithm has innovations in both the user policy network structure and the process of the evaluation network outputting V ζ (s(t)), specifically as follows:
[0141] Benefiting from the theoretical design of the parameter update mechanism for the policy network and the evaluation network, the MAPPO algorithm has strong training stability and optimization expectations. However, in the face of a changing training environment, the MAPPO algorithm often has large fluctuations in training effects due to unknown random factors in the environment. There are probably two reasons for this phenomenon:
[0142] On the one hand, since the policy network structure of MAPPO generally selects a multilayer perceptron neural network (MLP), this neural network structure only focuses on the observed state of the user in each current training round and does not care about the observed states experienced by the user in the past. Therefore, when facing newly emerging observed states, the policy network cannot extract effective information from the observed states experienced in the past to output actions that can help stabilize or improve the training effect for the agent;
[0143] On the other hand, due to the poor generalization ability of the evaluation network in the face of a dynamic environment, when there is a joint state that has not been evaluated before in the environment, it cannot give an appropriate output V ζ (s(t)), resulting in incorrect updates of the policy network parameters.
[0144] Considering the above two reasons, on the premise of considering the limited energy storage of user nodes, this algorithm improves the MAPPO algorithm as follows:
[0145] The new deep reinforcement learning algorithm based on the multi-agent proximal policy optimization algorithm in step 8, which is an improvement on the multi-agent proximal policy optimization algorithm MAPPO, specifically includes:
[0146] (1) Regarding the set of state values of all training rounds of each user in each episode as a time series, designing the policy network of each user as a long short-term memory neural network (LSTM) that can extract time series features to replace the MLP network structure in the policy network of the MAPPO algorithm. In each training round, the policy of the user is obtained by calculating the output of the LSTM network based on the user's observation state;
[0147] (2) Adding noise to the global state s(t) during the output process of the evaluation function, i.e., the global state value function V ζ (s(t)). The global state after adding noise is represented as follows:
[0148]
[0149] In the formula, the concat(·) function represents the concatenation function, whose role is to horizontally concatenate s(t) and the noise matrix x(t). is the global state after adding noise, and each element in x represents a normal distribution with a mean of 0 and a variance of , ε is a scaling factor, ε:U(a,b), and U(a,b) represents a uniform distribution with a numerical range between a and b. In addition, a noise factor is added during the calculation of the advantage function;
[0150] In the t-th round of the user's offline training stage, the policy network of the k-th user selects an appropriate action a k (t) based on the user's observation state o k (t). The action output process is represented as follows:
[0151] Forgetting gate: Determines which information in the previous cell state needs to be forgotten; the activation value f of the forgetting gate t is calculated by the current input o k (t) and the previous hidden state h′ t-1 through a fully connected layer with a sigmoid activation function. W if , W hf , b if , b hf are the parameters of the fully connected layer:
[0152] f t =σ(Wif o k (t) + b if + W hf h′ t-1 + b hf ) (28)
[0153] Input gate: Determines which information in the current input should be added to the cell state, and consists of two parts: one is to determine the acceptance weight i of the information through the sigmoid function t , and the other is to calculate the candidate state g through the tanh function t , W ii , W hi , b ii , b hi , W ig , W hg , b ig , b hg are the corresponding fully connected layer parameters respectively:
[0154] i t = σ(W ii o k (t) + b ii + W hi h′ t-1 + b hi ) (29)
[0155] g t = tanh(W ig o k (t) + b ig + W hg h′ t-1 + b hg ) (30)
[0156] Cell state update: Combines the results of the forget gate and the input gate to update the cell state c′ t :
[0157] c′ t = f t * c t-1 + i t * g t (31)
[0158] In the formula, * represents the bitwise multiplication operation, and c t-1 represents the cell state at the previous moment;
[0159] Output gate: Determines which information in the cell state should be passed to the hidden state at the next moment or used as the model output at the current moment; the activation value o of the output gate t is determined by the current input o k (t) and the hidden state h′ at the previous momentt-1 Calculated through a fully-connected layer with a sigmoid activation function, W io , W ho , b io , b ho are the parameters of the fully-connected layer, and then the final hidden state h′ is obtained by element-wise multiplication with the value of the cell state after passing through the tanh function t :
[0160] o t = σW io o k (t) + b io + W ho h′ t-1 + b ho ) (32)
[0161] h′ t = o t *tanh(c′ t ) (33)
[0162] The result o of the output gate t obtains the action output through sampling by the Categorical function after passing through a fully-connected layer with a RELU function;
[0163] For the k-th user, let the parameters of its target policy network be θ k , and the parameters of the behavioral policy network be θ′ k . The network structures of the target policy and the behavioral policy are the same, but the parameters are different. The policy probability value generated by the target policy network is expressed as π k (a k (t)|θ k , o k (t)), and the policy probability value generated by the behavioral policy network is expressed as π k (a k (t)|θ k ′, o k (t)). After the user interacts with the environment, r(t), V ζ (s(t)), a k (t), π k (a k (t)|θ k , o k (t)) and o k (t + 1) are stored in the cache (replay buffer). When entering the network parameter update process, mini-batch copies of the cache data are taken out from the cache for parameter update. In the above LSTM-NOISEMAPPO algorithm, the policy objective function of each user is expressed as:
[0164]
[0165] where ι(·) represents the observed state o of user k k (t) follows the distribution, V k (·) represents the state value function of the k-th user, where the gradient with respect to the parameter θ k is expressed as:
[0166]
[0167] In the formula, Q k (o k (t), a k (t)) represents the action value function of user k when the observed state is o k (t) and the action is a k (t);
[0168] It is very difficult to directly calculate formula (26). Therefore, an unbiased estimate of formula (35) is calculated to approximate the policy gradient, as shown in formula (36):
[0169]
[0170] Let represent the advantage function, where α and n adv are the noise factors respectively, and the policy gradient is expressed in the following form:
[0171]
[0172] According to the above formula (37) and introducing the idea of the clipping surrogate function, the loss function of the policy network is expressed in the following form:
[0173]
[0174] In the formula, r(θ) = π k (a k (t)|θ k , o k (t)) / π k (a k (t)|θ′ k , o k (t)) is the importance weight, ε is a hyperparameter, S represents the entropy of the policy, δ is a hyperparameter that controls the entropy coefficient, and clip(·) is the clipping function, that is:
[0175]
[0176] In the formula, x represents the variable;
[0177] The loss function of the critic network is in the following form:
[0178]
[0179] In the formula, represents the Huber function with parameter δ H where R represents the discounted reward for the user to interact with the environment until the end of a game, and value clipped is expressed as follows:
[0180]
[0181] In the formula, respectively represent the state value functions output by the comment network before and after update, and ε1 is the parameter of the clipping function.
[0182] The training process of the designed LSTM-NOISEMAPPO algorithm is shown in Algorithm 1.
[0183]
[0184]
[0185] In one embodiment, a semantic communication system resource allocation system for an image restoration task is provided. The system includes:
[0186] A first module for establishing a semantic communication system, including a base station and a number of users randomly distributed within the signal coverage area of the base station; a neural network for extracting image semantic information is stored on each user, and an edge computing node is installed at the base station, and a neural network for restoring images using semantic information is stored at this node;
[0187] A second module for the user to send pilot signals to the base station, and the different pilot signals are orthogonal to each other;
[0188] A third module for the base station to perform channel estimation on the uplink channel of each user based on the pilot signals sent by all users;
[0189] A fourth module for theoretically analyzing the signal-to-interference-plus-noise ratio (SINR) and the statistical characteristics of the SINR of the transmitted signal at the receiving end based on the uplink channel estimation results;
[0190] A fifth module for analyzing the transmission energy consumption of the user sending semantic information using the SINR of each user, and calculating the theoretical average symbol error performance of the transmitted signal of each user according to the modulation method;
[0191] The sixth module is used to obtain the mapping curve of the number of CPU cycles required for the user to extract semantic information and the base station to restore image information under different semantic extraction rates, perform numerical fitting on the mapping curve, and calculate the computational energy consumption for extracting semantic information and restoring image information based on the fitting result;
[0192] The seventh module is used to construct an optimization problem and simplify the problem to obtain the final semantic communication resource allocation problem that needs to be optimized;
[0193] The eighth module is used to design and build a new deep reinforcement learning algorithm based on the multi-agent proximal policy optimization algorithm to solve the semantic communication resource allocation problem. During the training process, an offline training of the user is carried out by constructing a dynamically changing training environment by updating the user's location and channel state in each interaction round until the training result converges.
[0194] For the specific limitations of the semantic communication system resource allocation system for the image restoration task, reference can be made to the limitations of the semantic communication system resource allocation method for the image restoration task in the above text, which will not be elaborated here. Each module in the above semantic communication system resource allocation system for the image restoration task can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0195] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the semantic communication system resource allocation method for the image restoration task are implemented.
[0196] For the specific limitations of each step, reference can be made to the limitations of the semantic communication system resource allocation method for the image restoration task in the above text, which will not be elaborated here.
[0197] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the semantic communication system resource allocation method for the image restoration task are implemented.
[0198] For the specific limitations of each step, reference can be made to the limitations of the semantic communication system resource allocation method for the image restoration task in the above text, which will not be elaborated here.
[0199] As a specific example, in one of the embodiments, the present invention is further verified and explained.
[0200] The computer graphics card configuration used in the simulation of this embodiment is GeForce RTX 4090 GPU, the CPU configuration is i9-14900K, the RAM configuration is 64GB, the software configuration is Python 3.12, and Pytorch 1.12.1. The large-scale fading model is adopted. , where r k represents the actual distance from the k-th user to the base station, r h represents the minimum distance between all users and the base station, and v represents the path loss coefficient. satisfies where let The modulation method of the user's transmitted signal selects 16-QAM. The size of the original image data to be transmitted by each user in each training round is obtained by referring to the image size of the MNIST dataset. Task k (t) is uniformly distributed between [6, 13] KB. The number of rounds per game in the training process, episode_length, is set to 100, the number of training games, episodes, is set to 45000, the dimension of the hidden state of the user policy network is set to 64, the sequence length of the observed state input is set to 1, and the learning rate for updating the parameters of the policy network and the evaluation network is set to 5×10 -4 , the clip range ε in the clip function clip(g) is set to 0.2, and the parameter δ of the Huber function is H set to 10. In the global state s(t), the parameter a of the scaling factor ε follows a uniform distribution U(0.95, 0.99), and b follows a uniform distribution U(1.01, 1.05). The parameters x in and the noise factors α, n in the advantage function adv are set using existing methods. The weight coefficients λ1, λ2, λ3, λ4 in the global reward r(t) are set to 1, 0.5, 0.1, and 0.1 respectively. The remaining simulation parameters are shown in Table 1.
[0201] Table 1 Simulation Parameters
[0202]
[0203]
[0204] To verify the good performance of the proposed algorithm demonstrated in the simulation, three algorithms (MAPPO algorithm, NOISE-MAPPO algorithm, and LSTM-MAPPO algorithm) are adopted as the comparison algorithms for the proposed LSTM-NOISE MAPPO algorithm in this embodiment.
[0205] Figure 4This is a graph showing the relationship between the average energy consumption of all users in each game under different algorithms and the number of training sessions in each training round. By observing the four simulation results, three phenomena can be found:
[0206] First, the energy consumption of the proposed algorithm and the comparison algorithms after convergence is almost in the same range. Taking the average of the data after training convergence, the mean values of the four algorithms after stabilization are 12.2038 mJ (proposed algorithm), 11.8784 mJ (MAPPO), 11.6352 mJ (LSTM MAPPO), and 12.1100 mJ (NOISE MAPPO) respectively. By comparing the data, it is found that the training effects of the proposed algorithm and the comparison algorithms have slightly decreased but the difference is not significant. It can be considered that the proposed algorithm has almost the same training effect as the comparison algorithms.
[0207] Second, there is an obvious large-scale fluctuation phenomenon in the energy consumption values of the MAPPO algorithm and the NOISE MAPPO algorithm during the initial training process (referring to episodes ≤ 10 4 ), while this phenomenon does not occur in the proposed algorithm and the LSTM MAPPO algorithm during the training process. Further analysis of the training data shows that when the proposed algorithm reaches the same variance level of the training results as the MAPPO algorithm and the NOISE MAPPO algorithm (episodes ≥ 10 4 ), the number of training sessions can be reduced by 80% (compared with MAPPO), 66.7% (compared with LSTM MAPPO), and 80% (compared with NOISE MAPPO) respectively. The results show that the training convergence speed of the proposed algorithm is significantly faster.
[0208] Third, taking the variance of the training results of the entire training process, the corresponding training result variance values of the four algorithms are 2.4597 (proposed algorithm), 4.4169 (MAPPO), 2.8430 (LSTM MAPPO), and 3.2299 (NOISE MAPPO). This shows that the proposed algorithm is more stable overall compared with all comparison algorithms, and the stability has increased by 44.31% (compared with MAPPO), 13.48% (compared with LSTM MAPPO), and 23.85% (compared with NOISE MAPPO) respectively. And in the middle and late stages of training (referring to episodes ≥ 3×10 4) The energy consumption of the proposed algorithm fluctuates even less compared to the comparison algorithms. The variance values of the corresponding training results of the four algorithms are 1.0187 (proposed algorithm), 2.0465 (MAPPO), 2.1683 (LSTM MAPPO), and 2.0850 (NOISE MAPPO), indicating that the proposed algorithm can make the network energy consumption more stable in the middle and late stages of training, with the stability improved by 50.22% (compared to MAPPO), 53.02% (compared to LSTM MAPPO), and 51.14% (compared to NOISE MAPPO) respectively.
[0209] Figure 5 Figure showing the relationship between the average delay per user per game in each training episode and the number of training games for different algorithms. Similarly, it can be seen that the proposed algorithm makes the delay performance of the system significantly more stable than the comparison algorithms in the middle and late stages of training (episodes ≥ 3×10 4 ) The proposed algorithm can make the delay performance of the system significantly more stable than the comparison algorithms in the middle and late stages of training. Analyzing the delay data of the system under the four algorithms, it can be found that the delay stability brought by the proposed algorithm to the network in the middle and late stages of training is 96.6% (compared to MAPPO), 96.7% (compared to LSTM MAPPO), and 96.8% (compared to NOISE MAPPO) higher than that of the comparison algorithms respectively. At the same time, the delay performance of the four algorithms after training stabilization is almost the same. In addition, the MAPPO algorithm and the NOISE MAPPO algorithm again show the phenomenon of large fluctuations in training effects in terms of delay performance. The proposed algorithm can achieve the same training effect as the MAPPO algorithm and the NOISE MAPPO algorithm when episodes = 10 3 while the MAPPO algorithm and the NOISE MAPPO algorithm require episodes ≥ 10 4 This is approximately consistent with the comparison results of energy consumption performance.
[0210] Figure 6 Figure showing the relationship between the average link rate per user per game in each training episode and the number of training games for different algorithms. Analyzing the data, the average link rate per round of all users in the network under the four algorithms are 13.627 Mb / s (proposed algorithm), 13.635 Mb / s (MAPPO), 13.64 Mb / s (LSTM MAPPO), and 13.63 Mb / s (NOISE MAPPO) respectively, and the variances are 6.0294×10 11 (proposed algorithm), 6.0134×10 11 (MAPPO), 6.0378×10 11 (LSTM MAPPO), 6.0435×10 11(NOISE MAPPO), the maximum and minimum values are 16.953 Mb / s and 10.645 Mb / s (the proposed algorithm), 17.32 Mb / s and 10.791 Mb / s (MAPPO), 17.178 Mb / s and 10.881 Mb / s (LSTM MAPPO), 16.927 Mb / s and 10.647 Mb / s (NOISE MAPPO). Through the comparative analysis of the above three indicators, it can be known that there is almost no difference in the performance of the proposed algorithm and the comparative algorithms in terms of network link performance.
[0211] Table 2 shows the mean and variance of the maximum average symbol error rate per round for all users. By comparing the data, it is found that the proposed algorithm also has a small performance improvement in both the mean and variance of the symbol error performance compared with the comparative algorithms. The improvement amplitudes of the mean are 4% (compared with MAPPO), 10.12% (compared with LSTM MAPPO), and 16.33% (compared with NOISE MAPPO) respectively. By comparing the variances, the improvement amplitudes of the stability are 4.11% (compared with MAPPO), 16.48% (compared with LSTM MAPPO), and 9.75% (compared with NOISE MAPPO) respectively.
[0212] Comparison of the network symbol error performance between the proposed algorithm and the comparative algorithms in Table 2
[0213]
[0214]
[0215] Through the performance analysis and comparison in the above four aspects, it can be seen that the proposed algorithm can fully reach the level of the comparative algorithms in terms of network energy consumption, delay, link rate, and symbol error performance, and has a significantly faster convergence speed and better convergence stability compared with the comparative algorithms. At the same time, it also shows a small performance improvement in the symbol error performance.
[0216] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A resource allocation method for a semantic communication system for image restoration tasks, characterized in that, The method includes the following steps: Step 1: Establish a semantic communication system, which includes a base station and several users randomly distributed within the signal coverage area of the base station; a neural network for extracting image semantic information is stored on each user, and an edge computing node is installed at the base station, and a neural network for restoring images using semantic information is stored on this node. Step 2: The users send pilot signals to the base station, and the different pilot signals are orthogonal to each other. Step 3: The base station estimates the uplink channel of each user based on the pilot signals sent by all users. Step 4: Based on the uplink channel estimation results, conduct a theoretical analysis of the signal-to-interference-plus-noise ratio (SINR) and the statistical characteristics of the SINR at the receiving end of the transmitted signal. Step 5: Analyze the transmission energy consumption of the users sending semantic information using the SINR of each user, and calculate the theoretical average symbol error performance of the transmitted signals of each user according to the modulation method. Step 6: Obtain the mapping curve of the number of CPU cycles required for users to extract semantic information and the base station to restore image information under different semantic extraction rates, perform numerical fitting on the mapping curve, and calculate the computational energy consumption for extracting semantic information and restoring image information based on the fitting results. Step 7: Construct an optimization problem, and simplify this problem to obtain the final semantic communication resource allocation problem that needs to be optimized; the optimization problem is: with the theoretical average symbol error performance of the transmitted signals of the users obtained in Step 5 as the constraint, minimize the sum of the computational energy consumption of all users. Step 8: Design and build a new deep reinforcement learning algorithm based on the multi-agent proximal policy optimization algorithm to solve the semantic communication resource allocation problem. During the training process, construct a dynamically changing training environment by updating the positions and channel states of the users in each interaction round to perform offline training on the users until the training results converge; the new deep reinforcement learning algorithm based on the multi-agent proximal policy optimization algorithm specifically includes: designing a long short-term memory neural network that can extract time series features to replace the MLP network structure in the policy network of the MAPPO algorithm; adding noise during the output process of the evaluation function.
2. The resource allocation method for the semantic communication system for image restoration tasks according to claim 1, wherein Step 1 further includes: Before performing the semantic communication task, train the semantic transceiver networks of both the semantic communication transmitter and receiver. After the training is completed, obtain the transceiver network parameters corresponding to specific semantic background knowledge; the semantic transceiver network includes the neural network for extracting image semantic information at the user side and the neural network for restoring images using semantic information at the edge computing node.
3. The resource allocation method of the semantic communication system for image restoration tasks according to claim 1, characterized in that In step 2, at the beginning of each channel coherence interval, M single-antenna users in the cell simultaneously send mutually orthogonal pilot signals to the base station, and the number of antennas at the base station is denoted as N r , the number of pilot signals is M, and the pilot dimension size is 1×τ p , the pilot vector sent by the k-th user is denoted as Φ k , and its elements are denoted as satisfying τ p is greater than or equal to M, and the transmission power of the pilot signal is the maximum transmission power of the user, denoted as p k .
4. The resource allocation method for the semantic communication system for image restoration tasks according to claim 3, wherein Step 3 specifically includes: Let represent the uplink fading between the N r antennas of the base station in the cell and the k-th user, expressed as: where h k is an N r ×1 vector, and each element follows a complex Gaussian distribution with a mean of 0 and a variance of 1, representing the small-scale fading between the user and each antenna of the base station; β k is a scalar representing the large-scale fading between the k-th user and the base station; After all users send pilot signals, the received signal at the base station side is represented as Y: Wherein, the dimension of Y is N r ×M, M is the number of pilot signals, and n is a complex additive Gaussian noise matrix of dimension N r ×M, the elements of which satisfy independent and identically distributed CN(0, σ 2 ), CN(0, σ 2 ) represents a complex Gaussian distribution with a mean of 0 and a covariance of σ 2 ; obtained from the orthogonality between different pilot vectors, Φ j represents the pilot sent by the jth user, and H represents the conjugate transpose; Let Right-multiply both sides of Equation (2), and we get: In the formula, The dimension size is N r ×1; Based on Equation (3), the MMSE method is used for channel estimation to obtain the uplink channel of the k-th user as 5. The resource allocation method for the semantic communication system for image restoration tasks according to claim 4, characterized in that Step 4 specifically includes: (1) After channel estimation, the received signal of the base station during the communication process is represented as y: where x k represents the transmission signal of the k-th user, and p′ k represents the signal transmission power of the k-th user, For any k-th user, at the base station side, the received signal is linearly combined by calculating the inner product of the weight vector v of dimension 1×N r and the received signal y, which is expressed as: k where p′ k' represents the signal transmission power of the k'-th user, x k' represents the transmitted signal of the k'-th user, represents the channel estimation value of the k'-th user; Based on Equation (6), the maximum SINR of the k-th user is expressed as γ k : wherein, is an identity matrix of dimension N r ×N r , C k′ is the mean square value of the channel estimation error of the k'-th user, expressed as: wherein, β k' is a scalar representing the large-scale fading between the k'-th user and the base station; (2) Analyze the probability density function of the SINR of any user's uplink channel through an asymptotic method Let the matrix have eigenvalues converging to λ1. According to the Marchenko-Pastur distribution criterion, the mean of λ1 is calculated as follows: Thus, Equation (7) is simplified to: wherein, represents the j-th element of the channel estimation vector of the k-th user, and the expression of φ is as follows: The statistical distribution of SINR in Equation (10) is characterized as the convolution of the two distributions on the right-hand side of the equation, which is expressed as: In the formula, N1 = M - 1 N2 = N r -M + 1 6. The resource allocation method for the semantic communication system for image restoration tasks according to claim 4, characterized in that Step 5 specifically includes: For various digital modulation systems, its average symbol error rate is represented by the Gaussian Q function: Among them, the Gaussian Q function is expressed as: The parameters a and b are constants determined by a specific modulation scheme; Substituting formula (14) into formula (13), we get: Among them, the closed - form expression of C(x) is as follows: Among them, Substitute C(x) into Equation (15) to obtain the expression of.
7. The resource allocation method for the semantic communication system for image restoration tasks according to claim 1, characterized in that Step 6 specifically includes: For each user, its energy consumption consists of two parts: the energy consumption \(E\) for extracting semantic information 1k and the energy consumption \(E\) for uplink transmission 2k ; (1) The energy consumption E for extracting semantic information 1k is as follows: E 1k = κy 1k (D k ,S k )f 2 (17) where y 1k (D k , S k ) represents the number of CPU cycles for the k-th user in the cell to extract semantic information, D k and S k respectively represent the data volume sizes before and after the k-th user extracts semantics, related to the semantic extraction rate ρ k ; f represents the computing capacity of the k-th user; κ represents the energy consumption coefficient; Obtain y by means of numerical fitting 1k (D k ,S k ) and ρ k The relationship between them is expressed as follows: y 1k (D k ,S k ) = η1ρ k +η2 (18) In the formula, η1>0 and η2>0 are constant parameters respectively; The time for the user to extract semantic information is: (2) The energy consumption E of the uplink transmission 2k is as follows: E 2k = t 2k p' k (20) Among them, the time required for uplink transmission is t 2k : where, Cap k represents the uplink channel capacity of the k-th user, and its expression is Cap k = Wlog2(1 + γ k ), where W is the uplink bandwidth allocated to each user, and S k is the size of the data volume after the k-th user extracts semantic information; (3) For the edge computing node, the energy consumption is the computing energy consumption when restoring the semantic information sent by the users in the cell, which is expressed as: where E 3k represents the energy consumption for decoding the information sent by the k-th user, represents the number of CPU cycles required for decoding the information sent by the k-th user, and respectively represent the size of the semantic information data volume after wireless channel transmission and the size of the data volume after semantic decoding. Since the neural network structures in the semantic encoding and semantic decoding processes are symmetric, it is considered that f 3k represents the computing capacity allocated to the k-th user; The corresponding time consumption is: In summary, the total energy consumption \(E\) of the \(k\)-th user for performing its semantic communication task k is expressed as: \(E\) k = \(E\) 1k + \(E\) 2k + \(E\) 3k .
8. The resource allocation method for the semantic communication system for image restoration tasks according to claim 7, wherein The optimization problem constructed in Step 7 is: Among them, the constraint condition C1 is the delay constraint for each user in the cell to perform semantic communication tasks, and t th represents the upper limit value of the delay constraint for each user, C2 represents the accuracy constraint for each user to perform semantic tasks, and u k (D k ,S k ) represents the semantic accuracy of the k-th user for semantic communication, and A k represents the lower limit value of the semantic accuracy constraint for the k-th user. C3 and C4 are the computing capacity constraints of the edge computing node, and F max represents the maximum computing capacity of the edge computing node. C5 is the signal transmission power constraint for each user, and p level1 , p level2 , p level3 , p level4 respectively represent the optional power. C6 is the link symbol error performance constraint for each user, is the upper limit value of the symbol error performance constraint for each user, C7 represents the link rate sum constraint for all users in the cell, and Γ th represents the lower limit value of the link rate sum constraint, The optimization problem P1 is equivalent to: where C2 constraint is the constraint condition for the semantic extraction rate of each user, and ρ th represents the lower limit value of the semantic extraction rate constraint for each user.
9. The resource allocation method for the semantic communication system for image restoration tasks according to claim 8, characterized in that, Step 8 specifically includes: For the optimization problem P2, the overall idea of the solution is: decompose the problem into two sub - problems, and by alternately solving the two sub - problems in turn, make P2 converge to the optimal value; First, fix the variables ρ k and p′ k When this is done, P2 is transformed into a sub-problem with only the single variable f 3k as follows: P3 is a convex optimization problem, which is solved using the convex optimization tool CVXPY; Regarding each user as an agent in deep reinforcement learning, under the condition that information can be shared among all users, the observed state of the k - th user in the t - th training episode is defined as Among them, Task k (t) represents the size of the original data that the k-th user needs to transmit in the t-th training round. Energy(t) represents the total energy consumption of all users in the (t - 1)-th training round. Max_Delay(t) represents the maximum value of the delay experienced by all users in the (t - 1)-th training round. Max_SER(t) represents the maximum value of the average symbol error rate of all user links in the (t - 1)-th training round. Define the action of the k-th user as Among them, p k (t) represents the power selection scheme of the user in the t-th training round, and ρ k (t) represents the semantic extraction rate selection scheme of the user in the t-th training round. Action - (t) represents the set of actions taken by other users except the user himself in the (t - 1)-th training round, that is, Action - (t) = {a1(t) ∪ a2(t)... a k-1 (t) ∪ a k+1 (t)... ∪ a M (t)}; The training adopts a centralized training method, and the global state is defined as the set of all user states, that is The reward value of each user in the training episode is expressed as: The second, third, and fourth terms on the right side of the r(t) equation are penalty terms, and λ1, λ2, λ3, and λ4 are the weight coefficients of the corresponding terms respectively. All users share the same reward value. In each round, each user uses its own policy network to output an action according to its respective observed state and uploads its own state to the base station. At the base station, the edge computing node calculates r(t) and uses the evaluation network to output the global state value function V ζ (s(t)), where ζ represents the parameters of the evaluation network.
10. The resource allocation method for the semantic communication system for image restoration tasks according to claim 9, wherein The new deep reinforcement learning algorithm based on the multi - agent proximal policy optimization algorithm in Step 8 is an improvement on the multi - agent proximal policy optimization algorithm MAPPO, which specifically includes: (1) Regarding the set of state values of all training episodes of each user as a time series, designing the policy network of each user as a long - short - term memory neural network capable of extracting time - series features to replace the MLP network structure in the policy network of the MAPPO algorithm. In each training episode, the policy of the user is obtained by calculating the output of the user's observed state through the LSTM network; (2) Add noise to the global state s(t) during the output process of the evaluation function, i.e., the global state value function V ζ (s(t)). The global state with added noise is represented as follows: as follows: In the formula, the concat(·) function represents the concatenation function, and its function is to horizontally concatenate s(t) and the noise matrix x(t). is the global state after adding noise, and each element in x represents a normal distribution with a mean of 0 and a variance of , ε is a scaling factor, ε: U(a, b), and U(a, b) represents a uniform distribution with a numerical range between a and b; in addition, a noise factor is added during the calculation of the advantage function. At the t-th round of the user offline training phase, the policy network of the k-th user selects an appropriate action a k (t) based on the observed state o k (t) of this user, and the action output process is expressed as follows: Forget gate: determines which information in the cell state at the previous moment needs to be forgotten; the activation value f of the forget gate t is calculated from the current input o k (t) and the previous hidden state h' t-1 through a fully connected layer with a sigmoid activation function, where W if , W hf , b if , and b hf are the parameters of the fully connected layer: f t = σ(W if o k (t) + b if + W hf h’ t-1 + b hf ) (28) Input gate: Determines which information in the current input should be added to the cell state, and consists of two parts: one is to determine the acceptance weight i of the information through the sigmoid function t , and the other is to calculate the candidate state g through the tanh function t , W ii , W hi , b ii , b hi , W ig , W hg , b ig , b hg are the corresponding fully connected layer parameters respectively: i t = σ(W ii o k (t) + b ii + W hi h’ t-1 + b hi ) (29) g t = tanh(W ig o k (t) + b ig + W hg h’ t-1 + b hg ) (30) Cell state update: Combine the results of the forget gate and the input gate to update the cell state c t ': c’ t = f t * c t-1 + i t * g t (31) In the formula, * represents the bitwise multiplication operation, and c t-1 represents the cell state at the previous moment; Output gate: determines which information in the cell state should be passed to the hidden state at the next time step or be used as the model output at the current time step; the activation value o of the output gate t is calculated from the current input o k (t) and the previous hidden state h′ t-1 through a fully connected layer with a sigmoid activation function, where W io , W ho , b io , and b ho are the parameters of the fully connected layer, and then the element-wise product with the value of the cell state after passing through the tanh function gives the final hidden state h t ': o t = σ(W io o k (t) + b io + W ho h’ t-1 + b ho ) (32) h’ t = o t *tanh(c’ t ) (33) Output the result o of the gate t After passing through the fully connected layer with the ReLU function, the action output is obtained by sampling through the Categorical function; For the k-th user, let the parameters of its target policy network be θ k , and the parameters of the behavioral policy network be θ k '. The network structures of the target policy and the behavioral policy are the same, but the parameters are different. The policy probability value generated by the target policy network is expressed as π k (a k (t)|θ k , o k (t)), and the policy probability value generated by the behavioral policy network is expressed as π k (a k (t)|θ k ', o k (t)). After the user interacts with the environment, r(t), V ζ (s(t)), a k (t), π k (a k (t)|θ k , o k (t)) and o k (t + 1) generated by this step of interaction will be stored in the cache. When entering the network parameter update phase, mini-batch copies of the cache data are retrieved from the cache for parameter update. In the LSTM-NOISEMAPPO algorithm, the policy objective function for each user is expressed as: where ι(·) represents the observed state o of user k k (t) follows the distribution, V k (·) represents the state value function of the k-th user, where the gradient with respect to the parameter θ k is expressed as: Where Q k (o k (t), a k (t)) represents the action value function of user k when the observation state is o k (t) and the action is a k (t); Approximating the policy gradient by calculating an unbiased estimate of formula (35), as shown in formula (36): Let denote the advantage function, where α and n adv are noise factors respectively, and the policy gradient is expressed in the following form: According to the above formula (37) and introducing the idea of the clipped surrogate function, the loss function of the policy network is expressed in the following form: where \(r(\theta)=\pi\) k (a k (t)|\theta k ,o k (t)) / \pi k (a k (t)|\theta' k ,o k (t)) is the importance weight, \(\varepsilon\) is a hyperparameter, \(S\) represents the entropy of the policy, \(\delta\) is a hyperparameter that controls the entropy coefficient, and \(clip(\cdot)\) is the clipping function, i.e.: In the formula, x represents a variable; The loss function of the critic network is in the following form: In the formula, represents the Huber function with parameter δ H , R represents the discounted reward for the user to interact with the environment until the end of a game, and value clipped is expressed as follows: wherein, respectively represent the state value functions of the comment network output before and after the update, and ε1 is the parameter of the clipping function.
Citation Information
Patent Citations
Edge computing unloading and resource allocation method based on multi-agent reinforcement learning
CN116321293A
Semantic communication system and user association and resource allocation joint optimization method in semantic communication system
CN117915484A