Resource allocation method for semantic communication system based on deep reinforcement learning

By employing a resource allocation method based on deep reinforcement learning for semantic communication systems, the power and user access of multi-cell semantic communication systems are optimized, solving the resource allocation problem of multi-cell semantic communication systems and achieving maximum system throughput and intelligent resource allocation.

CN116390238BActive Publication Date: 2026-03-20NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively solve the power allocation and user access problems in multi-cell semantic communication systems, and have not introduced deep reinforcement learning into wireless resource optimization.

Method used

A resource allocation method for semantic communication systems based on deep reinforcement learning is adopted. Data is collected through a multi-cell intelligent semantic processing center. The Transformer model and knowledge base KB are used to assist the system. Combined with DDQN & DDPG neural networks, sub-channel and power allocation are optimized, and a Semantic-JCP network is designed for resource allocation.

Benefits of technology

It maximizes the semantic throughput of multi-cell systems, avoids the problems of non-intelligent operation and slow computation of traditional methods, saves the overhead of channel state information acquisition, and provides a near-optimal resource allocation scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116390238B_ABST
    Figure CN116390238B_ABST
Patent Text Reader

Abstract

The application discloses a resource allocation method of a semantic communication system based on deep reinforcement learning. A multi-cell intelligent semantic processing center collects sending data and receiving data of communication through each base station, obtains a mapping table of BLEU scores of the semantic communication system changing with signal-to-noise ratio by using a Transformer model, and obtains a bit-to-semantic (B2M) conversion curve by fitting the mapping curve, so as to obtain semantic throughput of the system through a bit rate; an auxiliary system of a knowledge base is introduced; a multi-cell network user sub-channel allocation mode and a power allocation mode are determined, transmission bit rates of each cell are obtained, semantic throughput of a single cell is obtained through the B2M conversion curve; an optimization problem is obtained, a final semantic resource allocation scheme to be optimized is obtained by simplifying the problem; a resource allocation network Semantic-JCP is designed and built, and sub-channel allocation and power allocation under the condition that the semantic throughput of the multi-cell system is maximum are finally obtained after solving the resource allocation problem.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a resource allocation method of a semantic communication system based on deep reinforcement learning and belongs to the technical field of communication. BACKGROUND

[0002] In recent years, semantic communication has attracted extensive attention in the industry and academia and is one of the core challenges of the sixth generation (6G) wireless network. The data compression in semantic communication and the reliable semantic communication model are explored, and attempts are made to associate semantic communication with Shannon information theory. Considering that the real intention of each intelligent terminal in semantic communication cannot be determined, the semantic communication problem is described as a Bayesian game problem, and the most transmission strategy capable of minimizing the end-to-end average semantic error is described. Semantic communication only transmits the necessary information related to a specific task at the receiving end, thereby realizing a truly intelligent system and significantly reducing data traffic. This communication framework will be widely applied in the fields of industrial internet, intelligent transportation, video conferencing, online education, augmented reality (AR) and virtual reality (VR), and is expected to form a more sustainable communication network.

[0003] Significant progress has been made in understanding the mathematical basis of symbol transmission since Shannon. With the development of neural networks, the latest progress in deep learning and its applications provides important insights for the development of semantic communication. The first rigorous and comprehensive vision of an end-to-end semantic communication network based on new concepts of artificial intelligence (AI), causal reasoning, transfer learning and minimum description length theory is proposed.

[0004] So far, no research work has introduced deep reinforcement learning into the wireless resource optimization of a semantic communication system, and no consideration has been given to the multi-cell power allocation and user connection problem based on semantic information. SUMMARY

[0005] The application aims to provide a resource allocation method of a semantic communication system based on deep reinforcement learning to solve the power allocation and user access problem of a multi-cell semantic communication system.

[0006] To achieve the above-mentioned purpose, the application provides a resource allocation method of a semantic communication system based on deep reinforcement learning, mainly comprising the following steps:

[0007] S1, a multi-cell intelligent semantic processing center collects the sending data and receiving data of communication through each base station, obtains a mapping table of the BLEU score of the semantic communication system varying with the signal-to-noise ratio by using a Transformer model, and obtains a bit-to-semantic B2M conversion curve by using mapping curve fitting, so as to obtain the system semantic throughput T(η,c t ,p t), wherein η represents a knowledge matching coefficient, and the greater the value, the higher the degree of knowledge matching between the user and the base station, c t and p t respectively represent a subchannel allocation variable and a power allocation variable of time slot t;

[0008] S2, introduce the auxiliary system of the knowledge base KB, provide the background knowledge of a specific application field for the user, and use it for semantic reconstruction;

[0009] S3, determine the subchannel allocation mode and the power allocation mode of the user of the multi-cell network, and obtain the transmission bit rate X n (c t ,p t ) of each cell through the B2M conversion curve to obtain the semantic throughput T n (η,c t ,p t ) of a single cell, wherein n represents the cell index;

[0010] S4, obtain the optimization problem through the system judgment, and obtain the final semantic resource allocation scheme to be optimized through the simplification of the problem;

[0011] S5, design and build the resource allocation network Semantic-JCP of the multi-cell semantic communication system based on DDQN&DDPG, and finally obtain the subchannel allocation and power allocation under the condition that the semantic throughput of the multi-cell system is maximized after solving the resource allocation problem

[0012]

[0013] As a further improvement of the application, step S1 specifically comprises the following steps:

[0014] S11, for all users m in the semantic communication network, define the bit-message B2M function under the current channel condition as S m (·), which is used to convert the bit rate in Shannon's law into the message rate, and the performance of the function is determined by the message attribute, the semantic encoder and decoder, and the knowledge matching degree of the transmitting and receiving parties. The complete knowledge matching function between the base station and the user is introduced

[0015]

[0016] wherein 0<η m <1, and it is assumed that all η m are independently and identically distributed, satisfying wherein τ represents the mean value, and σ 2 is the variance;

[0017] ​S12. Obtain the message transmission rate that can be achieved by base station n and user m through subchannel l in the t-th time slot:

[0018]

[0019] in, This indicates that a resource allocation strategy c is used on this channel. t and p t If the bit transmission rate achievable in the nth cell is given by the following, then the total message transmission rate achievable in the nth cell is:

[0020]

[0021] S13. Using the Transformer model as the semantic conversion model, since the channel conditions are constantly changing, the BLEU score is used to evaluate the B2M conversion accuracy of the model under different signal-to-interference-plus-noise ratios (SINR).

[0022] As a further improvement of the present invention, in step S2, an auxiliary system of knowledge base KB is introduced into the user and the semantic processing center to provide the user with background knowledge in specific application fields of medicine, communication and chemistry related to the type of communication user for semantic reconstruction; the knowledge base matching coefficient is a self-defined number between 0 and 1, the larger the value, the more matched the relevant knowledge base, and the higher the semantic transmission rate under the same conditions.

[0023] As a further improvement of the present invention, step S3 specifically includes the following steps:

[0024] S31. A multi-cell intelligent semantic communication network consists of N base stations, each semantic communication base station being located at the center of its respective semantic communication cell. The set of N semantic communication base stations is represented as follows: Let each base station have L sub-channels, denoted as φ = {1, 2, ..., L}. There are M users in the multi-cell semantic communication network, and these users are randomly distributed in each cell. The set of users served by semantic communication base station j is denoted as ψ = {1, 2, ..., M}, n ∈ j. Each user is equipped with an antenna for transmitting and receiving semantic data, that is, the total number of antennas of the semantic base station is equal to the number of users. The total bandwidth of the multi-cell semantic communication system is B, and the bandwidth of each semantic communication sub-channel is w = B / L.

[0025] S32. Set the frequency reuse coefficient of the multi-cell semantic communication network to f = 1, and represent the sub-channel allocation matrix as C = [C1, C2, ..., C...]. N ],in The connection status of the m-th user and the n-th base station in subchannel l at time slot t is represented as follows:

[0026]

[0027] In a single time slot, each user in the system is allowed to use only one subchannel, and in order to avoid mutual interference between subchannels within a cell, the subchannels in a single cell are set to be mutually orthogonal, which is expressed as

[0028]

[0029] S33, define the large-scale fading at time slot t as The small-scale fading is Therefore, in a time slot, the large-scale fading is constant and the small-scale fading is subject to Gaussian distribution

[0030] S34, at this time, the channel gain matrix is where the element is Define the set of transmission powers of the base stations as P = [P1, P2, …, P N ], where Pn represents the transmission power vector of the nth base station, and the total transmission power of the base station n is limited to between 0 and P max ;

[0031] S35, define the received SINR of user m on the lth subchannel of the nth base station as:

[0032]

[0033] where, N0 is the noise power, is the interference from the useless signal, and therefore in the tth time slot, the transmission bit rate that can be achieved by the base station n and user m through the lth subchannel is:

[0034]

[0035] where Finally, the transmission bit rate of each cell can be obtained as:

[0036]

[0037] As a further improvement of the present application, step S4 specifically comprises the following steps:

[0038] S41, the system target is to adjust the power allocation and subchannel allocation scheme of each cell, so the optimization problem can be expressed as:

[0039]

[0040]

[0041]

[0042] This optimization problem is a multi-objective and non-convex programming problem, and η m The randomness of makes the problem a non-deterministic problem;

[0043] S42, since the random variable in P1 only exists in the objective, a new objective function is introduced and an additional constraint is added, so that P1 is equivalent to:

[0044]

[0045]

[0046]

[0047]

[0048] Where, Φ ―1 (α) is the inverse function of the standard normal distribution,

[0049]

[0050] As a further improvement of the present application, step S5 specifically comprises the following steps:

[0051] S51, define the semantic channel state of the agent as S, the algorithm estimates the semantic channel state value with the function V, estimates the semantic resource allocation action value with the advantage function A, and calculates the two functions with the adversarial neural network, when the input semantic channel state, one network outputs the value function V(s,β), wherein β is the network parameter, and the other network outputs the semantic action advantage function A(s i ,c i ,α), wherein α is the network parameter, therefore the Q-value function of the network is defined as

[0052] Q(s i ,c i ;α,β)=V(s i ,β)+A(s i ,c i ;α);

[0053] S52, in order to obtain a unique solution of the Q-value function, the semantic value Q function is converted to:

[0054]

[0055] S53, in order to make the network semantic estimation Q value more accurate, another deep neural network is introduced to calculate the semantic target Q value Q target (t), at time t, each agent shares a Qtarget (t) to reduce the training cost.

[0056] As a further improvement of the present application, step S5 further comprises the following steps:

[0057] S54, in the network training phase, the current channel gain is taken as the input of each cell network, and the semantic actor network selects a suitable semantic action A based on the state, and the policy set of all semantic agents is defined as μ = { μ 1, …, μ N}, θ = { θ 1, …, θ N} are the corresponding semantic resource allocation policy parameter sets, representing the optimal semantic resource allocation policy of agent i;

[0058] S55, the semantic critic network calculates the Q value of the current network semantic action A based on the current semantic channel state of all agents and its corresponding semantic action A, and the cumulative expected semantic reward of semantic agent i is

[0059]

[0060] wherein, represents expectation, and γ represents long-term discount rate;

[0061] S56, the target semantic function is defined as gradient

[0062]

[0063] wherein, the semantic state The semantic communication critic network is trained by estimating the semantic Q value and the target semantic Q value, which is expressed as the loss value in the following formula:

[0064]

[0065] wherein

[0066] is the target network semantic parameter corresponding semantic policy, and then the semantic actor network continuously updates its semantic communication resource allocation policy policy through the feedback given by the semantic critic network;

[0067] S57, in the network test phase, the semantic actor network in the multi-cell semantic agent only needs to select a suitable semantic resource allocation action A by analyzing the current local semantic communication channel state information given by the semantic critic network, and the semantic critic network does not need to know the semantic state information and potential semantic resource allocation actions of all agents.

[0068] The beneficial effects of the present application include:

[0069] I. The resource allocation method in the multi-cell semantic communication based on deep reinforcement learning, firstly simplifies the resource allocation problem in semantic communication, and then uses deep learning to allocate sub-channels and power in the multi-cell system. Compared with the traditional non-neural network algorithm, the resource allocation in the multi-cell system can be automatically adjusted to maximize the system semantic throughput.

[0070] II. The resource allocation method in the multi-cell semantic communication based on deep reinforcement learning adopts the algorithm combining DDQN and DDPG neural networks, which can solve both discrete sub-channel allocation schemes and continuous power allocation schemes.

[0071] III. The present application uses semantic throughput to measure semantic transmission rate, avoiding the inaccuracy of traditional bit rate measurement method for semantic information.

[0072] IV. The resource allocation method in the semantic communication based on deep reinforcement learning avoids the non-intelligent, heavy workload and slow calculation of the traditional resource allocation method, and solves the problems of difficult resource allocation and difficult migration in the multi-cell semantic communication system. At the same time, the method can avoid the high dependence of the model-based method on channel knowledge, save the cost of obtaining necessary channel model and its parameters, and realize the approximate optimal resource allocation method without channel state information. BRIEF DESCRIPTION OF DRAWINGS

[0073] Figure 1 is the specific structure process of the multi-cell semantic communication resource allocation system based on deep reinforcement learning of the present application.

[0074] Figure 2 is a schematic diagram of the multi-cell semantic communication system in the embodiment.

[0075] Figure 3 is the BLEU score of the Transformer model under different signal-to-interference ratios.

[0076] Figure 4 is the specific process of the deep reinforcement learning network in the embodiment.

[0077] Figure 5 is the average semantic throughput comparison of the algorithm of the present application and other different algorithms.

[0078] Figure 6 is the semantic throughput comparison after changing the semantic confidence of the present application. DETAILED DESCRIPTION

[0079] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be described in detail below with reference to the drawings and specific embodiments.

[0080] Here, it should be noted that, in order to avoid obscuring the present application due to unnecessary details, only structures and / or processing steps closely related to the solutions of the present application are shown in the drawings, and other details not closely related to the present application are omitted.

[0081] As shown in Figures 1 to 6 , the present application provides a resource allocation method of a semantic communication system based on deep reinforcement learning, mainly including the following steps:

[0082] S1, the multi-cell intelligent semantic processing center collects the sending data and receiving data of communication through each base station, obtains a mapping table of the BLEU score of the semantic communication system varying with the signal-to-noise ratio by using the Transformer model, and obtains a conversion curve of bits to semantics B2M by using the mapping curve fitting, and obtains the system semantic throughput T(η,c t ,p t ) through the bit rate, wherein η represents a knowledge matching coefficient, and the greater the value, the higher the degree of knowledge matching between the user and the base station, c t and p t represent the subchannel allocation variable and the power allocation variable of the time slot t respectively;

[0083] S2, an auxiliary system of a knowledge base KB is introduced to provide the user with background knowledge in a specific application field for semantic reconstruction thereof;

[0084] S3, the subchannel allocation mode and the power allocation mode of the multi-cell network user are determined to obtain the transmission bit rate X n (c t ,p t ) of each cell, and the semantic throughput T n (η,c t ,p t ) of a single cell is obtained through the B2M conversion curve, wherein n represents the cell index;

[0085] S4, an optimization problem is derived by judging the system, and the final semantic resource allocation scheme to be optimized is derived by simplifying the problem;

[0086] S5, a resource allocation network Semantic-JCP of the multi-cell semantic communication system based on DDQN&DDPG is designed and built, and after solving the resource allocation problem, the subchannel allocation and power allocation

[0087]

[0088] Wherein, step S1 specifically includes the following steps:

[0089] S11, for all users m e M in the semantic communication network, define its bit-message (B2M) function under the current channel condition as S m (·), which is used to convert the bit rate in Shannon's law into the message rate, and the performance of the function is determined by the message attribute, semantic codec, and the matching degree of the knowledge of the sender and receiver. In theory, the background knowledge base between the user and the base station can be completely matched, but it is not possible in actual application, so the complete knowledge matching function between the base station and the user is introduced

[0090]

[0091] Wherein, 0 < η m <1, assuming that all η m are independent and identically distributed, satisfying Where τ represents the mean, and σ 2 is the variance;

[0092] S12, get the message transmission rate that the base station n and the user m can achieve through the subchannel l at the tth time slot:

[0093]

[0094] Wherein, represents the bit transmission rate that can be obtained by using resource allocation strategy c t and p t on the channel. Then, the total message transmission rate that the nth cell can achieve is:

[0095]

[0096] S13, use the Transformer model as the semantic conversion model, since the channel condition is constantly changing, use the BLEU score as the evaluation of the B2M conversion accuracy of the model under different signal-to-interference-and-noise ratios SINR.

[0097] In step S2, the knowledge base KB auxiliary system is introduced in the user and the semantic processing center, which provides the user with the background knowledge related to the communication user type in the medical, communication, and chemical specific application fields, which is used for semantic reconstruction; The knowledge base matching coefficient is a number between 0 and 1 defined by oneself, and the larger the value is, the more matching the related knowledge base is, and the larger the semantic transmission rate is under the same condition.

[0098] Step S3 specifically includes the following steps:

[0099] ​S31. A multi-cell intelligent semantic communication network consists of N base stations, each semantic communication base station being located at the center of its respective semantic communication cell. The set of N semantic communication base stations is represented as follows: Let each base station have L sub-channels, denoted as φ = {1, 2, ..., L}. There are M users in the multi-cell semantic communication network, and these users are randomly distributed in each cell. The set of users served by semantic communication base station j is denoted as ψ = {1, 2, ..., M}, n ∈ j. Each user is equipped with an antenna for transmitting and receiving semantic data, that is, the total number of antennas of the semantic base station is equal to the number of users. Assuming that the total bandwidth of the multi-cell semantic communication system is B, then the bandwidth of each semantic communication sub-channel is w = B / L.

[0100] S32. Set the frequency reuse coefficient of the multi-cell semantic communication network to f = 1, and represent the sub-channel allocation matrix as C = [C1, C2, ..., C...]. N ],in The connection status of the m-th user and the n-th base station in subchannel l at time slot t is represented as follows:

[0101]

[0102] Within a single time slot, each user in the system is only allowed to use one sub-channel. Furthermore, to avoid interference between sub-channels within a single cell, the sub-channels within a single cell are set to be mutually orthogonal, as described below.

[0103]

[0104] S33. Define large-scale fading at time slot t as... Small-scale fading Within a time slot, large-scale fading is constant, while small-scale fading follows a Gaussian distribution.

[0105] S34, At this time, the channel gain matrix is: The element is Define the set of base station transmit powers as P = [P1, P2, ..., P...]. N ],in Let Pmax represent the transmit power vector of the nth base station, and let Pmax be the limit of the total transmit power of base station n between 0 and Pmax.

[0106] S35. Define the received SINR of user m on the l-th sub-channel of the n-th base station as:

[0107]

[0108] in, N0 is the noise power. The transmission bit rate that can be achieved by base station n and user m through subchannel l at the tth time slot is:

[0109]

[0110] where The transmission bit rate of each cell can be finally obtained as:

[0111]

[0112]

[0113] Step S4 specifically comprises the following steps:

[0114] S41, the system aims to adjust the power allocation and subchannel allocation scheme of each cell to improve the message transmission rate of the whole system. Therefore, the optimization problem can be expressed as:

[0115]

[0116]

[0117]

[0118] This optimization problem is a multi-objective and non-convex programming problem, and η m The randomness of makes the problem a non-deterministic problem.

[0119] S42, since the random variable in P1 only exists in the objective function, a new objective function is introduced and an additional constraint is added, which can make the solution simple without changing the original intention of the problem. At this time, P1 is equivalent to:

[0120]

[0121]

[0122]

[0123]

[0124] where, Φ ―1 (α) is the inverse function of the standard normal distribution,

[0125]

[0126] Step S5 specifically comprises the following steps:

[0127] S51, define the semantic channel state of the agent as S, the algorithm estimates the semantic channel state value with function V, estimates the semantic resource allocation action value with advantage function A, and calculates the two functions with an adversarial neural network. When the input semantic channel state is input, one network outputs the value function V(s, β), where β is the network parameter, and the other network outputs the semantic action advantage function A(s i i , α), where α is the network parameter, so the Q-value function of the network can be defined as

[0128] Q(s i , c i ; α, β) = V(s i , β) + A(s i , c i ; α).

[0129] S52, in order to obtain a unique solution of the Q-value function, the semantic value Q function is converted to:

[0130]

[0131] S53, in order to make the network semantic estimation Q value more accurate, another deep neural network is introduced to calculate the semantic target Q value Q target (t), called semantic target Q network. At time t, each agent shares a Q target (t) to reduce the training cost.

[0132] S54, in the network training stage, the current state of the multi-cell semantic communication cell, that is, the current channel gain, is taken as the input of the network of each cell. The semantic actor network selects the appropriate semantic action A based on the state. The policy set of all semantic agents is defined as μ = {μ1, …, μ N}, θ = {θ1, …, θ N} is the corresponding semantic resource allocation policy parameter set, represents the optimal semantic resource allocation policy of agent i.

[0133] S55, the semantic critic network calculates the Q value of the current network semantic action A based on the current semantic channel state of all agents and its corresponding semantic action A. The cumulative expected semantic reward of semantic agent i is

[0134]

[0135] where, represents the expectation, and γ represents the long-term discount rate.

[0136] S56, the target semantic function is defined as gradient​

[0137]

[0138] wherein the semantic state The semantic communication Critic network is trained by estimating the semantic Q value and the target semantic Q value, expressed as the loss value in the following formula:

[0139]

[0140] wherein

[0141] the target network semantic parameters The corresponding semantic policy, and then the semantic actor network constantly updates its semantic communication resource allocation policy policy through the feedback given by the semantic critic network.

[0142] S57, in the network test phase, the semantic actor network in the multi-cell semantic intelligent agent only needs to select the appropriate semantic resource allocation action A by analyzing the current local semantic communication channel state information given by the semantic critic network, and the semantic critic network does not need to know the semantic state information and the potential semantic resource allocation action of all intelligent agents.

[0143] In summary, the present application provides a resource allocation method of a semantic communication system based on deep reinforcement learning, which can maximize the semantic throughput of each unit per unit time; under the same conditions, it has greater semantic throughput, and the higher the semantic confidence, the smaller the maximum semantic throughput, overcoming the shortcomings of traditional resource methods that are not intelligent enough and need to adjust the allocation scheme at any time, while realizing the resource allocation of the multi-cell semantic communication system.

[0144] The above embodiments are only used to illustrate the technical solutions of the present application and not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A resource allocation method for a semantic communication system based on deep reinforcement learning, characterized in that, The main steps include: S1. The multi-cell intelligent semantic processing center collects transmitted and received data from each base station, uses the Transformer model to obtain a mapping table of the BLEU score of the semantic communication system as a function of the signal-to-noise ratio, and uses the mapping curve fitting to obtain the bit-to-semantic B2M conversion curve. The system semantic throughput is obtained through the bit rate. ,in This represents the knowledge matching coefficient; a higher value indicates a higher degree of knowledge matching between the user and the base station. and They represent time slots respectively. t Sub-channel allocation variables and power allocation variables; S2. Introduce a knowledge base (KB) auxiliary system to provide users with background knowledge for specific application domains for semantic reconstruction. S3. Determine the sub-channel allocation method and power allocation method for multi-cell network users, and obtain the transmission bit rate for each cell. The semantic throughput of a single cell is obtained through the B2M conversion curve. , where n represents the cell index; S4. By judging the system, we can find the optimization problem, and by simplifying the problem, we can find the final semantic resource allocation scheme that needs to be optimized. S5. Design and build the resource allocation network Semantic-JCP for a multi-cell semantic communication system based on DDQN & DDPG. After solving the resource allocation problem, the sub-channel allocation and power allocation that maximize the semantic throughput of the multi-cell system are finally obtained. 。 2. The resource allocation method according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11. For all users m∈M in the semantic communication network, define their bit-message B2M function under the current channel conditions as follows: This function is used to convert the bit rate in Shannon's theorem into the message rate. Its performance is determined by message attributes, semantic codecs, and the knowledge matching degree between the sender and receiver. A complete knowledge matching function between the base station and the user is introduced. ,but: Where, 0 < < 1, assuming all They are independent and identically distributed, satisfying... ,in This represents the mean. For variance; S12, obtain the data from the sub-channels of base station n and user m in the t-th time slot. achievable message transmission rate: in, This indicates that a resource allocation strategy is used on this channel. and If the bit transmission rate achievable in the nth cell is given by the following, then the total message transmission rate achievable in the nth cell is: S13. Using the Transformer model as the semantic conversion model, since the channel conditions are constantly changing, the BLEU score is used to evaluate the B2M conversion accuracy of the model under different signal-to-interference-plus-noise ratios (SINR).

3. The resource allocation method according to claim 1, characterized in that: In step S2, an auxiliary system of knowledge base KB is introduced into the user and semantic processing center to provide users with background knowledge in specific application fields such as medicine, communication, and chemistry related to the type of communication user for semantic reconstruction. The knowledge base matching coefficient is a self-defined number between 0 and 1. The larger the value, the better the relevant knowledge base matches, and the higher the semantic transmission rate under the same conditions.

4. The resource allocation method according to claim 2, characterized in that, Step S3 specifically includes the following steps: S31. A multi-cell intelligent semantic communication network consists of N base stations, each semantic communication base station being located at the center of its respective semantic communication cell. The set of N semantic communication base stations is represented as follows: Let each base station have L sub-channels, denoted as... In a multi-cell semantic communication network, there are M users randomly distributed across cells. The set of users served by semantic communication base station j is represented as ψ = {1, 2, ..., M}, n ∈ j. Each user is equipped with one antenna for transmitting and receiving semantic data, meaning the total number of antennas at the semantic base station equals the number of users. The total bandwidth of the multi-cell semantic communication system is B. Therefore, the bandwidth of each semantic communication sub-channel is... w=B / L ; S32. Set the frequency reuse coefficient of the multi-cell semantic communication network to f=1, and the sub-channel allocation matrix is ​​represented as follows: ,in This represents the sub-channel between the m-th user and the n-th base station at time slot t. The connection status is represented as follows: Within a single time slot, each user in the system is only allowed to use one sub-channel. Furthermore, to avoid interference between sub-channels within a single cell, the sub-channels within a single cell are set to be mutually orthogonal, as described below. S33. Define large-scale fading at time slot t as... Small-scale fading Then, within a time slot, the large-scale fading is constant, and the small-scale fading follows a Gaussian distribution. ; S34, At this time, the channel gain matrix is: , where the element is Define the set of base station transmit powers as follows: ,in Let Pmax represent the transmit power vector of the nth base station, and let Pmax be the limit of the total transmit power of base station n between 0 and Pmax. S35, Define the nth base station's... The received SINR of user m on each sub-channel is: in, N0 is the noise power. To avoid unwanted signal interference, in the t-th time slot, base station n and user m communicate via sub-channels. The achievable transmission bit rate is: in Finally, the transmission bit rate of each cell can be obtained: 。 5. The resource allocation method according to claim 4, characterized in that, Step S4 specifically includes the following steps: S41. The system objective is to adjust the power allocation and sub-channel allocation scheme for each cell, so the optimization problem can be expressed as: This optimization problem is a multi-objective and non-convex programming problem. The randomness of the problem makes it a nondeterministic problem; S42. Since the random variables in P1 exist only in the objective function, a new objective function is introduced and an additional constraint is added, making P1 equivalent to: in, It is the inverse function of the standard normal distribution. 。 6. The resource allocation method according to claim 5, characterized in that, Step S5 specifically includes the following steps: S51. Define the semantic channel state of the agent as S. The algorithm uses function V to estimate the semantic channel state value and the advantage function A to estimate the semantic resource allocation action value. An adversarial neural network is used to compute these two functions. When the semantic channel state is input, one network outputs a value function. ,in For network parameters, another network outputs a semantic action advantage function. ,in Since these are network parameters, the Q-value function of the network is defined as follows: ; S52. To obtain a unique solution to the Q-value function, the semantic value Q-function is transformed into: ; S53. To make the network semantic estimation Q-value more accurate, another deep neural network is introduced to calculate the semantic target Q-value. At time slot t, each agent shares one To reduce training costs.

7. The resource allocation method according to claim 6, characterized in that, Step S5 also includes the following steps: S54. During the network training phase, the current channel gain serves as the input to each cell network. The semantic actor network selects an appropriate semantic action A based on this state, and the policy set of all semantic agents is defined as follows: , Assign a set of strategy parameters to the corresponding semantic resources. The optimal semantic resource allocation strategy for agent i is represented. S55. The semantic critic network calculates the Q-value of the current semantic action A based on the current semantic channel state of all agents and their corresponding semantic actions A. The cumulative expected semantic reward for semantic agent i is... , in, Expressing expectations, Indicates the long-term discount rate; S56. Define the target semantic function as... gradient in, The semantic communication Critic network is trained by estimating the semantic Q-value and the target semantic Q-value, expressed as the loss value in the following formula: in , For target network semantic parameters The corresponding semantic policy is then used by the semantic actor network to continuously update its semantic communication resource allocation policy based on feedback from the semantic critic network. S57. During the network testing phase, the semantic actor network in a multi-cell semantic agent only needs to select the appropriate semantic resource allocation action A by analyzing the current local semantic communication channel state information given by the semantic critic network. The semantic critic network does not need to know the semantic state information and potential semantic resource allocation actions of all agents.