Transmission parameter optimization method of time division multiple access communication system and related equipment

By using a pre-trained transmission parameter optimization model in a smart reflector-assisted time division multiple access communication system, the system directly utilizes the location information of the mobile terminal to optimize time slot allocation and beamforming, solving the problem of high complexity in obtaining channel state information, achieving efficient transmission parameter decision-making, and improving system performance and communication quality of the mobile terminal.

CN121240109APending Publication Date: 2025-12-30GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511233863.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In intelligent reflector-assisted time division multiple access communication systems, existing technologies require precise acquisition of channel state information for each link, resulting in high computational complexity and large latency. In particular, the reliability of transmission parameter decision optimization is low when users move and channels change rapidly over time.

Method used

By using a pre-trained transmission parameter optimization model, the location information of the mobile terminal is directly utilized, and data matching is performed in combination with the pre-trained model to optimize time slot allocation, beamforming vector and phase shift sequence. This avoids dependence on real-time channel state information, reduces computational complexity and dynamically adjusts the transmission strategy.

Benefits of technology

Without needing to obtain instantaneous channel state information, it significantly improves the transmission performance of the time division multiple access communication system, ensuring communication fairness and service quality for mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121240109A_ABST
    Figure CN121240109A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a transmission parameter optimization method of a time division multiple access communication system and related equipment, a wireless communication system comprises a base station, an intelligent reflecting surface and a plurality of mobile terminals, and the method comprises the following steps: acquiring terminal positions of all the mobile terminals in a current time step, and acquiring a transmission parameter optimization model; inputting all terminal positions into a transmission parameter optimization model for data matching to obtain output optimized time slot allocation, optimized beam forming vectors and optimized phase shift sequences corresponding to a plurality of time slots; and allocating corresponding time slots to the plurality of mobile terminals in the current time step according to the optimized time slot allocation, and adjusting a beam forming vector between the base station and each mobile terminal in the current time step according to the optimized beam forming vector. And the reflection unit of the intelligent reflection surface of each time slot is adjusted in the current time step according to the optimized phase shift sequence, so that the transmission performance of the whole time division multiple access communication system can be reliably improved under the condition that instantaneous channel state information does not need to be acquired.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless communication, in particular to a transmission parameter optimization method of a time division multiple access communication system and related equipment. BACKGROUND

[0002] In recent years, in the next generation communication system, in order to meet the increasing number of users and high speed, high reliability communication demand, the intelligent reflecting surface (RIS) technical scheme has been widely concerned. Intelligent reflecting surface can intelligently regulate and control the propagation environment of wireless signal through a large number of low-cost passive reflecting units, so as to enhance signal coverage and improve system capacity. In the intelligent reflecting surface (RIS) assisted multi-user time division multiple access (TDMA) communication system, in order to improve the overall communication quality and efficiency, it is usually necessary to jointly optimize the system transmission parameters. These parameters mainly include: the time allocation scheme for allocating communication time slots to different mobile terminals and the passive reflecting unit phase shift sequence configured by the intelligent reflecting surface end to enhance the propagation of each time slot signal.

[0003] In the related art, in the communication system provided with the intelligent reflecting surface, the system transmission parameter optimization is usually performed according to the accurate channel state information (CSI) corresponding to the current time slot in the system at the beginning of each time slot, but since the intelligent reflecting surface usually contains hundreds of reflecting units, it is necessary to consume a large number of pilots and complex estimation algorithms to accurately obtain the channel state information of each link, which will consume valuable time-frequency resources and bring significant calculation delay. Especially in the scene of user movement and fast channel variation, frequent channel estimation will make the system overhead unbearable, resulting in low reliability of transmission parameter decision optimization of the system in the actual communication system provided with the intelligent reflecting surface. SUMMARY

[0004] The embodiments of the present application provide a transmission parameter optimization method of a time division multiple access communication system and related equipment, which can improve the transmission parameter decision optimization reliability between the base station and the mobile terminal in the time division multiple access communication system provided with the intelligent reflecting surface.

[0005] To achieve the above-mentioned purpose, the first aspect of the embodiments of the present application provides a transmission parameter optimization method of a time division multiple access communication system, the wireless communication system includes a base station, an intelligent reflecting surface and a plurality of mobile terminals, and the method comprises:

[0006] Obtaining the terminal positions of all the mobile terminals at the current time step, and obtaining a transmission parameter optimization model, the time step includes a plurality of time slots;

[0007] inputting all the terminal positions into the transmission parameter optimization model for data matching to obtain outputted optimized time slot allocation, optimized beamforming vectors and optimized phase shift sequences corresponding to the time slots;

[0008] allocating corresponding time slots for the mobile terminals in the current time step according to the optimized time slot allocation, adjusting the beamforming vectors between the base station and each mobile terminal in the current time step according to the optimized beamforming vectors, and adjusting the reflecting units of the intelligent reflecting surface of each time slot in the current time step according to the optimized phase shift sequences.

[0009] To achieve the above object, a second aspect of embodiments of the present application provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the transmission parameter optimization method of the time division multiple access communication system according to the first aspect when executing the computer program.

[0010] To achieve the above object, a third aspect of embodiments of the present application provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the transmission parameter optimization method of the time division multiple access communication system according to the first aspect.

[0011] The method for optimizing transmission parameters of a time division multiple access communication system and related equipment provided by the embodiments of the present application, the wireless communication system comprising a base station, a smart reflecting surface and a plurality of mobile terminals, the method comprising: first, obtaining the terminal positions of all mobile terminals at the current time step, and obtaining a transmission parameter optimization model, the time step comprising a plurality of time slots; then, inputting all terminal positions into the transmission parameter optimization model for data matching to obtain an output of optimized time slot allocation, an optimized beamforming vector and an optimized phase shift sequence corresponding to the plurality of time slots; finally, allocating corresponding time slots to the plurality of mobile terminals in the current time step according to the optimized time slot allocation, adjusting the beamforming vector between the base station and each mobile terminal in the current time step according to the optimized beamforming vector, and adjusting the reflecting elements of the smart reflecting surface in each time slot in the current time step according to the optimized phase shift sequence. The embodiments of the present application directly establish the mapping relationship between the mobile terminal positions and the optimal transmission parameters by using the pre-trained transmission parameter optimization model, thereby completely avoiding the dependence on real-time and accurate channel state information (CSI), directly using the more easily obtained real-time user position information as the decision basis, avoiding the complex channel estimation process and the high overhead and calculation delay caused thereby, and by inputting the user position into the pre-trained model, the optimized time slot allocation, the smart reflecting surface phase shift sequence and the beamforming vector scheme can be obtained, and the system is adjusted accordingly, which not only significantly reduces the complexity of resource optimization, but also dynamically adjusts the transmission strategy in real time according to the user movement, so that the system configuration can more effectively adapt to the macroscopic channel environment change caused by the change of user position, and finally the transmission performance of the entire time division multiple access communication system can be reliably improved without obtaining the instantaneous channel state information, especially the communication fairness and service quality of the mobile terminal are guaranteed.

[0012] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and attained by the structure particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a structural schematic diagram of a time division multiple access communication system provided by an embodiment of the present application.

[0014] Figure 2 is a flowchart of a method for optimizing transmission parameters of a time division multiple access communication system provided by another embodiment of the present application.

[0015] Figure 3 is a schematic diagram of the running framework of Algorithm 1 provided by another embodiment of the present application.

[0016] Figure 4Fig. 1 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application.

[0018] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in a manner different from the module division in the device or the order in the flowchart.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification is for describing the embodiments of the present application only and not intended to limit the present application.

[0020] In recent years, in the next generation communication system, in order to meet the increasing number of users and high speed, high reliability communication demand, the intelligent reflecting surface (RIS) technical scheme has received widespread attention. Intelligent reflecting surface can intelligently regulate and control the propagation environment of wireless signal through a large number of low-cost passive reflecting units, thereby enhancing signal coverage and improving system capacity. In the intelligent reflecting surface (RIS) assisted multi-user time division multiple access (TDMA) communication system, in order to improve the overall communication quality and efficiency, it is usually necessary to jointly optimize the system transmission parameters. These parameters mainly include: the time allocation scheme for allocating communication time slots to different mobile terminals and the passive reflecting unit phase shift sequence configured by the intelligent reflecting surface end to enhance the signal propagation of each time slot.

[0021] In the related art, in a communication system provided with an intelligent reflecting surface, the system transmission parameters are usually optimized according to the accurate channel state information (CSI) corresponding to the current time slot in the system at the beginning of each time slot, but since the intelligent reflecting surface usually contains hundreds of reflecting units, to accurately obtain the channel state information of each link, a large number of pilots and complex estimation algorithms are required, which will consume valuable time-frequency resources and bring significant calculation delay, especially in the scene of user movement and fast channel variation, frequent channel estimation will make the system overhead unbearable, resulting in low reliability of system transmission parameter decision optimization in the actual communication system provided with an intelligent reflecting surface.

[0022] In order to improve the optimization reliability of transmission parameter decision between the base station and the mobile terminal in the time division multiple access communication system provided with the intelligent reflecting surface, the embodiment of the application directly establishes the mapping relationship between the mobile terminal position and the optimal transmission parameter by using the pre-trained transmission parameter optimization model, thereby completely avoiding the dependence on real-time and accurate channel state information (CSI), directly using the more easily obtained user real-time position information as the decision basis, avoiding the complex channel estimation process and the high overhead and calculation delay caused thereby, and through inputting the user position into the pre-trained model, the optimized time slot allocation, intelligent reflecting surface phase shift sequence and beamforming vector scheme can be obtained instantaneously, and the system is adjusted accordingly, which not only significantly reduces the complexity of resource optimization, but also dynamically adjusts the transmission strategy in real time according to the user movement, so that the system configuration can more effectively adapt to the macro channel environment change caused by the user position change, and finally the transmission performance of the entire time division multiple access communication system can be reliably improved without obtaining the instantaneous channel state information, especially the communication fairness and service quality of the mobile terminal are guaranteed.

[0023] In order to better describe the transmission parameter optimization method of the time division multiple access communication system provided by the application, the time division multiple access communication system applied to the transmission parameter optimization method of the time division multiple access communication system is described first as follows. Referring to Figure 1 , it is a structural schematic diagram of a time division multiple access communication system provided by the embodiment of the application. As Figure 1 shown, the time division multiple access communication system is composed of a base station (BS), an intelligent reflecting surface RIS with M=M x ×M y reflective units and K mobile terminals carried by the mobile terminals, wherein the base station BS is equipped with N antennas, and each mobile terminal is equipped with one antenna. Let M={1, 2, L, M}, K={1, 2, L, K} and N={1, 2, L, N} represent the set of reflective units of the intelligent reflecting surface, the mobile terminal and the base station antenna of the base station respectively. In this illustrative scenario, the direct link between the base station and the mobile terminal MU k has poor communication quality due to the obstruction of the building. For this reason, the base station The signals can be sent to the smart reflector deployed on the outer wall of another building, and then reflected by the smart reflector to each mobile terminal. Among them, the smart reflector is controlled by the RIS controller connected to it, and the RIS controller can adjust the phase shift of multiple reflecting units on the smart reflector in real time according to the optimization strategy generated by the base station (as shown by the dashed line), so as to actively improve the quality of the reflection channel, and to ensure the communication performance of multiple mobile terminals. In addition, the positions of the base station and the smart reflector RIS remain fixed, while the terminal position of each mobile terminal can change as the mobile terminal moves, that is, the terminal position of the mobile terminal randomly changes within a given area δ at any time. By using high-precision positioning technology (such as GPS in outdoor scenes), the accurate position information of the base station BS, the smart reflector RIS and the mobile terminal can be obtained.

[0024] In the time division multiple access communication system of the present application, the information transmission from the base station BS to the mobile terminal is based on time slots, each time slot has a length of T seconds, and multiple time slots form a time step. In each time slot, the channels in the time division multiple access communication system are constant, but they can be independently changed between different time slots. Since the position of the mobile terminal changes over time, the channel coefficients of the link from all the antennas of the base station BS to the smart reflector RIS and the communication link from the smart reflector RIS to the kth mobile terminal in the ith time slot are represented as and (also the second reflection channel model and the first reflection channel model), where q k,i ={x k,i ,y k,i ,z k,i} represents the terminal position of the mobile terminal k, x k,i ,y k,i ,z k,i represent the horizontal and vertical coordinates of the mobile terminal k respectively. Considering that the position change speed of the mobile terminal is usually much lower than the channel change, its position can remain static within multiple time slots (i.e. a time step). Therefore, in order to simplify the representation, the time slot index i in q k,i will be omitted in the subsequent part of the present application.

[0025] Since the time division multiple access (TDMA) technology is adopted as the multiple access scheme in the time division multiple access communication system of the present application. Under this scheme, the system allocates exclusive communication time slots for each mobile terminal. Therefore, in any given time slot, only one target mobile terminal communicates with the base station, while other non-target mobile terminals are in a non-working state. Therefore, the set of all time slot allocation parameters (i.e. the selection of mobile terminals that can be allocated time slots) is set as τ = {τ1, τ2, L, τ K}, where τk Let τ represent the time slot allocation parameters for the k-th mobile terminal, satisfying τ k ∈(0,1),

[0026] Based on the characteristics of Time Division Multiple Access (TDMA), multi-user communication links in a TDMA communication system can be decoupled into a series of independent single-user communication links. Specifically, in each specific time slot, only the cascaded channel between the base station (BS), the intelligent reflector (RIS), and the target mobile terminal needs to be considered. At this time, the channel between the base station and other non-target mobile terminals can be ignored. Therefore, for a given mobile terminal k, its communication performance can be evaluated by performing independent Monte Carlo simulations on that single-user link.

[0027] In this application, it is assumed that neither the instantaneous channel state information (CSI) nor the statistical channel state information (CSI) of the time division multiple access communication system is available. However, since the location of the mobile terminal can be obtained, the phase shift of each reflector element of the intelligent reflector RIS can be determined by the location of the mobile terminal, Q = {q}. k Adjust Θ, k∈K}. Therefore, Θ Q =diag(θ) Q ) represents the phase shift matrix of the reflective unit of the intelligent reflector RIS, where θ Q =[θ Q,1 ,θ Q,2 ,L,θ Q,M ] T ∈£ M×1 Let θ represent the sequence of phase shift parameters of the reflective elements of the intelligent reflector RIS given Q. Q,m This represents the phase shift of the m-th element given Q, and the reflection coefficient amplitude of each element is set to 1. This represents a vector set of beamforming parameters between the base station and each mobile terminal, where K is the number of mobile terminals and N is the number of antennas at the base station. When the base station communicates with the k-th mobile terminal, the base station uses the beamforming parameter ω... k Actively control the transmitted signal, among which, ω n,k This represents the beamforming parameters of the nth antenna when the base station communicates with the kth mobile terminal. After weighting by the beamforming parameters, the transmitted signal forms a transmitted beam with a specific spatial directionality. The beamforming vector must satisfy the total transmitted power constraint. Where P B This represents the total transmit power of the base station (BS).

[0028] Furthermore, considering that the actual RIS phase shift is discrete, and the phase shift of the intelligent reflector RIS is quantized using D bits, then the set of discrete phase shift values ​​for each reflector element can be represented as Θ = {0, Δξ, ..., (2 D-1)Δξ}, where Δξ=π / 2 (D-1) .

[0029] Based on such Figure 1 The time division multiple access (TDMA) communication system shown below will be described in detail below with respect to the transmission parameter optimization method of the TDMA communication system in the embodiments of this application. (Refer to...) Figure 2 This is an optional flowchart of a transmission parameter optimization method for a time division multiple access communication system provided in this application embodiment. Figure 2 The method may include, but is not limited to, steps 100 to 300. It is also understood that this embodiment... Figure 2 The order of steps 100 to 300 is not specifically limited; the order of steps can be adjusted or certain steps can be added or removed according to actual needs. This method for optimizing transmission parameters in a time-division multiple access communication system can be applied to the information transmitting end or to processors or controllers connected to the time-division multiple access communication system.

[0030] Step 100: Obtain the terminal location of all mobile terminals at the current time step, and obtain the transmission parameter optimization model.

[0031] Step 100 is described in detail below.

[0032] In some embodiments, due to the complexity of the channel between the base station and multiple mobile terminals, and because each reflective unit and the information transmitter, as well as each reflective unit and the information receiver, are considered to have a communication link during the reflection of data information by the intelligent reflective surface, it is extremely difficult to obtain the instantaneous channel state between the information receiver and the intelligent reflective surface in real time, which requires particularly high hardware specifications and is also difficult to achieve.

[0033] Therefore, in this embodiment, the terminal locations of all mobile terminals at the current time step are obtained through a positioning system (such as GPS or network-side positioning technology). These terminal locations are the key input replacing traditional Channel State Information (CSI), as they macroscopically reflect the geometric relationship between the mobile terminal, the base station, and the smart reflector. Then, the terminal locations of all mobile terminals at the current time step are used to replace the instantaneous channel state at the current moment as the basis for optimizing the transmission parameters of the time-division multiple access (TDMA) communication system at the current moment. In this embodiment, the transmission parameters to be optimized in the TDMA communication system include the time slot allocation parameter τ assigned by the base station BS to each mobile terminal k. k The beamforming parameters ω of the base station (BS) for transmitting information to the corresponding mobile terminal in a time slot. k And the phase shift sequence θ of the intelligent reflector RIS in each time slot. Q (That is, the reflection phase shift of each reflection unit).

[0034] Simultaneously, the system loads or invokes a pre-built transmission parameter optimization model, which is typically a well-trained deep reinforcement learning model. It is worth noting that in this application, a time step consists of multiple time slots, meaning that a decision cycle (time step) is temporally divided into several shorter physical transmission units (time slots). This lays the foundation for subsequent finer dynamic adjustments under a single decision and for effective performance evaluation during the model training phase.

[0035] Furthermore, to obtain appropriate optimized transmission parameter decisions for the current time step, a transmission parameter optimization model is constructed in advance using a deep reinforcement learning framework, based on relevant data information from the time division multiple access communication system. This model determines appropriate transmission parameter optimization decisions based on the terminal locations of all mobile terminals, thereby reducing the information reception bit error rate at the information receiver. The construction of this transmission parameter optimization model will be further described below.

[0036] The process of obtaining the transmission parameter optimization model includes steps 110 to 150.

[0037] Step 110: Based on the phase shift parameter sequence of the smart reflector and the position parameters of multiple mobile terminals, generate the cascaded channel model corresponding to each mobile terminal in each time slot.

[0038] Step 110 will be described in detail below.

[0039] In some embodiments, in order to subsequently perform mathematical modeling of system performance, it is first necessary to construct a channel model of the communication link. This step focuses on the phase shift parameter sequence θ of any given set of smart reflectors. Q =[θ Q,1 ,θ Q,2 ,L,θ Q,M ] T and the location parameters q of multiple mobile terminals k,i ={x k,i ,y k,i ,z k,i Under the condition of}, the cascaded channel model h from the base station BS to the intelligent reflector RIS and then to the mobile terminal in each time slot is derived. k(i,Q). This cascaded channel model describes the equivalent channel response of the complete propagation path of a signal from the base station, through reflection by the smart reflector, and finally to a specific mobile terminal. Since the terminal location parameters of the mobile terminal determine the path loss and angle information of the two links from the base station to the smart reflector and from the smart reflector to the mobile terminal, and the phase shift parameter sequence determines the phase modulation method of the smart reflector on the incident signal, this cascaded channel model that reflects the overall end-to-end transmission characteristics can be obtained by mathematically multiplying the channel and phase shift matrices of these two links.

[0040] The following section will further describe how to construct this cascaded channel model.

[0041] The process involves generating a cascaded channel model for each mobile terminal in each time slot based on the phase shift parameter sequence of the smart reflector and the position parameters of multiple mobile terminals, including steps 111 to 112.

[0042] Step 111: Based on the location parameters of each mobile terminal, determine the first reflection channel model between each mobile terminal and the smart reflector.

[0043] Step 112: Based on the product of the first reflection channel model, the phase shift parameter sequence, and the second reflection channel model, the cascaded channel model is obtained.

[0044] Steps 111 to 112 are described in detail below.

[0045] In some embodiments, to construct a complete channel model, it is first necessary to model each independent link in the communication system. The core of this step is to first model the location parameters q of each mobile terminal. k,i ={x k,i ,y k,i ,z k,i Using the position parameters of the mobile terminal and the smart reflector, the first reflection channel model f between each mobile terminal and the smart reflector is determined. k (i,q k The model describes the path characteristics of the signal propagating from the smart reflector to each mobile terminal; and it uses the position parameters of the base station and the smart reflector to determine the second reflection channel model G(i) between the base station and the smart reflector, which describes the path characteristics of the signal propagating from the base station to the smart reflector. These position parameters are the basis for calculating key channel parameters such as path loss, propagation delay, and signal arrival / departure angles.

[0046] After obtaining the first reflection channel model f respectively k (i,q kAfter obtaining the first reflection channel model and the second reflection channel model G(i), they are combined to form a cascaded channel model describing end-to-end signal propagation. This step involves mathematically multiplying the determined first reflection channel model, the externally given sequence of phase shift parameters, and the second reflection channel model representing the channel characteristics between the smart reflector and the base station to obtain the final cascaded channel model. Where h k (i,Q)∈£ 1×N The phase shift parameter sequence participates in the calculation in the form of a diagonal matrix, where each diagonal element represents the phase shift applied to the corresponding reflective unit on the smart reflector. Through this series of product operations, the entire physical process of the system transmitting the signal from the base station, through the regulation of each unit on the smart reflector, and then reflecting it to the mobile terminal is accurately expressed as a unified and equivalent channel matrix, i.e., a cascaded channel model.

[0047] Through the implementation of steps 111 to 112 above, the complex communication environment is decomposed into multiple independent and easy-to-model channel segments. Then, through rigorous mathematical calculations, these segments are combined with the control effect of the intelligent reflector to form a comprehensive channel model that can accurately reflect the end-to-end signal transmission characteristics. This makes the channel modeling process clear and logical, and can accurately incorporate the two core variables, the position parameters of the mobile terminal and the phase shift parameter sequence of the intelligent reflector, into the model, providing a solid and reliable mathematical foundation for subsequent system performance analysis and optimization algorithm design.

[0048] Step 120: Based on beamforming parameters, time slot modulation symbols, noise power, and the cascaded channel model corresponding to each time slot, obtain the average bit error rate model.

[0049] In some embodiments, based on the established cascaded channel model, the system will further construct a key performance indicator model for evaluating communication reliability, namely the average bit error rate model. This model predicts the average error probability of data transmission by comprehensively considering multiple influencing factors. Specifically, it uses the beamforming vector ω set by the base station for the target terminal. k The actual transmitted time slot modulation symbol S i,j The inherent noise power in the communication system, and the cascaded channel model h corresponding to each time slot. k (i,Q) are integrated. Using these parameters, the signal-to-noise ratio (SNR) of the mobile terminal at the receiving end can be calculated first, and then the instantaneous bit error rate can be derived based on communication principles. Finally, by statistically averaging all possible channel states, an average bit error rate model that can characterize the long-term communication quality under specific transmission parameter configurations is obtained, as described below.

[0050] The average bit error rate model is obtained based on beamforming parameters, time slot modulation symbols, noise power, and the cascaded channel model corresponding to each time slot, including the following steps 121 to 123.

[0051] Step 121: Based on the product of the cascaded channel model, beamforming parameters, and time slot modulation symbols, plus the noise power, obtain the received signal model for each mobile terminal.

[0052] Step 122: Based on the expected value of the received signal model, subtract the product of noise power and system bandwidth, and then divide by the logarithm of noise power, system bandwidth and modulation order in turn to obtain the received signal-to-noise ratio model of each mobile terminal in each time slot.

[0053] Step 123: Based on the product of the transmission power corresponding to the information sent by the base station to each mobile terminal in each time slot and the received signal-to-noise ratio model, perform expectation processing to obtain the average bit error rate model.

[0054] Steps 121 to 123 are described in detail below.

[0055] In the time-division multiple access communication system of this application, it is assumed that M-ary modulation and Gray coding are used. Let T... s Let T be the time length of a modulation symbol. Then, the number of modulation symbols transmitted by T in each time slot is J = T / T. s .

[0056] To analyze the signal quality at the receiving end, a mathematical expression needs to be established to describe the signal actually received by the mobile terminal. This step involves integrating all elements along the signal transmission path to construct a received signal model Y for each mobile terminal. i,j,k Specifically, this model is based on the cascaded channel model h. k (i,Q), beamforming parameters ω k and the time slot modulation symbol S i,j The product of the two is added to the additive white Gaussian noise, as shown in the following formula (1).

[0057] Y i,j,k =h k (i,Q)ω k S i,j +n k (1)

[0058] Among them, the time slot modulation symbol S i,jThe base station transmits the j-th modulation symbol in the i-th time slot; beamforming parameters (usually a complex vector) weight this symbol to form a directional transmission beam; the cascaded channel model describes the changes in this beam after propagation through space; finally, a noise power term representing the effects of various interferences and thermal noise in the communication environment is added. This model can accurately describe the instantaneous shape of the composite signal composed of the desired signal and irrelevant noise on any mobile terminal antenna. It is the additive white Gaussian noise (AWGN) of the k-th mobile terminal, and N0 is the power spectral density (i.e., noise power) of the AWGN.

[0059] Since signal and noise are statistically independent, the total average power of the received signal model equals the sum of the average signal power and the average noise power. Therefore, we have... Where P s It is the average power of the signal, P noi It is the average power of noise. Let P represent the mathematical expectation operation, and let P be... noi =N0B. Therefore, the average signal power can be expressed as Assume the base station has a constant transmission power, denoted as P. B The received signal-to-noise ratio (SNR) of the k-th mobile terminal for decoding each symbol in the i-th time slot can be expressed as shown in the following formula (2).

[0060]

[0061] Assume the energy of each modulation symbol is constant E. s And there is P B =E s / T s The received signal-to-noise ratio can also be expressed as γ. k (i,Q)=E s |h k (i,Q)ω k | 2 / (T s N0B). Because the pulse shaping at the base station BS satisfies T s =1 / B (for example, using a raised cosine pulse with a roll-off factor β=1), so γ can be derived. k (i,Q)=E s |h k (i,Q)ω k | 2 / N0, which is equal to the symbol signal-to-noise ratio (SNR) received at the k-th mobile terminal in the i-th time slot (denoted as γ). s,k (i,Q)). Therefore, we have the following formula (3).

[0062]

[0063] Based on this, a received signal model Y was established. i,j,k Next, the system needs to extract a key metric for measuring signal clarity, namely the received signal-to-noise ratio (SNR). The goal of this step is to obtain the received SNR model γ for each mobile terminal in each time slot (e.g., the i-th time slot). b,k (i,Q). This process first calculates the expected value of the received signal model. This represents the total average power of the received signal. Then, based on this expected value, the product of noise power and system bandwidth N0B (i.e., total noise power) is subtracted to separate the pure average signal power. Next, this average signal power is divided successively by the noise power N0 and the system bandwidth B (to obtain the signal-to-noise ratio per hertz), and finally divided by the logarithm of the modulation order log2M (i.e., the number of bits carried by each symbol), finally obtaining the received signal-to-noise ratio model normalized to each bit as shown in the following formula (4).

[0064]

[0065] This model is a core physical quantity for evaluating the reliability of communication links; a higher value indicates better signal quality and a lower probability of decoding errors. Since the bit error rate (BER) is determined by the bit frequency response (SNR), the BER for the k-th mobile terminal decoding each symbol in the i-th time slot can be expressed as P. b,k (γ b,k (i,Q)).

[0066] Due to G(i) and f in the reflection link k (i,q k There exists a time-varying channel coefficient, γ b,k (i,Q) is a random variable, therefore, P b,k (γ b,k (i,Q) will change randomly in each time slot.

[0067] Therefore, after obtaining the received signal-to-noise ratio model, it is necessary to transform it into a more intuitive performance metric, namely the average bit error rate model. This step predicts the final error rate of data transmission by associating the signal-to-noise ratio with a specific modulation and coding scheme. Specifically, it involves calculating the transmission power P corresponding to the information sent by the base station to each mobile terminal in each time slot. b,k With the received signal-to-noise ratio model γ b,k The product of (i,Q) (this step is actually an intermediate process for calculating the instantaneous bit error rate, since the bit error rate is a function of the signal-to-noise ratio), and then the expected value is taken for all possible random channel changes, as shown in the following formula (5).

[0068]

[0069] The purpose of expectation processing is to eliminate the impact of instantaneous channel fluctuations, thereby obtaining an indicator that reflects the average performance over a longer period or under various channel implementations. This final average bit error rate model provides a direct and crucial performance evaluation basis for system optimization.

[0070] Through the implementation of steps 121 to 123 above, an accurate received signal model is constructed. Then, through a series of mathematical operations, the core received signal-to-noise ratio model is extracted from it. Finally, combined with the statistical averaging method, it is transformed into an intuitive average bit error rate model, ensuring that the final average bit error rate model can accurately reflect the comprehensive impact of various factors such as beamforming parameters, modulation methods, and channel conditions on communication reliability. This provides a solid and quantitative theoretical foundation for subsequent system performance optimization.

[0071] Step 130: Subtract the average bit error rate model from the first value, and then multiply by the time slot allocation parameter to obtain the throughput model for each mobile terminal.

[0072] Step 130 will be described in detail below.

[0073] In some embodiments, after obtaining the average bit error rate model, the system will use the communication reliability index (average bit error rate) ) and resource allocation efficiency index (time slot allocation τ) k By combining these factors, a performance metric that is closer to the actual user experience is constructed, namely the throughput model of each mobile terminal. Specifically, firstly, the average bit error rate model is subtracted from a "first value" (usually 1) representing ideal transmission to obtain the successful transmission rate; then, this successful transmission rate is multiplied one by one by the total number of bits that the base station can send and the time slot allocation parameters allocated to the user to obtain the throughput model as shown in the following formula (6).

[0074]

[0075] Here, L represents the total number of bits that the base station can transmit. The time slot allocation parameter represents the proportion of the total communication time for that user. Through this calculation, the throughput model not only reflects the accuracy of data transmission but also the amount of communication resources allocated to users, thus comprehensively measuring the amount of information effectively received by each user per unit of time.

[0076] Step 140: Based on the minimum throughput model that maximizes all mobile terminals, the transmission parameter optimization problem is obtained.

[0077] Step 140 will be described in detail below.

[0078] The objective of this application is to optimize the phase shift matrix, beamforming vector, and time allocation factor (i.e., the time slot allocation parameter τ allocated by the base station BS to each mobile terminal k) of the RIS without requiring Channel State Information (CSI). k The beamforming parameters ω of the base station (BS) for transmitting information to the corresponding mobile terminal in a time slot. k And the phase shift sequence θ of the intelligent reflector RIS in each time slot. Q ), to maximize the minimum throughput of all mobile terminals within a given area δ.

[0079] Based on this, the system will establish the ultimate goal of the entire optimization process, namely, to construct a formalized transmission parameter optimization problem. This optimization problem aims to achieve a balance between the overall system performance and fairness among users. Specifically, the throughput model of each mobile terminal is used as the basis, and the optimization objective is to maximize the minimum throughput among all mobile terminals, so as to obtain the transmission parameter optimization problem, as shown in the following formula (7).

[0080]

[0081] This "maximizing the minimum" criterion, also known as the Max-Min fairness criterion, aims to improve the experience of the worst-performing user in the system, thereby preventing some users from being unable to obtain service due to poor channel conditions. This ensures the service fairness and overall robustness of the entire network. This transmission parameter optimization problem has pointed the way for subsequent algorithm design and model training.

[0082] For the transmission parameter optimization problem (7), the traditional approach is to use CSI as the key performance indicator of the system (such as bit error rate) to derive an analytical expression for the optimization variables (such as RIS phase shift), and then apply traditional mathematical tools such as convex optimization to solve it. However, under the core setting of this application that does not require CSI, the preconditions for establishing this analytical expression are missing, which makes it impossible to obtain the analytical expression using traditional optimization methods. Therefore, the transmission parameter optimization problem (7) is a very challenging optimization problem. To solve this problem, the transmission parameter optimization problem (7) is reconstructed into a deep reinforcement learning (DRL) task and solved using the PPO algorithm.

[0083] Step 150: Construct an initial transmission parameter optimization model based on the transmission parameter optimization problem, and train the initial transmission parameter optimization model in multiple rounds to obtain the transmission parameter optimization model.

[0084] In some embodiments, the system enters the core machine learning model building and training phase. This step first designs and builds an initial transmission parameter optimization model based on the transmission parameter optimization problem defined in step 140. This model is typically a deep reinforcement learning agent with a specific network structure, whose state space corresponds to the terminal location and action space corresponds to the transmission parameters to be optimized. Subsequently, the system trains this initial transmission parameter optimization model through multiple rounds. During training, the model continuously interacts with a simulated communication environment, repeatedly trying different parameter combinations and receiving rewards or penalties based on the resulting throughput (i.e., the objective of the optimization problem), continuously adjusting its internal parameters. After sufficient training rounds, the model finally learns how to directly output the optimal parameters that maximize the minimum throughput model for all mobile terminals based on the input terminal location, thus obtaining a mature, efficient, and directly deployable final transmission parameter optimization model.

[0085] The following section first describes how to construct an initial transmission parameter optimization model.

[0086] The initial transmission parameter optimization model is constructed based on the transmission parameter optimization problem, including the following steps 151 to 155.

[0087] Step 151: Based on the time slot allocation parameters, beamforming parameters, and phase shift parameter sequence in the transmission parameter optimization problem, obtain the action space.

[0088] Step 151 will be described in detail below.

[0089] When using deep reinforcement learning methods to solve the transmission parameter optimization problem (7), the base station BS can be regarded as an intelligent agent. And make decisions based on the observation of the location of the mobile terminal. First, the transmission parameter optimization problem (7) is transformed into a Markov decision process (MDP) problem, and then a deep reinforcement learning algorithm based on PPO is proposed to solve the MDP problem.

[0090] An MDP problem consists of quadruples Composition. Among them, r and γ∈[0,1] represent the state space, action space, reward, and discount factor, respectively. State space It is the state space, that is, the set of states, and the action space. It is the action space, i.e., the intelligent agent. The set of actions, with a discount factor γ reflecting the proportion of future rewards at the current time step. Specifically, the actions, states, and rewards of its MDP are defined as follows.

[0091] In some embodiments, in order to transform the transmission parameter optimization problem into a Markov decision process (MDP) that can be solved by deep reinforcement learning, it is first necessary to define the set of operations that the agent (i.e., the optimization model) can perform, i.e., the action space. This step is based on the core optimization variable in the transmission parameter optimization problem, namely the slot allocation parameter τ. t ={τ 1,t ,τ 2,t ,L,τ K,t Beamforming parameter ω t ={ω 1,t ,ω 2,t ,L,ω K,t} and the phase shift parameter sequence θ Q,t ={θ Q,1,t ,L,θ Q,M,t}, to jointly construct and obtain the action space. Specifically, at the t-th time step, let ω t ={ω 1,t ,ω 2,t ,L,ω K,t}, ω k,t =[ω 1,k,t ,ω 2,k,t ,L,ω N,k,t ] T ω n,k,t Let θ represent the beamforming parameters of the nth antenna when the base station communicates with the kth mobile terminal; Q,t ={θ Q,1,t ,L,θ Q,M,t}, θ Q,m,t Let τ represent the phase shift parameter sequence of the m-th reflection unit of the RIS when the terminal position of the mobile terminal is given by Q; t ={τ 1,t ,τ 2,t ,L,τ K,t}, τ k,t This represents the time slot allocation parameters for the k-th mobile terminal. Configure the beamforming vector ω. n,k,t and time allocation factor τ k,t For continuous operation, the RIS phase shift is a discrete operation. Assume the beamforming vector ω for continuous operation... t The dimension is (N,K), and the time allocation factor τ for continuous actions. t Let the dimension be (K,1), and the cardinality of the discrete action be M. It is easy to calculate that the dimension of the RIS phase shift of the discrete action is 2. DM .

[0092] This action space fully describes all possible choices that the agent can make at any given decision moment. Specifically, it is a hybrid action space a that includes continuous variables (such as beamforming vectors and time slot allocation ratios) and discrete variables (such as quantized RIS phase shift values). t ={ω t ,θ Q,t ,τ t Its dimensionality and complexity directly reflect the challenge of the original optimization problem.

[0093] Step 152: Based on the location parameters of each mobile terminal and the distance statistics between all mobile terminals and the smart reflective surface, obtain the state space.

[0094] Step 152 will be described in detail below.

[0095] In some embodiments, after defining the action space, it is necessary to define the environmental information on which the agent bases its decisions, i.e., the state space. Since the CSI is unknown, the state definition can only rely on known location information. Because the base station and RIS locations are fixed, the time-varying mobile terminal's location information is equivalent to its distance from the RIS. Furthermore, since large-scale fading depends on distance and small-scale fading is unknown, the mobile terminal's location determines the channel quality of the link between the RIS and the user. This step constructs and obtains the state space based on the location parameters of each mobile terminal and the distance statistics between all mobile terminals and the smart reflector. Specifically, the terminal location parameter Q = {q} of the mobile terminal at time step t is... k The state is defined as a subset of the state. Furthermore, to more effectively summarize the spatial distribution characteristics of the mobile terminal group and accelerate the agent's learning process, the state incorporates the distance statistics D from all mobile terminals to the RIS at time step t. R,t ={D R,1,t D R,2,t ,L,D R,K,t In this design, distance statistics specifically include the mean and variance of the distances from all mobile terminals to the RIS center, where K is the number of users. The core advantage of this design lies in its direct extraction of the implicit geometric location information of the mobile terminal group into explicit numerical features that summarize its spatial distribution, thereby significantly reducing the agent's workload. The complexity of exploration. Specifically, the original mobile terminal's terminal location Q. t High dimensionality, intelligent agent Neural networks require extensive trial-and-error training to learn to calculate and perceive the spatial distribution of users from this initial vector, a highly inefficient learning process. Introducing distance statistics is equivalent to... Feature preprocessing was performed. For example, the distance variance statistic directly and explicitly provides the key information of the dispersion of the mobile terminal group as an input feature. so, Instead of learning how to calculate dispersion, we can directly learn the direct mapping relationship between high-variance states and action spaces. This shift from learning how to calculate features to learning how to utilize existing features can significantly simplify the function complexity that neural networks need to fit.

[0096] Among these, the location parameters of each mobile terminal provide the most direct and fine-grained environmental information. Introducing the distance statistics (such as mean and variance) between all mobile terminals and the intelligent reflective surface is an efficient feature engineering method. It can condense the spatial distribution characteristics (such as density and dispersion) of the dispersed user group into low-dimensional numerical features with clear physical meaning. This helps the agent to learn and understand universal optimization strategies under different user layouts more quickly, thereby accelerating convergence and improving the model's generalization ability.

[0097] At the same time, the reward r from the previous time step t-1 Adding the current state allows for Provide direct feedback on the effects of the previous decision, r t-1 This represents the reward at time step (t-1). This makes... It can perceive dynamic changes in system performance, thereby making more predictive optimization decisions at the current time step. Therefore, at time step t, the state space is determined by the user position Q at time step t. t The distance statistic D from the user to RIS at time step t. R,t and the reward r at time step t-1 t-1 Composition, that is

[0098] Step 153: Using the average bit error rate model of the throughput model in the transmission parameter optimization problem, based on the time slot error symbols corresponding to the signal received by each mobile terminal in each time slot and the number of time slot symbols to be transmitted in each time slot, the estimated bit error rate reward model for each time step is obtained. The bit error rate reward model is used to reduce the received bit error rate of each mobile terminal.

[0099] Step 153 will be described in detail below.

[0100] In some embodiments, in order to guide the agent to learn the correct optimization direction, it is necessary to design a feedback signal that can quantify the quality of its behavior, i.e., a reward model. This step utilizes the average bit error rate model of the throughput model in the transmission parameter optimization problem (7) and combines it with an online simulation method to construct it. Specifically, it calculates a real-time feedback based on the time slot error symbols corresponding to the signal received by each mobile terminal in each time slot and the number of time slot symbols to be transmitted in each time slot. The time slot error symbols refer to the symbols that fail to be correctly decoded in an actual or simulated transmission. By counting the number of these error symbols and comparing them with the total number of transmitted symbols, an effective estimate of the current average bit error rate can be quickly obtained. Based on this estimate, the system can generate an estimated bit error rate reward model for each time step. The design goal of this model is to make its value positively correlated with the system performance. Therefore, the bit error rate reward model is used to reduce the received bit error rate of each mobile terminal, thereby indirectly achieving the ultimate goal of maximizing throughput. The estimated bit error rate reward model for the t-th time step is defined as shown in the following formula (8).

[0101]

[0102] In the estimated bit error rate reward model (8), an appropriate constant factor λ and a log(·) function are used to amplify the difference between the minimum throughput obtained by taking different actions. Since the CSI is unknown, the average bit error rate required for the reward function cannot be accurately calculated using the analytical expression. However, This can be obtained using a Monte Carlo method. The following section will further describe how to determine the estimated bit error rate reward model based on this.

[0103] The estimated bit error rate reward model for each time step is obtained based on the bit error symbols corresponding to the signals received by each mobile terminal in each time slot and the number of time slot symbols to be transmitted in each time slot, including the following steps 1531 to 1535.

[0104] Step 1531: Based on multi-level modulation technology, determine the number of transmission time slots corresponding to the number of time slot symbols.

[0105] Step 1532: Accumulate the number of time slot symbols with error corresponding to the number of time slot symbols in all transmission time slots, and then divide by the number of transmission time slots and the number of time slot symbols in turn to obtain the average symbol error rate corresponding to each time step.

[0106] Step 1533: Based on the average symbol error rate divided by the logarithm of the modulation order, the estimated bit error rate model is obtained.

[0107] Step 1534: Based on the logarithmic processing of the estimated bit error rate model, multiply by a negative amplification factor to obtain the terminal bit error rate reward model for each mobile terminal at each time step.

[0108] Step 1535: Based on the minimum value among all terminal bit error rate reward models, obtain the estimated bit error rate reward model for each time step.

[0109] Steps 1531 to 1535 are described in detail below.

[0110] To accurately assess communication performance over a period of time, the size of the observation window used to calculate the bit error rate (BER) must first be determined. Based on the multi-level modulation technique employed, namely M-ary modulation, the total number of transmission time slots required to transmit a certain number of time slot symbols is then determined. Under M-ary modulation, the base station transmits L symbols to all mobile terminals, where M is the modulation order, and the base station... Transmit J = T / (T) in each time slot s ) symbols. Therefore, in order to transmit V symbols, the base station at I = V / J = LT s Transmission is maintained in / (Tlog2M) time slots, that is, the total number of time slot symbols is I and the number of time slot symbols is V.

[0111] Furthermore, at the t-th time step, the j-th symbol received at the k-th user in the i-th time slot is represented as... The time slot error symbol is shown in the following formula (9).

[0112]

[0113] After completing the transmission for the preset duration, the system needs to statistically analyze the transmission results to calculate a preliminary performance indicator. This step involves accumulating the number of time slot errors corresponding to the total number of time slot symbols across all transmission time slots. Then, divide by the number of transmission time slots and the number of time slot symbols in turn to obtain the average symbol error rate for each time step. As shown in formula (10) below.

[0114]

[0115] This ratio is called the average symbol error rate (SER), which directly reflects the frequency of errors in transmitted symbols under the current parameter configuration.

[0116] To align with more general performance metrics, the average symbol error rate obtained in the previous step needs to be converted to the bit error rate (BER). Utilizing the characteristics of Gray coding used in communication systems—that one symbol error corresponds to one bit error—the estimated average symbol error rate is converted into an estimate of the average BER. Therefore, the estimated bit error rate model is obtained by dividing the average symbol bit error rate by the logarithm of the modulation order, as shown in formula (11).

[0117]

[0118] In communication systems employing Gray coding, a symbol error typically corresponds approximately to a bit error; therefore, a simple conversion can be used. The logarithm of the modulation order (log2(M), where M is the modulation order) represents the number of bits carried by each symbol. Dividing the average symbol error rate by this value yields an approximate estimate of the average bit error rate. This result is the estimated bit error rate model, a more general measure of communication reliability.

[0119] Understandably, when the number of bits sent... Can be considered a true value Therefore, as long as F is large enough, this application can reliably estimate the P of all mobile terminals according to the bit error rate estimation model (11). b,k (Q t This allows the agent to receive a reward, r. t .

[0120] After obtaining the estimated bit error rate model, it needs to be transformed into a reward signal suitable for reinforcement learning training. This step is based on the logarithmic processing of the estimated bit error rate model, multiplied by a negative amplification factor, to obtain the terminal bit error rate reward model for each mobile terminal at each time step, as shown in formula (8) above. Directly using the bit error rate as the reward may lead to slow learning due to its small value and insignificant changes. Therefore, logarithmic processing can amplify the differences between different bit error rates, while multiplying by a negative amplification factor transforms the goal of minimizing the bit error rate into the goal of maximizing the reward, which is in line with the optimization habits of reinforcement learning algorithms. The terminal bit error rate reward model obtained after this processing can provide the agent with a clearer and more discriminative learning signal.

[0121] As shown in formula (8) above, after calculating the individual rewards for each mobile terminal, these individual rewards need to be integrated into a single reward value representing the overall performance of the system. This step is based on the minimum value among all terminal bit error rate reward models to obtain the estimated bit error rate reward model for each time step. The motivation for adopting this "minimum value" strategy is consistent with the Max-Min fairness criterion in the original optimization problem. This means that the overall reward of the system will be determined by the user with the worst performance. This design can strongly drive the agent to focus on and improve the performance of the weakest user, thereby avoiding the unfair phenomenon of some users having excellent performance while others have extremely poor performance, ensuring the balance and robustness of the overall service quality of the system.

[0122] Through the implementation of steps 1531 to 1535 above, starting with determining the amount of resources required for evaluation, the average symbol error rate is calculated by statistically analyzing actual bit error rates and converted into a more general estimated bit error rate model. Subsequently, through mathematical transformations, this model is shaped into a terminal reward suitable for reinforcement learning and possessing an amplifying effect. Finally, by employing a "minimum value" aggregation strategy, the rewards from multiple users are unified into a single global reward signal. This series of steps is logically clear and has well-defined objectives, providing a high-quality, effective feedback mechanism for deep reinforcement learning agents that guides them towards optimizing towards the goal of maximizing system fairness and throughput.

[0123] Step 154: Obtain the initial action parameters of the actor network and the initial comment parameters of the commenter network.

[0124] Step 155: Generate an initial transmission parameter optimization model based on the action space, state space, bit error rate reward model, initial action parameters, and initial comment parameters.

[0125] Steps 154 to 155 are described in detail below.

[0126] As mentioned earlier, the MDP action space corresponding to the transmission parameter optimization problem (7) is a hybrid action space composed of continuous beamforming parameters, time slot allocation parameters, and discrete RIS phase shift parameter sequences. Specifically, the number of discrete actions of RIS phase shift increases exponentially with the number of reflection units, resulting in a high dimensionality of the discrete action space; at the same time, the continuous action spaces of beamforming and time slot allocation parameters also have high dimensionality. This complex action space characteristic makes it difficult to apply traditional reinforcement learning methods. For example, value function methods represented by Deep Q-Networks (DQN) are designed to handle low-dimensional, discrete action spaces and cannot directly handle continuous actions. They are also difficult to optimize when facing high-dimensional discrete spaces due to the huge computational cost. Therefore, in order to effectively address the above challenges, this application proposes a PPO method based on action combination to solve this problem.

[0127] The following is an introduction to the PPO algorithm.

[0128] Unlike the value function-based DQN method, the PPO algorithm employs an actor-critic architecture. This architecture consists of an actor network for generating the optimization policy and a critic network for evaluating the quality of states, with parameters denoted by μ and μ, respectively. This indicates that the policy π is updated using a gradient-based method. μ Then update the state value function. First, assume that the state value model of the agent is as shown in the following formula (12).

[0129]

[0130] This function represents the state s t Below, follow strategy π μ The expected long-term reward that can be obtained. Then the optimization problem of PPO can be shown in the following formula (13).

[0131]

[0132] in, It is a new strategy π μ (a t |s t ) and old strategies The probability ratio between them, also known as the importance weight, μ old These are the policy network parameters before the update; It is the advantage function, used to evaluate the performance in state s. t Take action a t The degree of superiority or inferiority compared to the average level. In this application, the advantage function... It is represented by the time-series difference residual, as shown in the following formula (14).

[0133]

[0134] To further improve the stability of the training process and prevent performance crashes caused by excessively large policy update steps, this application adopts the PPO clipping mechanism. This mechanism clips the importance weights... By limiting it to a preset safety range [1-ε, 1+ε], the optimization objective of equation (13) becomes the following equation (15).

[0135]

[0136] Among them, the function Defined as clip(x,l,u) = max(l,min(x,u)), it is used to restrict x to the range [l,u], ε is a hyperparameter controlling the clipping range, and H(π) μ (·|s t The policy entropy, with coefficient β, is used to encourage the agent. Conduct thorough exploration.

[0137] In the algorithm implementation, this application employs mini-batch stochastic gradient descent to update the network parameters. Specifically, each time, G is randomly sampled from the empirical replay pool. B records {s t ,a t ,s t+1 ,r t |t=1,K,G B The actor and commentator networks are updated separately. The actor network is updated by maximizing L in equation (15). CLIP (s t ,a t ,μ) to update parameter μ, as shown in formula (16) below.

[0138]

[0139] Where, α μ Let represent the learning rate of the actor network. For the commenter network, its parameters are updated by minimizing the mean squared error. As shown in formula (17) below.

[0140]

[0141] in, This represents the learning rate of the commenter network and allows... The function representing the target state value.

[0142] Furthermore, the action space constructed based on the aforementioned transmission parameter optimization problem (7) contains high-dimensional continuous actions and large-scale discrete actions. This high-dimensional action space makes it difficult for traditional reinforcement learning methods to solve. To address this technical challenge, this application adopts an action combination method based on the PPO algorithm. That is, by combining some independent actions in the original MDP problem into a composite action, the dimensionality of the action space is reduced, and the output size of the actor network is correspondingly reduced. Specifically, in the MDP problem of this application, let ψ t ={θ Q,n,t |n∈M}, that is, treating the discrete phase shift of each reflection unit of RIS as a sub-action, and ψ tIt is a combination of all these sub-actions. With this design, the actor network no longer needs to output a specific phase shift value for each RIS phase shift sub-action, but instead outputs a probability distribution vector as shown in the following formula (18).

[0143] p(s t )={p1(s t ),p2(s t ),…,p M (s t (18)

[0144] Where, p m (s t )={p m,1 (s t ),…,p m,D (s t )},p m,d (s t ) indicates that in state s t The probability that the phase shift of the m-th reflection unit takes the d-th discrete value. That is, the actor network will output the policy. Where a t ={ω t ,p(s t ),τ t}

[0145] Subsequently, the intelligent agent Based on this probability distribution p(s) t By randomly sampling the phase shift of each reflecting element, the phase shift value θ of the m-th element can be obtained. Q,m,t By aggregating the phase shift values ​​of all units, the complete RIS phase shift matrix can be obtained.

[0146] Due to the phase shift θ of the m-th reflecting unit Q,m,t There are 2 possible values. D There are M actions in total, therefore the number of neurons in the discrete action output layer of the actor network is reduced from the original 2 DM Reduced to 2 after reconstruction D M, thus greatly simplifies the network structure and improves training efficiency. Therefore, the proposed PPO algorithm can efficiently solve the MDP problem.

[0147] Furthermore, in this application, a cosine annealing learning rate optimization strategy is introduced to improve the smoothness and stability of PPO training. This strategy enables a smooth transition of the learning rate from high to low, and by periodically resetting the learning rate, it helps the model escape local optima and enhances its global search capability. Specifically, let η... t The learning rate is the number of training steps t. Cosine annealing adjusts the learning rate using the following formula (19).

[0148]

[0149] Where T is the length of a complete cosine annealing cycle, and η max The maximum value that can be reached within one cosine annealing cycle is also the initial value, η. min It is the minimum value that can be achieved within one annealing cycle.

[0150] Furthermore, the accuracy of the value function estimation is crucial to the overall performance of the PPO algorithm, as it is directly used to evaluate the merits of the current strategy and to calculate the advantage function. Let... Parameters of the current commenter network For state s t Value estimation. Value estimation is crucial during the initial training phase or when the environment dynamically changes. Significant errors may exist, leading to a substantial difference between the current policy and the target value. If squared error loss is applied directly, these large errors will cause drastic updates to the value network parameters, affecting training stability. One of the core ideas of the PPO algorithm is to limit the step size of policy updates by pruning the objective function, thus preventing the old and new policies from deviating too much.

[0151] Similarly, this application prunes the value function to limit its variation in a single update, thereby stabilizing the learning process of the value function. Specifically, let... For state s t At that time, by the old commentator network parameters The obtained value estimate. Let R be... t Let R be the target value, and R be the target value. t Add the old value estimate to the advantage function estimate. The standard value function loss (mean squared error) is... To achieve the pruning of the value function, we first define a pruned value estimate V. clip (s t This cropping limits the new value estimate. Relative to the old value estimate The range of variation is shown in the following formula (20).

[0152]

[0153] Here, ε is a hyperparameter that controls the clipping range.

[0154] Therefore, in order to achieve complex decision-making and evaluation functions, it is first necessary to obtain the initial action parameters μ of the actor network and the initial comment parameters of the commenter network. The actor network, a key technical component, outputs an action policy based on the current state; while the commentator network is responsible for evaluating the merits of the actor's chosen action. Initial action parameters and initial commentator parameters typically refer to the random weights and biases of these neural networks before training begins, providing an initial starting point for subsequent learning and iterative optimization.

[0155] Then, all the core elements defined in the preceding steps are integrated. Specifically, this step, based on the established action space (defining what the model can do), state space (defining what the model can see), and bit error rate reward model (defining the learning objective and direction), combines these abstract definitions with concrete neural network entities (characterized by initial action parameters and initial comment parameters) to ultimately generate an initial transfer parameter optimization model. This generated initial transfer parameter optimization model is a fully structured, ready but untrained deep reinforcement learning agent that already possesses all the necessary components to autonomously learn optimization strategies through interaction with the environment.

[0156] Through the implementation of steps 151 to 155 above, by precisely defining the action space, cleverly designing the state space containing statistical features, and constructing a reward model based on online simulation, the original problem was successfully deconstructed into a standard framework for reinforcement learning. At the same time, by initializing the actor and commentator networks, the foundation for the model's learning process was laid, paving the way for obtaining an intelligent decision-making model that can perform global joint optimization without CSI through efficient training. This demonstrates a clear logical hierarchy and a powerful problem transformation capability.

[0157] The following section will further describe how to train the optimization model for this initial transmission parameter.

[0158] The process of training the initial transmission parameter optimization model multiple times to obtain the transmission parameter optimization model includes steps 156 to 159.

[0159] Step 156: In each training round, input the current training position parameters of each mobile terminal in the current training time step into the actor network to obtain the training action combination, and sample the current training action from the training action combination. Adjust the base station and smart reflector based on the current training action.

[0160] Step 157: In the current training time step, the control base station sends a training signal to each mobile terminal in each time slot, rewards the current bit error rate corresponding to the training signal received by each mobile terminal, and obtains the updated training position of each mobile terminal in the next training time step.

[0161] Step 158: Store the current training position, current training action, current bit error rate reward, and training experience information of the current training time step at the updated training position in the experience replay pool.

[0162] Step 159: Update the initial action parameters and initial comment parameters using the training experience information from the batch training time steps in the experience replay pool to obtain the updated actor network and commenter network, and use the actor network after multiple updates as the transfer parameters to optimize the model.

[0163] Steps 156 to 159 are described in detail below.

[0164] In some embodiments, the process is based on the PPO algorithm mentioned above. The training process begins with the interaction between the agent and the environment. In each round of training, the agent first observes the environmental state, that is, the current training position parameter q of each mobile terminal in the current training time step. k,i ={x k,i ,y k,i ,z k,i Input the actor network. After receiving the location information, the actor network will output a training action combination. This is typically a probability distribution describing the various possible actions. The agent then learns from this training action combination. Random sampling is performed to obtain a specific current training action a. t This exploratory sampling method is key to reinforcement learning, ensuring that the agent does not prematurely converge to a local optimum. Finally, based on this selected current training action, the system synchronously adjusts the base station (beamforming and timing) and the smart reflector (phase shift configuration), thereby translating the decision into actual operations in the physical world.

[0165] After performing action a t Afterwards, the environment needs to provide feedback; this step describes the "feedback" phase of the interaction. In the current training time step, the system controls the base station to send a preset training signal (such as Lbits data) to each mobile terminal in each time slot. After decoding these training signals, the receiver calculates the corresponding current bit error rate reward r based on its bit error rate using the above formula (8). t This reward directly reflects the quality of the action taken in the previous step. Simultaneously, to simulate a dynamic environment, the system obtains the updated training position (i.e., updated state s) for each mobile terminal in the next training time step. t+1This represents the natural evolution of the environmental state. Through this step, the agent not only receives an immediate evaluation of its actions but also observes the next state of the environment, thus completing a full interactive cycle of "state-action-reward-new state".

[0166] To fully utilize the information generated by each interaction, this valuable experience data needs to be recorded and stored. This step involves recording and storing the current training position (i.e., the current state s). t ), Current training action a t Current error rate reward r t And updating the training position (i.e., updating the state s) t+1 These four core pieces of information—state, action, reward, and next state—are packaged and stored as training experience information for the current training time step. Specifically, the system will package and store this four-tuple of experience information. t ,a t ,s t+1 ,r t}, stored in a data structure called the experience replay pool. The experience replay pool is a fixed-capacity first-in-first-out queue that accumulates a large amount of historical interaction data, providing rich and diverse training samples for subsequent model parameter updates. This breaks the temporal correlation between data and improves the stability and efficiency of training.

[0167] After accumulating sufficient empirical data, the system enters the model parameter learning and updating phase. This step utilizes training experience information from batch training time steps in the experience replay pool to update the initial action and comment parameters. Specifically, the system randomly selects a small batch of data from the experience replay pool and uses this data to calculate the loss function and gradient. Then, it updates the parameters of the actor and commenter networks using optimization algorithms such as gradient descent. This update process is repeated multiple times, allowing the network to learn the optimal strategy from a large amount of historical experience. After sufficient training and updates, the parameters of the actor network will gradually converge to the optimal or suboptimal state. At this point, the system uses this well-trained and updated actor network as the final transmission parameter optimization model, which can be directly used for online deployment.

[0168] The process of updating the initial action parameters and initial comment parameters using the training experience information from the batch training time steps in the experience replay pool to obtain the updated actor network and commenter network includes steps 1591 to 1595.

[0169] Step 1591: Based on the state value model, obtain the training state value corresponding to the training experience information. The state value model is generated based on the estimated bit error rate reward model.

[0170] Step 1592: Obtain the training time-series difference residual and training state value based on training experience information and training state value.

[0171] Step 1593: Obtain the importance weight of the update strategy, and truncate the importance weight of the update strategy based on the preset truncation parameter to obtain the truncated importance weight.

[0172] Step 1594: Update the initial action parameters using gradient descent based on the truncation importance weights.

[0173] Step 1595: Update the initial comment parameters using gradient descent based on the training state values.

[0174] Steps 1591 to 1595 are described in detail below.

[0175] In some embodiments, to evaluate the quality of historical actions, it is first necessary to estimate the long-term value of each state. This step is based on a state value model (12) to obtain the training state value corresponding to each state in the training experience information. This state-value model, known as the commentator network in the actor-commentator architecture, learns from a large amount of historical data to assign a value score to any input state. This score predicts the sum of future cumulative rewards obtainable from that state by following the current policy. Since the training objective of the commentator network is to predict the reward stream generated by the estimated bit error rate reward model as accurately as possible, it can be said that the state-value model is generated based on the estimated bit error rate reward model, providing a benchmark for subsequent calculations of action advantage.

[0176] After obtaining the state value estimate, the system will calculate a more refined action evaluation metric, namely the advantage function. This step is based on training experience information {s}. t ,a t ,s t+1 ,r t} and the training state value obtained in the previous step This is used to obtain the training time-series difference residuals and the training state values. Also known as the advantage value, it is calculated by subtracting the current state's training state value from the sum of the immediate reward and the training state value of the next state. This residual value intuitively measures whether taking a particular action in the current state is better or worse than the average level (i.e., the state value). Simultaneously, the system also calculates the target value used to update the commentator network, i.e., the training state value, which is typically composed of the immediate reward and the value estimate of the next state, representing a more accurate estimate of the current state's value.

[0177] To ensure the stability of policy updates, effective control of the update step size is necessary, which is the core of the Proximal Policy Optimization (PPO) algorithm. This step first obtains the importance weights of the updated policy. This weight is the ratio of the output probability of the new policy (updated actor network) to that of the old policy (previous actor network) for the same action. Then, the system will truncate the importance weight of the updated policy based on a preset truncation parameter ε (i.e., using formula (15)), that is, restrict it to a small interval (such as [1-ε, 1+ε]), thereby obtaining the truncated importance weight L. CLIP (s t ,a t This truncation process effectively prevents policy performance from collapsing due to excessively large single update steps, and is a key mechanism to ensure stable convergence of the PPO algorithm.

[0178] After calculating all necessary intermediate values, the system updates the actor network. This step is based on the truncated importance weights L. CLIP (s t ,a t (μ) and training time-series difference residuals The initial action parameter μ is updated using gradient descent based on the advantage value. Specifically, the objective function of the actor network, as shown in equation (16) above, is a function related to the truncation importance weights and the advantage value. By calculating the gradient of this objective function with respect to the actor network parameters and making a small update along the gradient direction, the probability of actions that bring positive advantage values ​​can be increased, while the probability of actions that bring negative advantage values ​​can be suppressed. In this way, the actor network's strategy will gradually optimize in the direction of obtaining higher rewards.

[0179] While updating the actor network, the commentator network responsible for value assessment also needs to be updated synchronously to keep up with policy changes and provide more accurate value estimates. This step is based on the training state value V. tar (s t ), to adjust the initial comment parameters Gradient descent updates are performed. The optimization objective of the commentator network is usually to minimize the mean squared error between its output value estimate and the more accurate training state value. As shown in Equation (17) above, by calculating the gradient of this error with respect to the commentator network parameters and performing gradient descent updates, the commentator network can make its evaluation of state value more and more accurate, thereby providing a more reliable benchmark and advantage estimate for the update of the actor network.

[0180] Through the implementation of steps 1591 to 1595 above, stable iteration and optimization of the policy are achieved by collaboratively updating the actor and commentator networks. First, the commentator network is used to calculate the training temporal difference residuals used to evaluate the quality of actions. Then, by introducing and truncating importance weights—a core technique—the update step size of the actor network is effectively constrained, ensuring the smoothness of the training process. Finally, more accurate training state values ​​are used to optimize the commentator network, forming a mutually reinforcing and stably convergent learning loop. This series of steps constitutes the robust and efficient theoretical foundation of the PPO algorithm and is the key to the successful training of a high-performance optimization model in this invention.

[0181] Based on the above description, the following is a schematic diagram of the training algorithm for the initial transmission parameter optimization model provided in this application.

[0182] Algorithm 1 is a PPO-based optimization algorithm for CSI-free intelligent reflective surface-assisted communication.

[0183] Input experience and return a value of G. B Reward discount factor γ and learning rate α μ and

[0184] Initialize the Actor network parameters μ and the Critic network parameters. Experience replay pool with capacity C and t cnt =0.

[0185] Output the Actor network parameters μ.

[0186]

[0187] Reference Figure 3 This is a schematic block diagram illustrating the training of an initial transmission parameter optimization model provided in an embodiment of this application. For example... Figure 3 As shown in the figure, at time step t, the agent observes the location information of all mobile terminals as the state s of the current time step. t , will state s t Input an Actor network, and the Actor network outputs the current policy. The agent follows the probability distribution of the policy. Sampling yields specific action a t and action a t Send to the RIS controller, the RIS controller and the base station according to action a t The phase shift, time slot allocation, and beamforming of each reflective element in the RIS are adjusted. Then, with the assistance of the RIS, the agent transmits signals to each mobile terminal. Each mobile terminal calculates the average bit error rate and its own reward r. k,t Feedback is given to the agent, and then the agent moves to the next position. The agent calculates the reward r according to equation (8). t And observe the next location information of all mobile terminals to obtain the state s t+1 After an interaction at a time step is completed, the agent saves the interaction data for the current time step {(s)}. t ,a t ,r t ,s t+1 )} and the probability of old moves And the state-value function is estimated using a Critic network. Calculate and save the target state value function V at each time step. t target and the dominant function A t When the agent interacts with the environment... step After one time step, use the saved t step Train a neural network with n data points and reuse these n data points. epoch Round: Each time, a small batch of data is randomly sampled, and then the probability ratio r is calculated. t (μ), trimming replacement target L CLIP (μ), error term of the value function And entropy regularization term L ENT (μ), and the total loss is obtained by weighted summation. Finally, the Adam optimizer is used to update the neural network parameters μ and When n epoch After each round of training, sampling continues using the updated network parameters until the network converges.

[0188] Through the implementation of steps 156 to 159 above, and through the closed-loop iteration of "action-feedback-storage-learning", the agent continuously learns and makes mistakes in its continuous interaction with the environment. By adopting the experience replay pool mechanism, the data utilization and training stability are effectively improved. Finally, through the collaborative updating of the actor network and the commentator network, the model can learn from scratch how to make the globally optimal decision that maximizes the system's fairness throughput based solely on user location information, thus obtaining a robust, efficient, and CSI-free intelligent optimization model.

[0189] Through the implementation of steps 110 to 150 above, an accurate physical channel and performance model is first established. Then, a transmission parameter optimization problem that balances efficiency and fairness is defined. Finally, through deep reinforcement learning, this complex optimization problem is transformed into an intelligent model that can make end-to-end decisions based on user location. This allows the final transmission parameter optimization model to make optimization decisions directly based on learned deep patterns without relying on real-time channel state information. This not only greatly reduces system overhead and complexity but also ensures robust and fair global resource optimization in dynamic multi-user scenarios.

[0190] Step 200: Input all terminal locations into the transmission parameter optimization model for data matching to obtain the output optimized time slot allocation, optimized beamforming vector, and optimized phase shift sequence corresponding to multiple time slots.

[0191] Step 200 is described in detail below.

[0192] In some embodiments, after acquiring the terminal location and the optimization model, the system executes the model's inference and decision generation process. Specifically, the system integrates all acquired terminal location information into a state vector and feeds it as input to the transmission parameter optimization model. Upon receiving the input, the model performs a rapid forward propagation calculation, a process known as data matching, which aims to calculate the optimal action strategy in real time based on the current terminal location state. The output of the model is a set of jointly optimized transmission parameters that comprehensively cover all the key configurations required for efficient communication in this time step. These parameters include: an optimized time slot allocation that determines the proportion of communication time each mobile terminal can occupy; an optimized beamforming vector that guides the base station antenna array to adjust the phase and amplitude of the transmitted signal, thereby accurately focusing the signal energy onto the target user; and an optimized phase shift sequence set for each reflector element in the intelligent reflector in different time slots to intelligently control the signal reflection path.

[0193] Step 300: Allocate corresponding time slots to multiple mobile terminals in the current time step according to the optimized time slot allocation; adjust the beamforming vector between the base station and each mobile terminal in the current time step according to the optimized beamforming vector; and adjust the reflection unit of the smart reflector of each time slot in the current time step according to the optimized phase shift sequence.

[0194] Step 300 is described in detail below.

[0195] In some embodiments, after the system calculates the complete set of optimized parameters based on the model, it enters the physical application and execution phase of the parameters. The network controller or scheduler within the system strictly adheres to the optimized time slot allocation scheme, assigning dedicated communication time slots to multiple mobile terminals in the current time step. Simultaneously, the base station's signal processing unit adjusts the signal weights applied to each mobile terminal in the current time step according to the optimized beamforming vector, thereby forming customized high-gain beams pointing towards target users in different time slots at the physical layer. Furthermore, the smart reflector controller also performs synchronized phase shift state configuration of the reflector units of the smart reflector in each time slot in the current time step according to the optimized phase shift sequence. Through this series of synchronized adjustments to temporal, spatial, and frequency domain resources, the virtual decisions output by the model are efficiently transformed into actual configurations on various physical devices in the communication system.

[0196] In addition, to verify the reliability of the time-division multiple access communication subway transmission parameter optimization method provided in this application, this embodiment performs algorithm complexity analysis.

[0197] The optimization scheme based on PPO proposed in this application has an algorithm complexity mainly consisting of two stages: prediction and training.

[0198] 1) Prediction Complexity: The system needs to generate actions in real time based on the current state, and its computational complexity mainly stems from the forward propagation of the actor network. Specifically, assume the actor network contains L... A Layers, with N neurons per layer. A,i The prediction complexity is shown in formula (21).

[0199]

[0200] 2) Training Complexity: Training complexity consists of several parts. The core overhead comes from updating the parameters of both the actor and commentator networks, a process based on the backpropagation algorithm. Assume the commentator network has L... C Layers, actor networks have L A The number of neurons in the 'tanh' activation layer in the two networks are respectively and If the backpropagation of a single 'tanh' neuron requires 6 floating-point operations, then the complexity of the process is shown in Equation (22).

[0201]

[0202] The training process itself requires a prediction step, the complexity of which has been described above. Furthermore, in the sixth step of the algorithm in this application, L bits of information need to be broadcast to the user to calculate the reward; the complexity of this step is denoted as... Therefore, the total complexity of the entire training process can be summarized as shown in the following formula (23).

[0203]

[0204] When training the PPO algorithm in a real physical system, C delay It can be measured by the maximum delay of transmitted bits, and is calculated as shown in the following formula (24).

[0205]

[0206] Where R and c represent the system's transmission rate and the speed of light, respectively, and d TR This represents the straight-line distance from the base station (BS) to the RIS. This represents the straight-line distance from RIS to the k-th user.

[0207] In summary, through the implementation of the above steps, the technical solution provided by this invention achieves end-to-end joint optimization of time slot allocation, base station beamforming, and intelligent reflector phase shift in a time-division multiple access communication system using only the terminal location information of the mobile terminal. This method completely avoids the dependence on channel state information (CSI) in traditional schemes, thus fundamentally solving the technical problems of high signaling overhead, computational delay, and resource consumption caused by complex channel estimation. Since the decision model directly outputs the final joint optimization parameters, the system can achieve fast and reliable resource allocation, exhibiting superior robustness and adaptability, especially in dynamic scenarios such as high-speed user movement. Ultimately, while ensuring fairness among users, it significantly improves the throughput and spectral efficiency of the entire wireless communication system.

[0208] This application also provides an electronic device, including:

[0209] At least one memory;

[0210] At least one processor;

[0211] At least one program;

[0212] The program is stored in memory, and the processor executes at least one program to implement the transmission parameter optimization method for the time division multiple access communication system described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0213] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0214] The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0215] The memory 402 can be implemented in the form of ROM (Read-Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 402 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402, and the processor 401 calls and executes the transmission parameter optimization method of the time division multiple access communication system of the embodiments of this application.

[0216] Input / output interface 403 is used to implement information input and output;

[0217] The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0218] Bus 405 transmits information between various components of the device (e.g., processor 401, memory 402, input / output interface 403, and communication interface 404);

[0219] The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.

[0220] This application embodiment also provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the above-described method for optimizing transmission parameters of a time-division multiple access communication system.

[0221] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In some embodiments, memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0222] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for optimizing transmission parameters of a time division multiple access communication system, characterized by, The wireless communication system includes a base station, a smart reflecting surface and a plurality of mobile terminals, and the method comprises: obtaining the terminal positions of all the mobile terminals at the current time step, and obtaining a transmission parameter optimization model, the time step comprising a plurality of time slots; inputting all the terminal positions into the transmission parameter optimization model for data matching to obtain an output of an optimized time slot allocation, an optimized beamforming vector and an optimized phase shift sequence corresponding to the plurality of time slots; allocating corresponding time slots to the plurality of mobile terminals in the current time step according to the optimized time slot allocation, adjusting the beamforming vector between the base station and each mobile terminal in the current time step according to the optimized beamforming vector, and adjusting the reflecting units of the smart reflecting surface in each time slot in the current time step according to the optimized phase shift sequence.

2. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 1, characterized in that, The transmission parameter optimization model comprises: based on the phase shift parameter sequence of the smart reflecting surface and the position parameter of the plurality of mobile terminals, generating a corresponding cascade channel model of each mobile terminal in each time slot; based on the beamforming parameter, the time slot modulation symbol, the noise power, and the cascade channel model corresponding to each time slot, obtaining an average bit error rate model; based on a first numerical value minus the average bit error rate model, multiplied by the time slot allocation parameter, obtaining a throughput model of each mobile terminal; based on maximizing the minimum throughput model of all mobile terminals, obtaining a transmission parameter optimization problem; based on the transmission parameter optimization problem, constructing an initial transmission parameter optimization model, and training the initial transmission parameter optimization model for multiple rounds to obtain the transmission parameter optimization model.

3. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 2, characterized in that, Based on the phase shift parameter sequence of the smart reflecting surface and the position parameter of the plurality of mobile terminals, generating a corresponding cascade channel model of each mobile terminal in each time slot, comprising: based on the position parameter of each mobile terminal, determining a first reflection channel model between each mobile terminal and the smart reflecting surface; based on the product of the first reflection channel model, the phase shift parameter sequence and a second reflection channel model, obtaining the cascade channel model, the second reflection channel model being a channel model between the smart reflecting surface and the base station.

4. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 2, wherein, Based on the beamforming parameter, the time slot modulation symbol, the noise power, and the cascade channel model corresponding to each time slot, obtaining an average bit error rate model, comprising: based on the product of the cascade channel model, the beamforming parameter and the time slot modulation symbol, and adding the noise power, obtaining a received signal model of each mobile terminal; based on the expected value of the received signal model, subtracting the product of the noise power and the system bandwidth, and then sequentially dividing by the logarithm of the noise power, the system bandwidth and the modulation order, obtaining a corresponding received signal-to-noise ratio model of each mobile terminal in each time slot; based on the product of the transmission power corresponding to the information sent by the base station to each mobile terminal in each time slot and the received signal-to-noise ratio model, and then performing expectation processing, obtaining the average bit error rate model.

5. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 2, wherein, The initial transmission parameter optimization model is constructed based on the transmission parameter optimization problem, comprising: Based on the time slot allocation parameter, the beamforming parameter and the phase shift parameter sequence in the transmission parameter optimization problem, an action space is obtained; Based on the position parameter of each mobile terminal and the distance statistical value between all mobile terminals and the intelligent reflecting surface, a state space is obtained; Based on the average bit error rate model of the throughput model in the transmission parameter optimization problem, an estimated bit error rate reward model of each time step is obtained based on the time slot error code symbol corresponding to the received signal of each mobile terminal at each time slot, and the time slot symbol quantity to be transmitted in each time slot, the bit error rate reward model being used to reduce the receiving bit error rate of each mobile terminal; An initial action parameter of an actor network and an initial comment parameter of a commentator network are obtained; The initial transmission parameter optimization model is generated based on the action space, the state space, the bit error rate reward model, the initial action parameter and the initial comment parameter.

6. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 5, characterized in that, The estimated bit error rate reward model of each time step is obtained based on the time slot error code symbol corresponding to the received signal of each mobile terminal at each time slot, and the time slot symbol quantity to be transmitted in each time slot, comprising: Based on the multi-ary modulation technology, the transmission time slot number corresponding to the time slot symbol quantity is determined; All time slot error code symbols corresponding to all time slot symbol quantities in all transmission time slots are accumulated, and then divided by the transmission time slot number and the time slot symbol quantity in turn to obtain the average symbol bit error rate corresponding to each time step; Based on the average symbol bit error rate divided by the logarithm of the modulation order, an estimated bit error rate model is obtained; Based on the logarithm processing of the estimated bit error rate model, and then multiplied by a negative amplification factor, a terminal bit error rate reward model corresponding to each mobile terminal at each time step is obtained; Based on the minimum value in all terminal bit error rate reward models, the estimated bit error rate reward model of each time step is obtained.

7. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 5, wherein, The transmission parameter optimization model is obtained by training the initial transmission parameter optimization model for multiple rounds, comprising: In each round of training, the current training position parameter of each mobile terminal in the current training time step is input into the actor network to obtain a training action combination, and a current training action is sampled from the training action combination, and the base station and the intelligent reflecting surface are adjusted based on the current training action; In the current training time step, the base station is controlled to send a training signal to each mobile terminal in each time slot, a current bit error rate reward corresponding to each mobile terminal receiving the training signal is obtained, and an updated training position of each mobile terminal in the next training time step is obtained; The current training position, the current training action, the current bit error rate reward and the updated training position are training experience information of the current training time step, and the training experience information is stored in an experience replay pool. The initial action parameter and the initial comment parameter are updated by using the training experience information of the batch training time steps in the experience replay pool to obtain an updated actor network and a critic network, and the updated actor network is used as the transmission parameter optimization model.

8. The method of optimizing transmission parameters of a time division multiple access communication system according to claim 7, characterized in that, The updating of the initial action parameter and the initial comment parameter by using the training experience information of the batch training time steps in the experience replay pool to obtain an updated actor network and a critic network comprises: Based on a state value model, a training state value corresponding to the training experience information is obtained, and the state value model is based on the estimated bit error rate reward model; Based on the training experience information and the training state value, a training time difference residual and a training state value are obtained; An update strategy importance weight is obtained, and the update strategy importance weight is truncated based on a preset truncation parameter to obtain a truncated importance weight; The initial action parameter is updated by gradient descent based on the truncated importance weight; The initial comment parameter is updated by gradient descent based on the training state value. 9.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to implement the transmission parameter optimization method of the time division multiple access communication system in any one of claims 1 to 8.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the transmission parameter optimization method of the time division multiple access communication system in any one of claims 1 to 8.