Transmission parameter optimization method of multi-carrier communication system and related equipment

By utilizing real-time user location information and deep reinforcement learning models to optimize subcarrier allocation and reflector phase shift in multi-carrier communication systems, the problem of low reliability in resource optimization in intelligent reflector systems is solved, achieving more efficient transmission performance and user experience assurance.

CN121240108APending Publication Date: 2025-12-30GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511233613.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In intelligent reflector-assisted multicarrier communication systems, existing technologies rely on transient channel state information that is difficult to obtain and has a high overhead, resulting in low reliability of resource optimization decisions. Especially in scenarios where users move and channels change rapidly over time, frequent channel estimation can make the system overhead unbearable, affecting the reliability of the communication system and the user experience.

Method used

By acquiring real-time user location information and utilizing a pre-trained deep reinforcement learning model, subcarrier allocation, intelligent reflector phase shift sequence, and power allocation are optimized, avoiding complex channel estimation processes and dynamically adjusting transmission strategies to adapt to changes in user location.

Benefits of technology

It significantly reduces the complexity of resource optimization, improves the transmission performance of the system, and ensures communication fairness and service quality for mobile users. In particular, it can reliably improve the overall performance of the communication system even without the need to obtain instantaneous channel state information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121240108A_ABST
    Figure CN121240108A_ABST
Patent Text Reader

Abstract

According to the transmission parameter optimization method and related equipment of the multi-carrier communication system, the multi-carrier communication system comprises a base station, an intelligent reflecting surface and a plurality of mobile terminals, the method comprises the steps that terminal positions of all the mobile terminals in the current time step are acquired, a transmission parameter optimization model is acquired, and the time step comprises a plurality of time slots; inputting all terminal positions into a transmission parameter optimization model for data matching to obtain output optimized subcarrier allocation, optimized phase shift sequences corresponding to a plurality of time slots and optimized subcarrier power; and allocating corresponding subcarriers to the plurality of mobile terminals according to the optimized subcarrier allocation, adjusting a reflection unit of the intelligent reflection surface of each time slot in the current time step according to the optimized phase shift sequence, and adjusting power allocation of each subcarrier in the current time step according to the optimized subcarrier power. The transmission performance of the whole multi-carrier communication system can still be reliably improved under the condition that instantaneous channel state information does not need to be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technology, and in particular to methods and related equipment for optimizing transmission parameters of multi-carrier communication systems. Background Technology

[0002] In recent years, in next-generation multi-carrier communication systems, intelligent reflector (RIS) technology has received widespread attention to meet the growing demand for high-speed and high-reliability communication and to meet the increasing number of users. RIS, through a large number of low-cost passive reflective elements, can intelligently regulate the propagation environment of wireless signals, thereby enhancing signal coverage and increasing system capacity. In such systems where RIS assists base stations in communicating with users, maximizing system communication performance requires joint optimization of system resources. These resources typically include: the phase shift configuration of the RIS, the communication power between the base station and the user, and so on.

[0003] In related technologies, in communication systems equipped with intelligent reflectors, system resources are typically optimized at the beginning of each time slot based on the precise Channel State Information (CSI) corresponding to the current time slot. However, since intelligent reflectors usually contain hundreds or thousands of reflective elements, accurately obtaining the channel state information for each link requires a large number of pilots and complex estimation algorithms. This consumes valuable time and frequency resources and introduces significant computational delays. Especially in scenarios where users move and channels change rapidly over time, frequent channel estimation can make the system overhead unbearable, resulting in low reliability of resource optimization decisions in actual communication systems equipped with intelligent reflectors. Summary of the Invention

[0004] This application provides a method and related equipment for optimizing transmission parameters in a multi-carrier communication system, which can improve the reliability of resource parameter optimization between the base station and the mobile terminal in a multi-carrier communication system equipped with an intelligent reflector.

[0005] To achieve the above objectives, a first aspect of this application proposes a method for optimizing transmission parameters in a multi-carrier communication system, the multi-carrier communication system including a base station, a smart reflector, and multiple mobile terminals, the method comprising:

[0006] Obtain the terminal location of all mobile terminals at the current time step, and obtain the transmission parameter optimization model, wherein the time step includes multiple time slots;

[0007] All terminal locations are input into the transmission parameter optimization model for data matching to obtain the output optimized subcarrier allocation, optimized phase shift sequence corresponding to multiple time slots, and optimized subcarrier power;

[0008] According to the optimized subcarrier allocation, corresponding subcarriers are allocated to multiple mobile terminals; according to the optimized phase shift sequence, the reflection unit of the intelligent reflector of each time slot is adjusted in the current time step; and according to the optimized subcarrier power, the power allocation of each subcarrier is adjusted in the current time step.

[0009] In some embodiments, obtaining the transmission parameter optimization model includes:

[0010] Based on the phase shift parameter sequence of the intelligent reflector and the position parameters of the multiple mobile terminals, a frequency response model corresponding to each mobile terminal in each subcarrier is generated;

[0011] Based on the subcarrier power parameters, noise power, channel bandwidth, frequency response model, and subcarrier allocation parameters of each subcarrier, the average bit error rate model is obtained.

[0012] The throughput model of each mobile terminal is obtained by subtracting the average bit error rate model from the first value, and then multiplying it by the subcarrier allocation parameter and the transmission symbol rate.

[0013] Based on the throughput minimization model that maximizes all mobile terminals, the transmission parameter optimization problem is obtained;

[0014] An initial transmission parameter optimization model is constructed based on the aforementioned transmission parameter optimization problem, and the initial transmission parameter optimization model is trained in multiple rounds to obtain the transmission parameter optimization model.

[0015] In some embodiments, generating a frequency response model for each mobile terminal in each subcarrier based on the phase shift parameter sequence of the smart reflector and the position parameters of the multiple mobile terminals includes:

[0016] Based on the location parameters of each mobile terminal, a direct link channel model between each mobile terminal and the base station is determined, and a first reflection channel model between each mobile terminal and the smart reflector is determined.

[0017] The frequency response model is obtained by multiplying the Fourier transform matrix elements, the first reflection channel model, the second reflection channel model, and the phase shift parameter sequence, and then multiplying the Fourier transform matrix elements with the direct link channel model. The second reflection channel model is the channel model between the smart reflector and the base station.

[0018] In some embodiments, obtaining the average bit error rate model based on the subcarrier power parameters, noise power, channel bandwidth, frequency response model, and subcarrier allocation parameters for each subcarrier includes:

[0019] The received signal-to-noise ratio model for each mobile terminal in each subcarrier is obtained by multiplying the subcarrier power parameter of each subcarrier by the square of the frequency response model, and then dividing by the noise power and the channel bandwidth.

[0020] The average bit error rate model is obtained by multiplying the subcarrier power parameter corresponding to each mobile terminal on each subcarrier with the received signal-to-noise ratio model, and then performing expectation processing.

[0021] In some embodiments, constructing an initial transmission parameter optimization model based on the transmission parameter optimization problem includes:

[0022] Based on the subcarrier power parameters, subcarrier allocation parameters, and phase shift parameter sequence in the transmission parameter optimization problem, the action space is obtained;

[0023] The state space is obtained based on the location parameters of each mobile terminal;

[0024] Using the average bit error rate model of the throughput model in the transmission parameter optimization problem, based on the time slot error symbols received by each mobile terminal at each subcarrier and the number of time slot symbols to be transmitted in each time slot, an estimated bit error rate reward model for each time step is obtained. The bit error rate reward model is used to reduce the received bit error rate of each mobile terminal.

[0025] Obtain the initial action parameters of the actor network and the initial comment parameters of the commenter network;

[0026] The initial transmission parameter optimization model is generated based on the action space, the state space, the bit error rate reward model, the initial action parameters, and the initial comment parameters.

[0027] In some embodiments, the step of obtaining the estimated bit error rate reward model for each time step based on the time slot error symbols received by each mobile terminal at each subcarrier and the number of time slot symbols to be transmitted in each time slot includes:

[0028] Based on multi-level modulation technology, the number of transmission time slots corresponding to the number of time slot symbols is determined;

[0029] The average symbol error rate for each time step is obtained by summing the number of time slot symbols corresponding to the total number of time slots in all the transmission time slots, and then dividing by the number of transmission time slots and the number of time slot symbols in turn.

[0030] The estimated bit error rate model is obtained by dividing the average symbol bit error rate by the logarithm of the modulation order.

[0031] Based on the logarithmic processing of the estimated bit error rate model, and then multiplied by a negative amplification factor, the terminal bit error rate reward model corresponding to each mobile terminal at each time step is obtained.

[0032] The estimated bit error rate reward model for each time step is obtained based on the minimum value among all the terminal bit error rate reward models.

[0033] In some embodiments, the step of training the initial transmission parameter optimization model multiple times to obtain the transmission parameter optimization model includes:

[0034] In each training round, the current training position parameters of each mobile terminal in the current training time step are input into the actor network to obtain a training action combination, and the current training action is sampled from the training action combination. The base station, the smart reflector and each mobile terminal are adjusted based on the current training action.

[0035] In the current training time step, the base station is controlled to send a training multi-carrier signal to each of the mobile terminals, and the current bit error rate reward corresponding to the training multi-carrier signal received by each mobile terminal is obtained, and the updated training position of each mobile terminal in the next training time step is obtained.

[0036] The training experience information of the current training position, the current training action, the current bit error rate reward, and the current training time step at the updated training position is stored in the experience replay pool.

[0037] The initial action parameters and the initial comment parameters are updated using the training experience information from the batch training time steps in the experience replay pool to obtain the updated actor network and commenter network, and the actor network after multiple updates is used as the model for optimizing the transmission parameters.

[0038] In some embodiments, updating the initial action parameters and the initial comment parameters using the training experience information from the batch training time steps in the experience replay pool to obtain updated actor networks and commenter networks includes:

[0039] Based on the state value model, the training state value corresponding to the training experience information is obtained, and the state value model is generated based on the estimated bit error rate reward model;

[0040] Based on the training experience information and the training state value, the training temporal difference residual and the training state value are obtained;

[0041] Obtain the importance weight of the update strategy, and truncate the importance weight of the update strategy based on the preset truncation parameter to obtain the truncated importance weight.

[0042] The initial action parameters are updated using gradient descent based on the truncation importance weights.

[0043] The initial comment parameters are updated using gradient descent based on the training state values.

[0044] To achieve the above objectives, a second aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the transmission parameter optimization method for a multi-carrier communication system as described in the first aspect.

[0045] To achieve the above objectives, a third aspect of this application provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the transmission parameter optimization method for the multi-carrier communication system described in the first aspect.

[0046] The method and related equipment for optimizing transmission parameters in a multi-carrier communication system proposed in this application include a base station, a smart reflector, and multiple mobile terminals. The method includes: first, obtaining the terminal locations of all mobile terminals at the current time step, and obtaining a transmission parameter optimization model, wherein the time step includes multiple time slots; then, inputting all terminal locations into the transmission parameter optimization model for data matching to obtain the output optimized subcarrier allocation, optimized phase shift sequences corresponding to multiple time slots, and optimized subcarrier power; finally, allocating corresponding subcarriers to multiple mobile terminals according to the optimized subcarrier allocation, adjusting the reflection units of the smart reflector in each time slot according to the optimized phase shift sequence in the current time step, and adjusting the power allocation of each subcarrier according to the optimized subcarrier power in the current time step. This application's embodiments no longer rely on the difficult-to-obtain and costly instantaneous channel state information, but instead directly utilize the more readily available real-time user location information as the basis for decision-making. This avoids the complex channel estimation process and the resulting high overhead and computational delay. By inputting the user's location into a pre-trained model, the optimized subcarrier allocation, intelligent reflector phase shift sequence, and power allocation scheme can be obtained instantaneously, and the system can be adjusted accordingly. This not only significantly reduces the complexity of resource optimization but also enables real-time and dynamic adjustment of transmission strategies based on user movement. This allows the system configuration to more effectively adapt to changes in the macroscopic channel environment caused by changes in user location. Ultimately, without the need to obtain instantaneous channel state information, the transmission performance of the entire multi-carrier communication system can still be reliably improved, especially ensuring communication fairness and service quality for mobile users.

[0047] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the structure of a multi-carrier communication system provided in an embodiment of this application.

[0049] Figure 2 This is a flowchart of a method for optimizing transmission parameters in a multi-carrier communication system, provided in another embodiment of this application.

[0050] Figure 3 This is a schematic diagram of the operating framework of Algorithm 1 provided in another embodiment of this application.

[0051] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0053] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit the scope of this application.

[0055] In recent years, in next-generation multi-carrier communication systems, intelligent reflector (RIS) technology has received widespread attention to meet the growing demand for high-speed and high-reliability communication and to meet the increasing number of users. RIS, through a large number of low-cost passive reflective elements, can intelligently regulate the propagation environment of wireless signals, thereby enhancing signal coverage and increasing system capacity. In such systems where RIS assists base stations in communicating with users, maximizing system communication performance requires joint optimization of system resources. These resources typically include: the phase shift configuration of the RIS, the communication power between the base station and the user, and so on.

[0056] In related technologies, in communication systems equipped with intelligent reflectors, system resources are typically optimized at the beginning of each time slot based on the precise Channel State Information (CSI) corresponding to the current time slot. However, since intelligent reflectors usually contain hundreds or thousands of reflective elements, accurately obtaining the channel state information for each link requires a large number of pilots and complex estimation algorithms. This consumes valuable time and frequency resources and introduces significant computational delays. Especially in scenarios where users move and channels change rapidly over time, frequent channel estimation can make the system overhead unbearable, resulting in low reliability of resource optimization decisions in actual communication systems equipped with intelligent reflectors.

[0057] To improve the reliability of resource parameter optimization between base stations and mobile terminals in multi-carrier communication systems equipped with intelligent reflectors, this application embodiment no longer relies on the difficult-to-obtain and costly instantaneous channel state information. Instead, it directly utilizes the more readily available real-time user location information as the decision-making basis, avoiding the complex channel estimation process and its associated high overhead and computational latency. By inputting the user's location into a pre-trained model, the optimized subcarrier allocation, intelligent reflector phase shift sequence, and power allocation scheme can be obtained instantaneously. Based on this, the system can be adjusted, significantly reducing the complexity of resource optimization. Furthermore, it can dynamically adjust the transmission strategy in real time according to the user's movement, enabling the system configuration to more effectively adapt to changes in the macroscopic channel environment caused by changes in user location. Ultimately, without the need to obtain instantaneous channel state information, it can still reliably improve the transmission performance of the entire multi-carrier communication system, especially ensuring the fairness of communication and quality of service for mobile users.

[0058] To better describe the transmission parameter optimization method for the multi-carrier communication system provided in this application, the following first describes a multi-carrier communication system to which the transmission parameter optimization method is applied. (Refer to...) Figure 1 This is a schematic diagram of a multi-carrier communication system provided in an embodiment of this application. Figure 1 As shown, the multi-carrier communication system includes a base station. A smart reflector (RIS, including M=M) x ×M y (one passive reflector unit) and multiple mobile terminals located within the service area Z. (i.e., MU1 to MU) k In this illustrative scenario, the base station With mobile terminal MU k The direct transmission link between them suffers from poor communication quality due to obstruction by buildings. Therefore, the base station... The signal can be sent to a smart reflector deployed on the exterior wall of another building, and then the smart reflector reflects the signal to each mobile terminal. The smart reflector is controlled by a connected RIS controller, which can adjust the phase shift of multiple reflective units on the smart reflector in real time according to the optimization strategy generated by the base station (as shown by the dotted line), thereby actively improving the quality of the reflection channel to ensure the communication performance of multiple mobile terminals.

[0059] Similar to traditional multi-carrier communication systems, the total bandwidth of the multi-carrier communication system provided in this application is divided equally into N orthogonal subcarriers, using a set To express. Furthermore, let Indicates base station The power allocated to the nth subcarrier. Assume the base station... The total available transmission power is The power allocation should satisfy make Indicates base station The zero-filled L-tap baseband equivalent multipath channel of the direct link to the k-th mobile terminal, where the superscript T is the transpose of the vector, and This represents the location information of the k-th mobile terminal in the i-th time slot; x k,i ,y k,i ,z k,i These represent the k-th mobile terminal MU respectively. k The coordinates on the X, Y, and Z axes. Furthermore, the change in the mobile terminal's position is typically slower than the channel change, meaning the mobile terminal's position can remain constant for multiple time slots. Therefore, in this application, the terminal position corresponding to other time slots after the first time slot in each time step is omitted.

[0060] Based on such Figure 1 The multi-carrier communication system shown below will be described in detail below, along with the method for optimizing the transmission parameters of the multi-carrier communication system in the embodiments of this application. (Refer to...) Figure 2 This is an optional flowchart of a method for optimizing transmission parameters in a multi-carrier communication system provided in this application embodiment. Figure 2 The method may include, but is not limited to, steps 100 to 300. It is also understood that this embodiment... Figure 2 The order of steps 100 to 300 is not specifically limited; the order of steps can be adjusted, or some steps can be added or removed, depending on actual needs. This method for optimizing the transmission parameters of a multi-carrier communication system can be applied to the information transmitting end, or to processors or controllers connected to the multi-carrier communication system, etc.

[0061] Step 100: Obtain the terminal location of all mobile terminals at the current time step, and obtain the transmission parameter optimization model.

[0062] Step 100 is described in detail below.

[0063] In some embodiments, due to the complexity of the channel between the base station and multiple mobile terminals, and because each reflective unit and the information transmitter, as well as each reflective unit and the information receiver, are considered to have a communication link during the reflection of data information by the intelligent reflective surface, it is extremely difficult to obtain the instantaneous channel state between the information receiver and the intelligent reflective surface in real time, which requires particularly high hardware specifications and is also difficult to achieve.

[0064] Therefore, in this embodiment, the terminal locations of all mobile terminals at the current time step are obtained through a positioning system (such as GPS or network-side positioning technology). These terminal locations are the key input replacing traditional Channel State Information (CSI), which macroscopically reflects the geometric relationship between the user, the base station, and the smart reflector. Then, the terminal locations of all mobile terminals at the current time step are used to replace the instantaneous channel state at the current moment as the basis for optimizing the transmission parameters of the multi-carrier communication system at the current moment. In this embodiment, the transmission parameters to be optimized in the multi-carrier communication system include the base station... Assigned to each mobile terminal MU k Subcarrier allocation w n,k Base stations For each subcarrier (e.g., the nth subcarrier), the subcarrier power p in each time slot n And the phase shift sequence of the intelligent reflector RIS in each time slot (i.e., the reflection phase shift of each reflector unit).

[0065] Simultaneously, the system loads or invokes a pre-built transmission parameter optimization model, which is typically a well-trained deep reinforcement learning model. It is worth noting that in this application, a time step consists of multiple time slots, meaning that a decision cycle (time step) is temporally divided into several shorter physical transmission units (time slots). This lays the foundation for subsequent finer dynamic adjustments under a single decision and for effective performance evaluation during the model training phase.

[0066] Furthermore, to obtain appropriate optimized transmission parameter decisions for the current time step, a transmission parameter optimization model is constructed in advance using a deep reinforcement learning framework, based on relevant data from the multi-carrier communication system. This model determines suitable transmission parameter optimization decisions based on the terminal locations of all mobile terminals, thereby reducing the information reception bit error rate at the information receiver. The construction of this transmission parameter optimization model will be further described below.

[0067] The process of obtaining the transmission parameter optimization model includes steps 110 to 150.

[0068] Step 110: Based on the phase shift parameter sequence of the smart reflector and the position parameters of multiple mobile terminals, generate the frequency response model of each mobile terminal in each subcarrier.

[0069] Step 110 will be described in detail below.

[0070] In some embodiments, the first step is to construct the physical layer mathematical representation of the communication link in a multi-carrier communication system. Let... Indicates base station With the m-th reflective element on the intelligent reflective surface RIS The L1 tap baseband equivalent channel between them is the second reflection channel model. Similarly, let... Represents the m-th reflecting element The L2 tap baseband equivalent channel between the RIS and the mobile terminal k is the first reflection channel model. In the smart reflector RIS, each element rescatters the received signal with an independent reflection coefficient. Specifically, let... This represents the sequence of phase shift parameters for the intelligent reflector RIS, where each... Including the amplitude coefficient β corresponding to each reflecting unit m ∈[0,1] and reflection phase shift θ m ∈[-π,π). For simplicity, the amplitude coefficient of all reflecting elements is fixed at β in this application. m =1, therefore, the phase shift parameter sequence of the smart reflector is Assuming the phase shift of the intelligent reflector RIS is quantized using D bits, then the set of discrete phase shift values ​​for each reflector element can be represented as Θ={0,Δθ,L,(2 D -1)Δθ}, where Δθ=π / 2 (D-1) .

[0071] In this process, the phase shift parameter sequence based on the smart reflector and location parameters of multiple mobile terminals Generate the frequency response model z for each mobile terminal (e.g., the kth one) in each subcarrier (e.g., the nth one). k,n This frequency response model, in a physical sense, represents the equivalent channel gain from the base station to a specific mobile terminal via the smart reflector, given a phase shift configuration of the smart reflector, as described below.

[0072] The process involves generating a frequency response model for each mobile terminal in each subcarrier based on the phase shift parameter sequence of the smart reflector and the position parameters of multiple mobile terminals, including steps 111 to 112.

[0073] Step 111: Based on the location parameters of each mobile terminal, determine the direct link channel model between each mobile terminal and the base station, and determine the first reflection channel model between each mobile terminal and the smart reflector.

[0074] Step 112: Based on the product of the Fourier transform matrix elements, the first reflection channel model, the second reflection channel model, and the phase shift parameter sequence, plus the product of the Fourier transform matrix elements and the direct link channel model, the frequency response model is obtained. The second reflection channel model is the channel model between the intelligent reflector and the base station.

[0075] Steps 111 to 112 are described in detail below.

[0076] In some embodiments, firstly, based on the position parameter q of each mobile terminal at a certain time step (e.g., the t-th time step), k,t Determine the direct link channel model between each mobile terminal and the base station. And determine each mobile terminal MU k First reflection channel model between each reflection unit of the smart reflective surface Specifically, the system utilizes known geometric information, namely the mobile terminal's position parameter q. k,t This process calculates the channel characteristics of each independent path constituting end-to-end communication. The direct link channel model describes the path of signal propagation directly from the base station to the mobile terminal; its model parameters (such as path loss and phase delay) are mainly determined by the distance and relative orientation between the two. Similarly, the first reflection channel model describes the characteristics of the path of signal propagation from the smart reflector to the mobile terminal; its model parameters are determined by the positional relationship between the smart reflector and the mobile terminal. In this way, this step decomposes the complex end-to-end channel into multiple simpler channel segments that can be mathematically represented based on location information.

[0077] Therefore, the base station via the m-th reflecting element The cascaded channel of the intelligent reflector RIS-mobile terminal k can be represented as: Where L3 = L1 + L2 - 1. Definition Zero-filled base station - Channel models for intelligent reflector RIS and intelligent reflector RIS-mobile terminal k, where each Therefore, the cascaded channel of all intelligent reflective surfaces and the k-th mobile terminal can be used... Represented as: h k =V k φ. Therefore, in conjunction with the base station - Direct channel and base station of mobile terminal k -RIS-Mobile terminal k-cascaded channel, base station The superimposed channel impulse response (CIR) between the mobile terminal and the terminal can be expressed as:

[0078] Next, based on a cyclic prefix (CP) length of N cp (N CP Impulse response under OFDM modulation of ≥max(L,L3)) frequency response It can be represented as: Can be written as Therefore, based on the elements of the Fourier transform matrix (i.e., DFT matrix F) N (each element in the first reflection channel model) Second Reflection Channel Model and phase shift parameter sequence The product, plus the elements of the Fourier transform matrix. With direct link channel model h d,k The product of these two factors yields the frequency response model as follows: The calculations here follow the principle of superposition of wireless signals, meaning the total signal received by the mobile terminal is the vector sum of the directly connected signal and the reflected signal. Specifically, the total response of the reflection path is the concatenated product of the channel from the base station to the smart reflector (i.e., the second reflection channel model), the phase modulation applied by the smart reflector itself (i.e., the phase shift parameter sequence), and the channel from the smart reflector to the mobile terminal (i.e., the first reflection channel model). This total reflection path response is then added to the directly connected path response. Finally, by multiplying by the elements of the Fourier transform matrix, this wideband channel model is projected into the frequency domain, thus obtaining the accurate equivalent channel gain on each independent subcarrier, i.e., the frequency response model.

[0079] Through the coordinated implementation of steps 111 and 112 above, based on easily obtainable terminal location parameters, and by combining geometric modeling and signal propagation principles, the direct measurement of instantaneous channel state information (CSI) in traditional methods is replaced. This completely avoids the huge system overhead and delay caused by channel estimation. It not only scientifically superimposes the direct path with the reflection path controlled by the intelligent reflector, but also extends the model to multi-carrier scenarios through Fourier transform. This provides a solid, reliable, and computationally inexpensive mathematical foundation for subsequent system performance analysis and resource optimization under CSI-free conditions.

[0080] Step 120: Based on the subcarrier power parameters, noise power, channel bandwidth, frequency response model, and subcarrier allocation parameters for each subcarrier, obtain the average bit error rate model.

[0081] Step 120 will be described in detail below.

[0082] In some embodiments, the frequency response model z of the physical layer is further... k,n This is transformed into a key communication quality metric. This step is based on the subcarrier power parameter p for each subcarrier (e.g., the nth subcarrier). n Noise power Channel bandwidth B, frequency response model z k,n and subcarrier allocation parameter w n,k The average bit error rate model is obtained. Because wireless channels experience rapid fading, the average bit error rate model here is a performance metric obtained by statistically averaging the channel fading. It reflects the average communication reliability over a period of time, thus avoiding dependence on instantaneous channel conditions, as described below.

[0083] In the multi-carrier communication system of this application, T is defined. S Let T be the duration of a transmission symbol, and T be the duration of each time slot. Then, in each time slot, the base station... The number of symbols transmitted is J = T / T s .

[0084] Let S n If the base station transmits symbols on the nth subcarrier in each time slot, then the signal received at the mobile terminal k is as shown in the following formula (1).

[0085]

[0086] in Let N0 represent the additive white Gaussian noise independently distributed at user k, i.e., the noise power. Let N0 represent the power spectral density of the AWGN, B represent the channel bandwidth of the system, and p0 represent the noise power. n For base stations The power allocated to each subcarrier, i.e., the subcarrier power parameter.

[0087] Additionally, the subcarrier allocation parameter w n,k The allocation of subcarriers to each mobile terminal is shown in the following formula (2).

[0088]

[0089] Furthermore, based on the subcarrier power parameters, noise power, channel bandwidth, frequency response model, and subcarrier allocation parameters for each subcarrier, an average bit error rate model is obtained, including the following steps 121 to 122.

[0090] Step 121: Based on the subcarrier power parameter of each subcarrier multiplied by the square of the frequency response model, and then divided by the noise power and channel bandwidth, obtain the received signal-to-noise ratio model of each mobile terminal in each subcarrier.

[0091] Step 122: Based on the product of the subcarrier power parameter corresponding to each mobile terminal on each subcarrier and the received signal-to-noise ratio model, perform expectation processing to obtain the average bit error rate model.

[0092] Steps 121 to 122 are described in detail below.

[0093] In some embodiments, to quantify the quality of the signal received by each mobile terminal on a specific subcarrier in a multi-carrier communication system, it is necessary to construct a received signal-to-noise ratio (SNR) model. In this process, based on the subcarrier power parameter p for each subcarrier... n Multiply by the frequency response model z k,n Square the result, then divide by the noise power N0 and the channel bandwidth B to obtain the received signal-to-noise ratio (SNR) model γ for each mobile terminal in decoding each subcarrier. n,k As shown in the following formula (3).

[0094]

[0095] This model provides a direct and crucial quantitative basis for subsequent evaluation of communication link performance.

[0096] Next, based on the received signal-to-noise ratio model γ obtained in the previous step... n,k Furthermore, an average bit error rate model that can measure long-term communication reliability was derived. The core of this step lies in establishing a mathematical relationship between the quality of the received signal and the probability of error in the final data decoding. Specifically, this model... Based on each subcarrier of each mobile terminal, the corresponding subcarrier power parameter P n,k With the received signal-to-noise ratio model γ n,k (i) Perform a product operation, and use the result as a core factor P for evaluating the probability of bit errors. n,k (γ n,k (i)). However, due to the time-varying characteristics of the wireless channel, the instantaneous received signal-to-noise ratio model fluctuates randomly. Therefore, in order to obtain a stable and statistically significant performance index, this step introduces expectation-taking processing. Expectation-taking processing is a key mathematical operation that eliminates the influence of random fluctuations by statistically averaging the instantaneous bit error rate performance under all possible channel states. Ultimately, it generates an average bit error rate model that can stably characterize the long-term performance of the system under specific subcarrier allocation parameters and power configurations, as shown in the following formula (4).

[0097]

[0098] By executing steps 121 and 122 above, controllable transmission parameters (such as subcarrier power parameters) and uncontrollable environmental factors (represented by the frequency response model) are transformed into an intuitive signal quality metric (received signal-to-noise ratio model). This instantaneous quality metric is further sublimated into a statistically significant and predictable long-term performance indicator (average bit error rate model), laying a solid theoretical foundation for subsequent transmission parameter optimization. This allows the optimization algorithm to have a clear and calculable optimization objective, enabling it to predict system performance under different parameter combinations based on the model, and ultimately find the optimal transmission strategy that minimizes the average bit error rate, without having to perform physical transmission tests for every possibility. This greatly improves the efficiency and feasibility of system optimization.

[0099] Step 130: Subtract the average bit error rate model from the first value, and then multiply by the subcarrier allocation parameter and the transmission symbol rate to obtain the throughput model of each mobile terminal.

[0100] Step 140: Based on the minimum throughput model that maximizes all mobile terminals, the transmission parameter optimization problem is obtained.

[0101] Step 150: Construct an initial transmission parameter optimization model based on the transmission parameter optimization problem, and train the initial transmission parameter optimization model in multiple rounds to obtain the transmission parameter optimization model.

[0102] Steps 130 to 150 are described in detail below.

[0103] In some embodiments, the communication reliability metric (average bit error rate) is further transformed into a user experience metric (i.e., throughput) that better meets business needs. In this process, the average bit error rate model is subtracted from a first value (i.e., 1), and then multiplied by the subcarrier allocation parameter w. n,k Given the symbol transmission rate R of mobile terminal k, we can obtain the throughput model C for each mobile terminal. k As shown in the following formula (5).

[0104]

[0105] The first value here is usually a constant (e.g., 1), so that the result of "first value minus average bit error rate" is inversely proportional to the bit error rate. That is, the lower the bit error rate, the higher this value, representing a higher effective data transmission rate. Multiplying this value by the subcarrier allocation parameter (a binary variable indicating whether a subcarrier is allocated to a user) accurately constructs the throughput model that each mobile terminal can obtain on the subcarrier allocated to it.

[0106] Next, the optimization objective will be set to be for any given Make all MU kTo maximize the minimum throughput, the transmission parameter optimization problem is obtained based on the minimum throughput model that maximizes all mobile terminals, as shown in the following formula (6).

[0107]

[0108] This objective is often referred to as the "maximize minimum fairness" criterion. Its core idea is to focus on the worst-performing mobile terminal in the system and strive to improve its throughput, thereby ensuring the basic quality of service for all mobile terminals and preventing communication interruptions for some users due to uneven resource allocation. By mathematizing this criterion, a formalized and well-defined transmission parameter optimization problem is obtained.

[0109] Since the position of the mobile terminal k is time-varying, and the goal is to optimize the phase shift of each reflection unit and the power p of each subcarrier in the absence of CSI, the optimization is achieved by... n and subcarrier allocation w n,k To improve the bit error rate performance for all users and enable MU k Located in a given area

[0110] In general, the transmission parameter optimization problem (6) can be solved by the following two methods. The first method is to first derive the objective function C at each position Q. k The closed-form expression of (Q) is obtained, and then the problem is solved using traditional discrete optimization methods. However, this method requires the base station to know the CSI of the system, which contradicts the starting point of this invention. Furthermore, in some complex communication scenarios, even if the CSI information is known, the objective function C... k The closed-form expression for (Q) is also difficult to obtain. The second method is to obtain the objective function C based on Monte Carlo simulation. k The method estimates the value of (Q), and then exhaustively searches for any position Q to obtain the optimal phase shift of the RIS reflector element. The computational complexity of this method is O(Q). Where Z is the region The number of internal positions Q and the theoretically infinite number of possible values ​​for Z make the computational complexity of this method unacceptable.

[0111] Therefore, the transmission parameter optimization problem (6) is a very complex optimization problem. To address these challenges, this application proposes an action decomposition-based PPO algorithm to solve this problem. In this process, based on the transmission parameter optimization problem (6), reinforcement learning techniques are used to construct an initial transmission parameter optimization model, and the initial transmission parameter optimization model is trained in multiple rounds to obtain a transmission parameter optimization model. Specifically, "constructing the initial model" refers to designing the structure of a deep neural network (such as an Actor-Critic network), which takes the terminal position as input and transmission parameters (phase shift, power, subcarrier allocation) as output. Then, in a simulation environment, the model is trained in multiple rounds by continuously trying different parameter combinations and receiving rewards or penalties according to the objective of the optimization problem (i.e., maximizing the minimum throughput). Finally, through this iterative learning, a mature transmission parameter optimization model that can quickly output an approximate optimal solution based on the input position is obtained.

[0112] The following section first describes how to construct the initial transmission parameter optimization model.

[0113] The initial transmission parameter optimization model is constructed based on the transmission parameter optimization problem, including the following steps 151 to 155.

[0114] Step 151: Based on the subcarrier power parameters, subcarrier allocation parameters, and phase shift parameter sequences in the transmission parameter optimization problem, obtain the action space.

[0115] Step 152: Obtain the state space based on the location parameters of each mobile terminal.

[0116] Step 153: Using the average bit error rate model of the throughput model in the transmission parameter optimization problem, based on the number of time slot error symbols received by each mobile terminal at each subcarrier and the number of time slot symbols to be transmitted in each time slot, obtain the estimated bit error rate reward model for each time step. The bit error rate reward model is used to reduce the received bit error rate of each mobile terminal.

[0117] Steps 151 to 153 are described in detail below.

[0118] Since the objective function of the transmission parameter optimization problem (6) has no closed-form expression and is a discrete optimization problem, traditional methods are difficult to solve it. This application adopts a data-driven DRL method to address this challenge. Specifically, the transmission parameter optimization problem (6) is reformulated as an MDP problem. It is considered an intelligent agent, which makes corresponding transmission parameter decisions by observing the terminal locations of all mobile terminals in order to maximize rewards.

[0119] In this application, a model-free DRL method is used to solve this problem. The corresponding MDP problem consists of quadruples. Composition. Among them, It is the state space, which represents the set of all possible states in the environment; is the action space, which represents the set of all actions that the agent can perform; r is the reward, the value of which is determined by the state s and the action a; γ∈[0,1] is a discount factor that reflects the impact of the reward obtained by the agent in the future time slot on the reward of the current state.

[0120] In some embodiments, in order to transform the transmission parameter optimization problem (6) into a paradigm solvable through reinforcement learning, it is first necessary to define all possible operations that the agent (i.e., the optimization model) can perform. This includes all the variable parameters that need to be decided in the transmission parameter optimization problem (6), specifically the subcarrier power parameter p used to adjust the subcarrier power of each subcarrier in each time slot. t ={p 1,t ,p 2,t ,L,p N,t Subcarrier allocation parameter w, used to determine the correspondence between users and subcarriers. t And the phase shift parameter sequence θ used to control beamforming of the smart reflector. t ={θ 1,t ,θ 2,t ,L,θ M,t}, together, form the action space a at time step t. t ={θ t ,p t ,w t Action space The dimension is 2 DM ·K n In intellectual property terminology, the action space refers to the set of all possible actions a reinforcement learning agent can take at any given moment. It fully defines the scope and capability of the model in making decisions.

[0121] After defining the agent's capabilities, it is necessary to provide it with environmental information upon which its decisions are based. A state space is constructed by using the location parameters of each mobile terminal as the core representation of the environment. The definition of the state should directly present key information about the current environment. In the multi-carrier communication system considered in this application, although the CSI is unknown, the base station can be determined. The location of the intelligent reflector RIS and the mobile terminal MU. Due to the base station The positions of the intelligent reflector RIS and the mobile terminal MU are fixed, and the time-varying position information of the mobile terminal MU is equivalent to the distance information between the intelligent reflector RIS and the mobile terminal MU. However, since large-scale fading of the wireless channel is determined by the distance between the transmitter and receiver, the CSI of the link between the intelligent reflector RIS and the mobile terminal MU is determined by the position of the mobile terminal MU. That is, the time-varying position information of the base station reflects the channel quality of the wireless channel between the RIS and the base station. Therefore, in this application, the easily obtainable position parameters of the mobile terminals are used to replace the traditional channel state information, which is difficult to estimate. At the t-th time step, the state space is the position of all mobile terminals, denoted as s. t =Q t ={q 1,t ,q 2,t ,L,q K,t}, where q k,t This represents the position of the k-th user at the t-th time step in the MDP. According to the definition of position, the state dimension is 3K. The state space is a key technical term; it represents the set of all possible states of the environment and is the basis for the agent's observation and decision-making. Therefore, through this step, the agent can perceive the geographical distribution of various mobile terminals in the communication environment and use this as input to formulate its optimization strategy.

[0122] It is worth noting that the t-th time step in the MDP problem refers to the t interactions between the agent and the environment. Unlike most existing technologies where one time step corresponds to one time slot, in this application, during the training phase, one time step in the MDP problem needs to include as many time slots as possible to accurately count the throughput of each mobile user. Specific details will be described in detail in the definition of reward below. Furthermore, during the prediction phase, when the base station transmits information to each mobile terminal, the trained neural network can quickly output the optimal phase shift of the RIS reflector, the optimal subcarrier allocation, and the optimal subcarrier power allocation based on the location information of each mobile terminal. In this case, the duration of one time step can be set to a relatively short time.

[0123] To guide the agent to learn effective optimization strategies, a clear feedback mechanism, i.e., a reward function, must be designed. In this application, the average bit error rate (BER) model established in the transmission parameter optimization problem (6) and used as the basis for the throughput model is combined with the actual or simulated communication results to obtain the estimated BER reward model for each time step. Specifically, the reward is calculated based on the statistical analysis of the number of time slot BER symbols actually generated by each mobile terminal when receiving signals, relative to the number of time slot BER symbols to be transmitted in each time slot. Since throughput is related to the average BER of the mobile terminal, in state s t =Q tThe estimated bit error rate reward model for the t-th time step is defined as shown in formula (7).

[0124]

[0125] In the estimated bit error rate reward model (7), an appropriate positive constant factor η and the log(·) function are used to amplify the difference between the minimum mean bit error rates obtained from different actions, thereby improving the learning efficiency and stability of the algorithm. To obtain the estimated bit error rate reward model r... t All need to be calculated average bit error rate That is, given Q t The average bit error rate across all mobile terminals at time step t. However, due to the unknown CSI, it is difficult to derive this value. The closed-form expression. However, it can be noted that when Q... t Given, This can be obtained using the Monte Carlo method, and therefore, the following will further describe how to determine the estimated bit error rate reward model based on this.

[0126] The following section will further describe how to construct the reward model for obtaining this estimated bit error rate.

[0127] The estimated bit error rate reward model for each time step is obtained based on the number of time slot error symbols received by each mobile terminal at each subcarrier and the number of time slot symbols to be transmitted in each time slot, including the following steps 1531 to 1535.

[0128] Step 1531: Based on multi-level modulation technology, determine the number of transmission time slots corresponding to the number of time slot symbols.

[0129] Step 1532: Accumulate the number of time slot symbols with error corresponding to the number of time slot symbols in all transmission time slots, and then divide by the number of transmission time slots and the number of time slot symbols in turn to obtain the average symbol error rate corresponding to each time step.

[0130] Step 1533: Based on the average symbol error rate divided by the logarithm of the modulation order, the estimated bit error rate model is obtained.

[0131] Step 1534: Based on the logarithmic processing of the estimated bit error rate model, multiply by a negative amplification factor to obtain the terminal bit error rate reward model for each mobile terminal at each time step.

[0132] Step 1535: Based on the minimum value among all terminal bit error rate reward models, obtain the estimated bit error rate reward model for each time step.

[0133] Steps 1531 to 1535 are described in detail below.

[0134] In some embodiments, assuming a base station With the assistance of the intelligent reflector RIS, F information bits are sent to each mobile terminal, and the position of MU is determined by Q. t It is given at the t-th time step.

[0135] To accurately assess communication performance over a period of time, it is first necessary to determine the size of the observation window used to calculate the bit error rate. Based on the multi-level modulation technique employed, namely M-ary modulation, the total number of transmission time slots required to transmit a certain number of time slot symbols is determined. Under M-ary modulation, the base station... To transmit G = F / (log2M) a ) symbols, where M a It is the modulation order, and the base station Transmit J = T / (T) in each time slot s J symbols. Therefore, in order to transmit J symbols, In I=G / J=FT s / (Tlog2M a Transmission is maintained in 10 time slots, i.e., the total number of time slot symbols is I and the number of time slot symbols is J.

[0136] Furthermore, in the t-th time step, the mobile terminal MU k The j-th symbol on the n-th subcarrier in the i-th time slot received is denoted as Y. i,j,k,n,t The time slot error symbol is shown in the following formula (8).

[0137]

[0138] After determining the observation window, statistical processing of the raw bit error data within the window is required to obtain a preliminary performance index. At this point, the total number of bit error symbols Y corresponding to the number of transmitted time slot symbols is accumulated across all determined transmission time slots. i,j,k,n,t The sum of these values ​​is then divided sequentially by the number of transmission time slots and the number of symbols in each time slot to obtain the average symbol error rate for each time step. As shown in the following formula (9).

[0139]

[0140] The average symbol error rate directly reflects the average probability that the symbols transmitted by the system will be erroneous within the current time step, and is a direct measure of the original reliability of the communication link.

[0141] Next, assuming the modulation symbols use Gray code mapping, in order to obtain an index that better meets the final performance evaluation criteria of the communication system, it is necessary to convert the symbol-level error rate into the bit-level error rate. At this point, based on the average symbol bit error rate (9), by dividing it by the modulation order M... a The logarithm is used to obtain the estimated bit error rate model. As shown in the following formula (10).

[0142]

[0143] Modulation order M a The logarithm of the modulation order (e.g., for 16-QAM, the logarithm of a modulation order of 16 is 4 to the base 2) represents the number of bits carried by each modulation symbol. This conversion is based on the assumption of coding methods such as Gray code, that a symbol error typically leads to only one or a few bit errors. Therefore, the resulting estimated bit error rate model (usually referring to Bit Error Rate, BER) more accurately reflects the data transmission quality ultimately perceived by the user.

[0144] Understandably, when the number of bits sent... Can be considered a true value Therefore, as long as F is large enough, this application can reliably estimate the bit error rate of all mobile terminals based on the bit error rate estimation model (10). This allows the agent to receive a reward r. t .

[0145] In obtaining bit error rate Next, it needs to be transformed into a form suitable as a reinforcement learning reward function so that the algorithm can learn effectively. As shown in formula (7) above, logarithmic processing is performed based on the estimated bit error rate model, and then the result is multiplied by a negative amplification factor to obtain the terminal bit error rate reward model r for each mobile terminal at each time step. k,t The purpose of performing logarithmic processing is to amplify the differences between different bit error rates, making the model more sensitive to small improvements in performance; while multiplying by a negative amplification factor ensures that a lower bit error rate corresponds to a higher reward value, which is consistent with the optimization objective of maximizing reward. At the same time, this factor can adjust the amplitude of the reward signal to optimize the stability and convergence speed of the learning process.

[0146] As shown in formula (7) above, since the system needs to serve multiple mobile terminals simultaneously, the optimization objective usually focuses on the fairness and overall performance of the system, especially the experience of the worst-performing user. Therefore, by comparing the terminal error rate reward model r of all mobile terminals... k,tThe minimum value among all rewards is selected to obtain the estimated bit error rate reward model for each time step. This operation embodies the optimization principle of "maximizing the minimum" or the "weakest link" principle, where the overall performance of the system is determined by the worst-performing user. By using the minimum value among all rewards as the final global reward signal, the optimization model can be driven to prioritize improving the communication quality of the worst-performing user, thereby enhancing the fairness and robustness of the entire multi-user system.

[0147] By performing steps 1531 to 1535 above, preliminary performance indicators (i.e., average symbol error rate) are extracted from the underlying physical phenomena through precise statistics. Then, these indicators are transformed into more practical bit-level performance indicators (i.e., estimated bit error rate model). Next, through mathematical transformation, they are shaped into a reward form suitable for machine learning (i.e., terminal bit error rate reward model). Finally, through the principle of fairness, the rewards of multiple users are integrated into a global optimization signal (i.e., estimated bit error rate reward model). This ensures that the reward signal not only accurately reflects the communication quality but also effectively guides the optimization model to learn towards the goal of improving the performance of the worst user and ensuring the overall fairness of the system, thus solving the core problem of reward design in multi-user scenarios.

[0148] Step 154: Obtain the initial action parameters of the actor network and the initial comment parameters of the commenter network.

[0149] Step 155: Generate an initial transmission parameter optimization model based on the action space, state space, bit error rate reward model, initial action parameters, and initial comment parameters.

[0150] Steps 154 to 155 are described in detail below.

[0151] As before, the MDP action space dimension corresponding to problem (P1) is 2. DM ·K n ·n, which increases exponentially with the number of RIS reflection units. This is because traditional deep reinforcement learning algorithms (such as DQN (Deep Q Network)) cannot handle MDP problems with large action space dimensions, while the recently proposed deep reinforcement learning algorithm—PPO algorithm—can effectively address such problems.

[0152] The following is an introduction to the PPO algorithm.

[0153] Unlike the value function-based DQN method, PPO is a policy-based method. It consists of an actor network for outputting the policy and a critic network for estimating the value function (where the network parameters of the actor network, i.e., the action parameters, are...). The network parameters of the commenter network (i.e., the comment parameters ω) are updated using a gradient-based method. Then update the value function. Suppose the state value model of the agent is as shown in the following formula (11).

[0154]

[0155] Its meaning is from the current state s t Begin by following the strategy The expected value of the reward obtained. Then the optimization problem of PPO can be shown in the following formula (12).

[0156]

[0157] in, Indicating a new strategy and old strategies The probability ratio is also called the importance weight. Represents the parameters before each round of training. The advantage function is defined in this application as the time difference residual, as shown in the following formula (13).

[0158]

[0159] And the advantage function A π The sign of A determines the direction of optimization: if A π A value greater than 0 indicates that the current action is better than the average action, and the objective function will encourage increasing the probability of that action; if A... π If the value is less than 0, it means that the current action is worse than the average action, and the objective function will suppress the probability of that action.

[0160] Furthermore, to improve training stability, the PPO method in this application employs PPO truncation to limit the magnitude of policy updates and avoid performance fluctuations caused by excessive differences between the old and new policies. The optimization objective is constrained to a safe region [1-ε, 1+ε]. Therefore, the optimization objective in (12) becomes as shown in formula (14).

[0161]

[0162] Among them, the function The definition is shown in formula (15).

[0163]

[0164] Where ε is a hyperparameter controlling the PPO cutoff range, it changes with gradient updates regardless of the advantage function. Positive and negative values ​​can ensure that the target strategy does not deviate too far from the behavioral strategy. It is the policy entropy with coefficient η, used to ensure that the PPO method fully explores.

[0165] In the PPO algorithm, the parameter updates of the actor network and the commentator network employ a mini-batch stochastic gradient descent method, where each update randomly draws C from the experience replay pool. B records {s t ,a t ,s t+1 ,r t |t=1,K,C B The parameters of the actor network and the commentator network are updated. For the actor network, the parameters are updated by maximizing the parameters in equation (14). To update parameters As shown in formula (16) below.

[0166]

[0167] in Let represent the learning rate of the actor network. For the commenter network, the parameter ω is updated by minimizing the mean squared error as shown in Equation (17).

[0168]

[0169] Where, α ω V represents the learning rate of the commenter network. tar (s t ) represents the target state value function.

[0170] Furthermore, to address the issue of excessively high dimensionality in the established MDP action space, this application employs an action combination method in the PPO algorithm. This involves combining smaller, independent actions from the original MDP problem into a single action, thereby reducing the dimensionality of the action space and consequently decreasing the number of neurons in the actor network's output layer. Specifically, in the MDP problem of this application, let... That is, each reflective unit of the intelligent reflective surface RIS is phase-shifted. Subcarrier power allocation and subcarrier allocation Consider it as an action, and Ψ t This is the combination of corresponding actions. Then, the actor network does not need to output π(a) t |s t Only output is required. As shown in formula (18) below.

[0171]

[0172] According to the combination action Ψ tThe definition is based on the phase shift of the m-th reflecting unit. There are 2 possible values. D There are M such small actions, therefore the actor network requires 2^M output layer neurons. D M. Therefore, by using the action combination method, the number of neurons in the output layer of the actor network can be reduced from 2. DM Reduced to 2 D M, thus greatly reducing the network size. Similarly, the total action space is reduced from 2... DM ·K n n reduced to 2 D M+nK+n. Therefore, the PPO algorithm proposed in this application can be used to solve the MDP problem in a large-scale discrete action space.

[0173] Therefore, in order to achieve complex decision-making and evaluation functions, it is first necessary to obtain the initial action parameters of the actor network. The initial action parameters and initial comment parameters ω of the commentator network. The actor network, a key technical component, outputs an action policy based on the current state; while the commentator network is responsible for evaluating the merits of the actor's chosen action. The initial action parameters and initial comment parameters typically refer to the random weights and biases of these neural networks before training begins, providing an initial starting point for subsequent learning and iterative optimization.

[0174] Then, all the core elements defined in the preceding steps are integrated. Specifically, this step, based on the established action space (defining what the model can do), state space (defining what the model can see), and bit error rate reward model (defining the learning objective and direction), combines these abstract definitions with concrete neural network entities (characterized by initial action parameters and initial comment parameters) to ultimately generate an initial transfer parameter optimization model. This generated initial transfer parameter optimization model is a fully structured, ready but untrained deep reinforcement learning agent that already possesses all the necessary components to autonomously learn optimization strategies through interaction with the environment.

[0175] By performing steps 151 to 155 above, the decision boundary (action space), perception ability (state space), and learning motivation (bit error rate reward model) of the agent are systematically defined, and it is equipped with a "brain" for decision-making and evaluation (an actor-commenter network initialized by initial action parameters and initial comment parameters). This makes the final initial transmission parameter optimization model a foundation for subsequent data-driven training to discover the complex nonlinear mapping relationship between the mobile terminal location and the optimal transmission parameters, making it possible to solve this type of optimization problem without a closed solution.

[0176] The following section will further describe how to train the optimization model for this initial transmission parameter.

[0177] The process of training the initial transmission parameter optimization model multiple times to obtain the transmission parameter optimization model includes steps 156 to 159.

[0178] Step 156: In each training round, input the current training position parameters of each mobile terminal in the current training time step into the actor network to obtain the training action combination, and sample the current training action from the training action combination. Adjust the base station, smart reflector and each mobile terminal based on the current training action.

[0179] Step 157: In the current training time step, the control base station sends a training multi-carrier signal to each mobile terminal, based on the current bit error rate reward corresponding to the training multi-carrier signal received by each mobile terminal, and obtains the updated training position of each mobile terminal in the next training time step.

[0180] Step 158: Store the current training position, current training action, current bit error rate reward, and training experience information of the current training time step at the updated training position in the experience replay pool.

[0181] Step 159: Update the initial action parameters and initial comment parameters using the training experience information from the batch training time steps in the experience replay pool to obtain the updated actor network and commenter network, and use the actor network after multiple updates as the transfer parameters to optimize the model.

[0182] Steps 156 to 159 are described in detail below.

[0183] In some embodiments, the process is based on the PPO algorithm mentioned above. The training process begins with the interaction between the agent and the environment. In the current training time step t of each round of training, the system first uses the current training location parameters of each mobile terminal as input. Feeding is provided to the actor network. The actor network then feeds its current state parameters s. t Output a training action combination that represents the probability distribution of all possible actions. Subsequently, in order to strike a balance between exploration and exploitation, the system will extract from this training action combination Random sampling is performed to obtain a specific current training action a. t Finally, the system will strictly follow this current training action to make actual adjustments to the controllable units in the multi-carrier communication system. These adjustments include configuring the base station's transmission parameters, setting the phase shift sequence of the smart reflector, and allocating corresponding subcarriers to each mobile terminal.

[0184] After the agent performs an action and adjusts the system parameters, it is necessary to observe the consequences of that action in the environment. In the current training time step, the system controls the base station to send training multi-carrier signals (i.e., OFDM signals) to each mobile terminal. After receiving the signal, each mobile terminal calculates the corresponding performance index based on its decoding result. The system then uses this index to generate the current bit error rate reward r using formula (7). t This reward provides direct, quantitative feedback on the quality of the current training action. Simultaneously, to simulate a dynamic environment and drive the training process, the system obtains the updated training position of each mobile terminal in the next training time step. This represents the state transition in the environment, i.e., the new state s. t+1 .

[0185] To enable the agent to learn from past experiences, the data generated during the interaction process needs to be recorded and stored. This step constitutes all the key information of a complete interaction event, namely the current training position s. t Current training action a t Current error rate reward r t And update training positions s t+1 Together, they are encapsulated into training experience information for the current training time step {s t ,a t ,s t+1 ,r t The system will then store this training experience information, which contains the complete causal chain, in a pool called the experience replay pool. In the data structure, the experience replay pool is a key technical component that can store a large amount of historical experience, breaking the temporal correlation between data, thereby improving the stability and efficiency of training.

[0186] When the experience replay pool After accumulating sufficient data, the actual optimization and updates of the neural network begin. This step involves retrieving data from the experience replay pool. A batch of training experience information (i.e., batch training time steps) is randomly selected and used to calculate the loss function and gradient, thereby synchronously updating the initial action parameters and initial comment parameters to obtain the updated actor network and commenter network. This update process is repeated many times, with each update slightly improving the network's decision-making ability and evaluation accuracy. Finally, after numerous rounds of interaction, storage, and update cycles, the repeatedly updated actor network is used as the final output, i.e., the transfer parameter optimized model.

[0187] The process of updating the initial action parameters and initial comment parameters using the training experience information from the batch training time steps in the experience replay pool to obtain the updated actor network and commenter network includes steps 1591 to 1595.

[0188] Step 1591: Based on the state value model, obtain the training state value corresponding to the training experience information.

[0189] Step 1592: Obtain the training time-series difference residual and training state value based on training experience information and training state value.

[0190] Step 1593: Obtain the importance weight of the update strategy, and truncate the importance weight of the update strategy based on the preset truncation parameter to obtain the truncated importance weight.

[0191] Step 1594: Update the initial action parameters using gradient descent based on the truncation importance weights.

[0192] Step 1595: Update the initial comment parameters using gradient descent based on the training state values.

[0193] Steps 1591 to 1595 are described in detail below.

[0194] In some embodiments, in order to evaluate the experience replay pool The long-term value of each state needs to be obtained by using the state value model (11) mentioned above to obtain the training state value corresponding to each state stored in the training experience information. State-value models, technically implemented using a critic network, primarily predict the expected value of future cumulative rewards obtained by following the current policy from a given state. Importantly, the training objective of this state-value model, in this application, is driven and calibrated by the actual reward signal generated by the bit error rate reward model, thereby ensuring that its evaluation results remain consistent with the system's ultimate optimization objective (i.e., reducing the bit error rate).

[0195] After obtaining the evaluation of the state value, it is necessary to calculate key metrics to guide network updates. This step is based on training experience information {s} retrieved from the experience replay pool. t ,a t ,s t+1 ,r t} and the training state value obtained in the previous step The training time-series difference residual (TD-Error) is calculated using the above formula (13). Using the right half of the equation in formula (17) above, we obtain the training state value V. tar (st (Target Value). Training time-series residual is a core reinforcement learning concept. It represents the difference between the single-step prediction of the state value (i.e., the immediate reward plus the value estimate of the next state) and the current state value estimate. This residual directly reflects the relative quality of the current action choice. Simultaneously, the calculated training state value serves as a supervisory signal or "target label" when the commentator network updates.

[0196] To ensure the stability and security of policy updates, the update step size needs to be limited. This step first obtains a parameter called the update policy importance weight, which represents the probability ratio of the new policy to the old policy (i.e., the policy that generated this batch of training experience information). Subsequently, based on a preset truncation parameter ε, the importance weights of the update strategy are truncated using formula (14) to obtain the truncated importance weights. Truncation is a hallmark technique of the Proximal Policy Optimization (PPO) algorithm. Its purpose is to limit the magnitude of policy updates to a reliable small range, effectively preventing policy performance from collapsing due to excessively large single update steps, and significantly improving the robustness of training.

[0197] After calculating all the necessary intermediate values, the actor network can be optimized. This step is based on the truncated importance weights obtained in the previous step. and training time-series difference residuals We jointly construct a specific objective function (i.e., the Clipped Objective of PPO), and then apply gradient descent to the initial action parameters. To update, the initial action parameters are adjusted using the formula (16) above. Update the policy. The truncated importance weights are used here to adjust the direction and magnitude of the gradient update, ensuring that significant updates are only performed when the policy improves and changes are minor, thus achieving smooth and effective policy improvement.

[0198] While updating the actor network, the commentator network also needs to be optimized simultaneously to enable it to more accurately evaluate state values. This step is based on the training state value V. tar (s t The initial comment parameters ω are then compared with the current value estimate output by the commenter network, as shown in Equation (17) above, and the mean squared error loss between the two is calculated. Subsequently, the initial comment parameters ω are updated using gradient descent. This process aims to continuously bring the value assessment capability of the commenter network closer to the real training state values ​​generated by actual interactions, thereby providing more accurate guidance for the updating of the actor network.

[0199] By executing steps 1591 to 1595, the critic network evaluates the merits of states and actions and calculates the critical training temporal difference residuals. Then, the PPO algorithm's unique truncation mechanism ensures the safety of policy updates. Finally, precise gradient updates are performed on both the decision-making actor network and the evaluation critic network. This collaborative update method under the Actor-Critic architecture not only ensures that the policy steadily progresses towards maximizing long-term rewards (i.e., minimizing the bit error rate), but also effectively avoids performance oscillations during training through PPO's conservative update strategy. This allows for efficient and reliable training of an initial transmission parameter optimization model into a high-performance final model.

[0200] Based on the above description, the following is a schematic diagram of the training algorithm for the initial transmission parameter optimization model provided in this application.

[0201] Algorithm 1: Location-Information-Based CSIOFDM RIS-Assisted Communication System Optimization Algorithm

[0202] Input the reward discount factor γ, and the experience return will be C. B and learning rate and α ω .

[0203] Initialize Actor network parameters The Critic network has parameters ω and an empirical replay pool of size C.

[0204] Output Actor network parameters

[0205] 1) for episode do 2) Initialize the environment and obtain the initial state. ; 3) for t do 4) cnt += 1; 5) Based on the status The Actor network outputs a strategy after combining actions. Then on Sampling to obtain action ; 6) Based on the action By adjusting phase shift optimization, power, and subcarrier allocation, the optimal allocation is obtained. 7) The intelligent agent transmits OFDM signals to the user. 8) The OFDM receiver reconstructs the signal, the mobile user calculates the average bit error rate, and the reward is calculated. Feedback is given to the agent and the agent moves to the next position; 9) The agent calculates the reward according to equation (7). And observe the next state. ; 10) Store experience information To the experience replay pool ; 11) if cnt % C == 0 then 12) Use The empirical information in the equations is calculated according to equations (11) and (14) respectively. and ; 13) for epoch do 14) Disorder Experience information; 15) for k do 16) From Extracting small batches of experience information The parameters are updated using equations (13) and (14) respectively. and ; 17) end for 18)end for 19) Clear experience data; 20)end for 21) end for Reference Figure 3 This is a schematic block diagram illustrating the training of an initial transmission parameter optimization model provided in an embodiment of this application. For example... Figure 3 As shown, in the first At each time step, the agent observes the location information of all mobile users as the state of the current time step. , will state Input an Actor network, and the Actor network outputs the current policy. The agent follows the probability distribution of the policy. Sampling yields specific actions and actions Send to the RIS controller, the RIS controller responds according to the action. The phase shift, subcarrier allocation, and subcarrier power allocation of each reflector in the RIS are adjusted. Then, with the assistance of the RIS, the agent transmits OFDM signals to each mobile user. Each mobile user calculates the average bit error rate and their respective reward. Feedback is given to the agent, which then moves to the next position. The agent calculates the reward according to equation (7). And observe the next location information of all mobile users to obtain the status. Once an interaction at a time step is completed, the agent saves the interaction data for that time step. and the probability of old actions And the state-value function is estimated through a Critic network. Calculate and save the target state value function at each time step. and dominance function When an intelligent agent interacts with its environment After a certain time step, use the saved data. Train a neural network with data and reuse that data. Round: Each time, a small batch of data is randomly sampled, and then the probability ratio is calculated. Cutting off alternative targets Error term of value function And entropy regularization term The total loss is obtained by weighted summation. Finally, the Adam optimizer is used to update the parameters of the neural network. and .when After the first round of training is completed, sampling continues using the updated network parameters. , Until the network converges.

[0208] By executing steps 156 to 159 above, an untrained initial transmission parameter optimization model is autonomously and iteratively optimized into a mature model with expert-level decision-making capabilities through extensive simulation interactions with the environment. This series of steps defines in detail the complete loop of data acquisition, experience storage, and model updates. Utilizing an experience replay pool mechanism, this process not only solves the problem of strong correlation in training data but also greatly improves data utilization. Ultimately, through multiple rounds of learning, the resulting transmission parameter optimization model deeply understands the complex mapping relationship between the mobile terminal's location and optimal communication parameters. Thus, in practical applications, only real-time location information needs to be input to quickly output an efficient system configuration scheme.

[0209] Steps 110 to 150 above constitute a complete and rigorous systematic method for generating intelligent decision-making models. This method starts with establishing a location-based physical channel model, progressively building a reliability model (average bit error rate) and a performance model (throughput), and defining a system-level optimization problem that balances fairness and efficiency. Finally, by employing training techniques such as deep reinforcement learning, the solution process for this complex, non-convex optimization problem is solidified into an efficient transmission parameter optimization model. The fundamental contribution of this series of steps is that it successfully transforms a traditional optimization problem dependent on CSI into an intelligent decision-making problem that can be learned and solved through data-driven methods, providing crucial technical support for achieving CSI-free adaptive communication system optimization.

[0210] Step 200: Input all terminal locations into the transmission parameter optimization model for data matching to obtain the output optimized subcarrier allocation, optimized phase shift sequence corresponding to multiple time slots, and optimized subcarrier power.

[0211] Step 200 is described in detail below.

[0212] In some embodiments, the system enters the decision generation phase. The core of this phase is to transform the acquired raw data into specific control commands. Specifically, the system uses the locations of all terminals as input data and feeds it into a loaded transmission parameter optimization model. The model performs a series of complex inference calculations (i.e., a data matching process) to find a set of system parameters that maximizes system performance (e.g., maximizes the minimum rate for all users) based on the current spatial distribution of users. After the calculations are complete, the model generates a complete, jointly optimized output, which includes three key parts: first, optimized subcarrier allocation, which clarifies which mobile terminal each subcarrier frequency resource should be allocated to in a multi-carrier system; second, optimized phase shift sequences corresponding to multiple time slots, which specify an optimal phase shift configuration for the intelligent reflector reflector unit for each time slot within the current time step, enabling fine-grained and dynamic adjustment of the reflected beam within a decision cycle; and third, optimized subcarrier power, which specifies the transmit power that the base station should use on each allocated subcarrier.

[0213] Step 300: Assign corresponding subcarriers to multiple mobile terminals according to the optimized subcarrier allocation, adjust the reflection unit of the intelligent reflector of each time slot in the current time step according to the optimized phase shift sequence, and adjust the power allocation of each subcarrier in the current time step according to the optimized subcarrier power.

[0214] Step 300 is described in detail below.

[0215] In some embodiments, the system enters the decision execution phase, applying the generated optimized parameters to the actual physical communication equipment. This process is parallel: First, the base station's resource scheduler allocates corresponding subcarriers to multiple mobile terminals based on the optimized subcarrier allocation, completing the partitioning of frequency domain resources. Second, the base station sends instructions to the controller of the intelligent reflector via the control link, adjusting the reflection units of the intelligent reflector in each time slot according to the optimized phase shift sequence in the current time step, so that the reflector can form the optimal reflection path in different time slots to serve the target user. Finally, the base station's power control unit adjusts the power allocation of each subcarrier in the current time step according to the optimized subcarrier power. Through these three synchronously executed adjustments, the transmission resources of the entire communication system are configured entirely according to the strategy output by the optimized model within the current time step, thereby achieving intelligent control of the wireless environment.

[0216] Through the implementation of steps 100 to 300 above, the method proposed in this application constructs an end-to-end optimized link from the terminal location to the system transmission parameters. Its core advantage lies in completely eliminating the reliance on traditional instantaneous channel state information (CSI). By directly utilizing more readily available user location information as decision input, this method fundamentally avoids the complex and high-overhead channel estimation process in traditional schemes, significantly reducing system overhead and decision latency. Simultaneously, the pre-trained transmission parameter optimization model can instantaneously output jointly optimized subcarrier allocation, dynamic phase shift sequences, and power allocation schemes based on real-time changes in user location. This achieves dynamic, efficient, and reliable adaptive optimization of the transmission parameters of the multi-carrier communication system, thereby effectively improving the overall system communication performance and resource utilization in complex mobile scenarios.

[0217] In addition, to verify the reliability of the implementation of this application, an algorithm complexity analysis was performed as described below.

[0218] The complexity of deep reinforcement learning algorithms mainly comes from training and prediction, which will be discussed separately below. Based on Algorithm 1, assume that the Actors and Critics in the PPO network have respectively and I ω The number of neurons in the i-th layer is denoted as follows: and

[0219] 1) Prediction complexity: The prediction stage uses a pre-trained neural network, which does not require gradient calculation. The computational complexity mainly comes from the forward propagation. For the PPO algorithm, the prediction stage only needs to use the action selection strategy output by the Actor network. Therefore, the computational complexity can be expressed as shown in the following formula (19).

[0220]

[0221] In fact, the complexity of forward propagation in an Actor network is very low and can be basically ignored.

[0222] 2) Training Complexity: During the training phase, it is first necessary to interact with the environment to obtain empirical data, and then use the empirical data to train the network. When training the network, both the Actor network and the Critic network need to perform forward propagation to calculate the loss function, and then backpropagate to calculate the gradient and update the network parameters. Therefore, the computational complexity can be approximately expressed as the following formula (20).

[0223]

[0224] in, It represents the computational complexity of the Actor network's forward propagation output of the action policy and the Critic network's forward propagation estimation of the state-value function when the agent interacts with the environment. It is the computational complexity of the base station sending bit information to the mobile user and estimating the bit error rate. It is the computational complexity of calculating the loss function during the forward propagation of the Actor and Critic networks during training. This represents the computational complexity of backpropagation to calculate gradients and update network parameters. They are represented by equations (21) to (22) below.

[0225]

[0226] In computer simulations, matrix multiplication is used to simulate the base station transmitting signals to mobile users. Approximately In real-world environments, It is approximately represented by the following formula (24).

[0227]

[0228] Where R is the transmission rate, c is the speed of light, d1 is the distance from the base station to the RIS, and d 2,k From RIS to MU k The distance.

[0229] This application also provides an electronic device, including:

[0230] At least one memory;

[0231] At least one processor;

[0232] At least one program;

[0233] The program is stored in memory, and the processor executes at least one program to implement the transmission parameter optimization method for the multi-carrier communication system described above in this application. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0234] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0235] The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0236] The memory 402 can be implemented in the form of ROM (Read-Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 402 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and is called and executed by the processor 401 to execute the transmission parameter optimization method of the multi-carrier communication system of the embodiments of this application.

[0237] Input / output interface 403 is used to implement information input and output;

[0238] The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0239] Bus 405 transmits information between various components of the device (e.g., processor 401, memory 402, input / output interface 403, and communication interface 404);

[0240] The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.

[0241] This application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described method for optimizing transmission parameters of a multi-carrier communication system.

[0242] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0243] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0244] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0245] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0246] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0247] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0248] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0249] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.

[0250] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0251] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0252] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0253] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method of optimizing transmission parameters of a multicarrier communication system, characterized in that, The multi-carrier communication system includes a base station, an intelligent reflecting surface and a plurality of mobile terminals, and the method comprises: obtaining terminal positions of all the mobile terminals at a current time step, and obtaining a transmission parameter optimization model, the time step comprising a plurality of time slots; inputting all the terminal positions into the transmission parameter optimization model for data matching to obtain an output optimization subcarrier allocation, an optimization phase shift sequence corresponding to a plurality of the time slots and an optimization subcarrier power; allocating corresponding subcarriers to a plurality of the mobile terminals according to the optimization subcarrier allocation, adjusting a reflecting unit of the intelligent reflecting surface of each of the time slots in the current time step according to the optimization phase shift sequence, and adjusting power allocation of each of the subcarriers in the current time step according to the optimization subcarrier power.

2. The method of Claim 1, wherein The transmission parameter optimization model comprises: generating a frequency response model corresponding to each of the mobile terminals in each subcarrier based on a phase shift parameter sequence of the intelligent reflecting surface and a position parameter of a plurality of the mobile terminals; obtaining an average bit error rate model based on a subcarrier power parameter, a noise power, a channel bandwidth, a frequency response model and a subcarrier allocation parameter of each of the subcarriers; obtaining a throughput model of each of the mobile terminals based on a first numerical value minus the average bit error rate model, multiplied by a subcarrier allocation parameter and a transmission symbol rate; obtaining a transmission parameter optimization problem based on maximizing a minimum of the throughput model of all the mobile terminals; constructing an initial transmission parameter optimization model based on the transmission parameter optimization problem, and performing multiple rounds of training on the initial transmission parameter optimization model to obtain the transmission parameter optimization model.

3. The method for optimizing transmission parameters of a multi-carrier communication system according to claim 2, characterized in that, The frequency response model corresponding to each of the mobile terminals in each subcarrier based on a phase shift parameter sequence of the intelligent reflecting surface and a position parameter of a plurality of the mobile terminals comprises: determining a direct link channel model between each of the mobile terminals and the base station based on the position parameter of each of the mobile terminals, and determining a first reflection channel model between each of the mobile terminals and the intelligent reflecting surface; obtaining the frequency response model based on a product of a Fourier transform matrix element, the first reflection channel model, a second reflection channel model and the phase shift parameter sequence, and adding a product of the Fourier transform matrix element and the direct link channel model, the second reflection channel model being a channel model between the intelligent reflecting surface and the base station.

4. The method for optimizing transmission parameters of a multi-carrier communication system according to claim 2, characterized in that, The average bit error rate model based on a subcarrier power parameter, a noise power, a channel bandwidth, a frequency response model and a subcarrier allocation parameter of each of the subcarriers comprises: obtaining a received signal-to-noise ratio model corresponding to each of the mobile terminals in each of the subcarriers based on a product of the subcarrier power parameter of each of the subcarriers and the frequency response model, and dividing by the noise power and the channel bandwidth; obtaining the average bit error rate model based on a product of the subcarrier power parameter corresponding to each of the mobile terminals on each of the subcarriers and the received signal-to-noise ratio model, and performing expectation processing.

5. The method for optimizing transmission parameters of a multi-carrier communication system according to claim 2, characterized in that, The initial transmission parameter optimization model is constructed based on the transmission parameter optimization problem, including: Based on the subcarrier power parameter, the subcarrier allocation parameter and the phase shift parameter sequence in the transmission parameter optimization problem, an action space is obtained; Based on the position parameter of each mobile terminal, a state space is obtained; Based on the average bit error rate model of the throughput model in the transmission parameter optimization problem, an estimated bit error rate reward model of each time step is obtained based on the time slot error code symbol received by each mobile terminal at each subcarrier, and the number of time slot symbols to be transmitted at each time slot, the bit error rate reward model being used to reduce the receiving bit error rate of each mobile terminal; An initial action parameter of an actor network and an initial comment parameter of a critic network are obtained; The initial transmission parameter optimization model is generated based on the action space, the state space, the bit error rate reward model, the initial action parameter and the initial comment parameter.

6. The method for optimizing transmission parameters of a multi-carrier communication system according to claim 5, characterized in that, The estimated bit error rate reward model of each time step is obtained based on the time slot error code symbol received by each mobile terminal at each subcarrier, and the number of time slot symbols to be transmitted at each time slot, including: Based on the multi-ary modulation technology, the number of transmission time slots corresponding to the number of time slot symbols is determined; All the time slot error code symbols corresponding to all the number of time slot symbols in all the number of transmission time slots are accumulated, and then divided by the number of transmission time slots and the number of time slot symbols in turn to obtain the average symbol error rate corresponding to each time step; Based on the average symbol error rate divided by the logarithm of the modulation order, an estimated bit error rate model is obtained; Based on the logarithmic processing of the estimated bit error rate model, and then multiplied by a negative amplification factor, a terminal bit error rate reward model corresponding to each mobile terminal at each time step is obtained; Based on the minimum value in all the terminal bit error rate reward models, the estimated bit error rate reward model of each time step is obtained.

7. The method for optimizing transmission parameters of a multi-carrier communication system according to claim 5, characterized in that, The transmission parameter optimization model is obtained by training the initial transmission parameter optimization model for multiple rounds, including: In each round of training, the current training position parameter of each mobile terminal in the current training time step is input into the actor network to obtain a training action combination, and a current training action is sampled from the training action combination, based on which the base station, the intelligent reflecting surface and each mobile terminal are adjusted; In the current training time step, the base station is controlled to send a training multicarrier signal to each mobile terminal, based on which a current bit error rate reward corresponding to the reception of the training multicarrier signal by each mobile terminal is obtained, and an updated training position of each mobile terminal in the next training time step is obtained; The current training position, the current training action, the current bit error rate reward and the updated training position are training experience information of the current training time step, and the training experience information is stored in an experience replay pool. The initial action parameter and the initial comment parameter are updated by using the training experience information of the batch training time steps in the experience replay pool to obtain an updated actor network and a critic network, and the updated actor network is used as the transmission parameter optimization model.

8. The method for optimizing transmission parameters of a multi-carrier communication system according to claim 7, characterized in that, The updating of the initial action parameter and the initial comment parameter by using the training experience information of the batch training time steps in the experience replay pool to obtain an updated actor network and a critic network comprises: Based on a state value model, a training state value corresponding to the training experience information is obtained, and the state value model is based on the estimated bit error rate reward model; Based on the training experience information and the training state value, a training time difference residual and a training state value are obtained; An update policy importance weight is obtained, and the update policy importance weight is truncated based on a preset truncation parameter to obtain a truncated importance weight; The initial action parameter is updated by gradient descent based on the truncated importance weight; The initial comment parameter is updated by gradient descent based on the training state value. 9.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to implement the transmission parameter optimization method of the multi-carrier communication system in any one of claims 1 to 8.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the transmission parameter optimization method of the multi-carrier communication system in any one of claims 1 to 8.