Retransmission system and method for link self-adaption in mobile environment

By using reinforcement learning neural networks in 5G networks for link adaptation, adaptively selecting MCS and adapting to new environments, the problem that existing algorithms cannot adapt to complex 5G network environments is solved and link throughput is improved.

CN120017230AActive Publication Date: 2025-05-16AIRUI COMM SYST (XIAMEN) CO LTD

Patent Information

Application Number
CN202510301364.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-16
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Link adaptive algorithms in existing 5G networks cannot effectively adapt to complex actual environments, especially in the case of retransmission mechanisms, feedback delays and channel switching, resulting in improper MCS decision-making and reduced throughput.

Method used

Reinforcement learning neural network is used for link adaptation. By constructing and initializing the Actor network and TargetActor network, combining deep reinforcement learning technology, adaptively selecting the appropriate MCS, and reconstructing the experience pool during channel switching, adjusting the exploration rate to adapt to the new environment.

Benefits of technology

It improves link throughput, can more effectively adapt to complex 5G network environments and channel changes, and solves the problem of performance degradation of traditional algorithms in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017230A_ABST
    Figure CN120017230A_ABST
Patent Text Reader

Abstract

The invention discloses a method for carrying out link self-adaption in a retransmission system and a mobile environment, which is characterized by comprising the following steps of: constructing and initializing a reinforcement learning neural network; acquiring information, wherein the information comprises channel condition information and HARQ process information fed back by the current user; monitoring whether a channel environment is switched; constructing an observation state vector according to the information; inputting the observation state vector into a reinforcement learning neural network, and outputting an MCS corresponding to the current HARQ process; transmitting a data packet containing the MCS to the user and receiving feedback information of the user; judging whether the current HARQ process is ended or not according to the feedback information, if so, calculating a retransmission conversion throughput award of the current HARQ process, performing state, action and award alignment integration, storing the state, action and award into an experience pool, and then entering a subsequent step; and training the reinforcement learning neural network based on a gradient descent method according to the plurality of empirical data. The problem that a traditional algorithm cannot adapt to an actual 5G network and a complex and changeable environment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless network, and in particular to a retransmission system and a method for performing link adaptation in a mobile environment. Background Art

[0002] The fifth generation mobile communication technology (5G) is a new generation of communication technology with high speed, low latency and large connection. One of the key technologies is link adaptation technology (LA). In the 5G network, the base station will convert the CQI (Channel Quality Indicator, CQI) fed back by the user into an appropriate MCS for data packet transmission to adapt to the changing channel environment. However, the 4-bit CQI only provides a coarse-grained estimate of the downlink channel quality, has a quantization error, and is limited by the periodically sent CQI signal. The base station's assessment of the channel will quickly become outdated, so they cannot provide a more timely channel quality estimate. At the same time, the complex retransmission mechanism and processing feedback delay will further affect the performance of each algorithm. In addition, when the channel changes dramatically, the above problems will suffer more serious effects, and incorrect MCS decisions will lead to a sharp drop in throughput.

[0003] In order to make MCS decisions, the existing methods are mainly traditional inner loop and outer loop link adaptation technologies and algorithms based on deep reinforcement learning. Traditional inner loop link adaptation directly maps CQI to an MCS according to a mapping table prepared in advance, but this method has quantization errors and generally selects a more conservative MCS, which cannot improve throughput. When the channel changes dramatically, the CQI expires rapidly and the inner loop performance degrades seriously. As a classic link adaptation algorithm, OLLA has the advantage of being able to continuously adjust the estimation deviation based on HARQ feedback information, so that the algorithm can adapt to dynamic channels, but the performance is greatly affected by the algorithm parameters TargetBLER and Step, and the optimal parameter configurations under different channels vary greatly, so it is not suitable for complex actual environments.

[0004] The link adaptation technology DRLLA based on deep reinforcement learning directly stores and trains when receiving HARQ feedback, but the actual 5G network has a complex retransmission mechanism and a certain processing feedback delay, which will seriously affect the performance of the algorithm. On the one hand, the retransmission mechanism may cause the same HARQ process to receive multiple NACKs. If the negative rewards are directly calculated without distinction and stored in the experience pool, the agent will receive a low reward, resulting in a more conservative selection of MCS. On the other hand, the delay in processing feedback is manifested in that the ACK / NACK feedback received in the next time slot is not the judgment result of the MCS of the current time slot, and the state, reward, and action experience constructed in the algorithm have no alignment operations, which will cause training confusion in the actual system and result in random selection of MCS. In addition, when the actual environment sends changes (such as LOS and NLOS switching due to movement and occlusion), the experience pool before the switch is not applicable to the new environment and if the exploration rate is not adjusted in time, the algorithm cannot fully adapt to the new environment and may converge to a suboptimal solution. In summary, this type of algorithm is no longer applicable in the face of complex actual environments. Summary of the invention

[0005] The present invention proposes a retransmission system and a method for link adaptation in a mobile environment to solve the problem that the existing algorithm is not applicable to the complex actual 5G network environment.

[0006] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0007] The present invention provides a method for performing link adaptation in a retransmission system and a mobile environment, comprising:

[0008] Build and initialize a reinforcement learning neural network;

[0009] Acquiring information, wherein the information includes channel status information and HARQ process information fed back by the current user;

[0010] Monitoring whether the channel environment is switched according to the information, if yes, reconstructing the experience pool of the reinforcement learning neural network and then proceeding to the subsequent steps, if no, directly proceeding to the subsequent steps;

[0011] constructing an observation state vector based on the information;

[0012] Determine whether the current HARQ process is in retransmission or a new process, if the current HARQ process is in retransmission, the MCS of the current HARQ process is consistent with the MCS used by the first HARQ process, if the current HARQ process is a new process, input the observed state vector into the reinforcement learning neural network, and output the MCS corresponding to the current HARQ process;

[0013] Transmitting a data packet containing the MCS to the user and receiving feedback information from the user;

[0014] Determine whether the current HARQ process is ended according to the feedback information. If so, calculate the retransmission discount throughput reward of the current HARQ process, and perform state, action, and reward alignment and integration according to the throughput reward, store them in the experience pool, and then proceed to the subsequent steps. If not, continue the current HARQ process and directly proceed to the subsequent steps;

[0015] Extract a number of experience data from the experience pool, train the reinforcement learning neural network based on the gradient descent method according to the number of experience data, obtain the trained reinforcement learning neural network, and use the trained reinforcement learning neural network to output the MCS corresponding to the next HARQ process.

[0016] Specifically, construct and initialize the neural network, including:

[0017] Constructing an Actor network and a TargetActor network, wherein the Actor network is used to output the MCS selected by the current HARQ process, and the TargetActor network is used to periodically replicate the weight of the Actor network, and both the Actor network and the TargetActor network include a gated recurrent unit and a neural network of several fully connected layers;

[0018] Initialize the Actor network and the TargetActor network, and assign default weights to the Actor network and the TargetActor network respectively.

[0019] Specifically, inputting the observed state vector into the reinforcement learning neural network includes:

[0020] After the observed state vector is input into the gated recurrent unit for processing, the value of each MCS is output by a plurality of the fully connected layers;

[0021] Combined according to the value The strategy makes an MCS decision for HARQ process i at time t The formula is:

[0022] in, the action selection strategy exploration rate for the reinforcement learning neural network, is the value of MCS, is the network default weight of the Actor network, is the current state vector, and a is the action.

[0023] Specifically, the retransmission discounted throughput reward of the HARQ process is calculated, and the state, action, and reward are aligned and integrated according to the throughput reward and then stored in the experience pool, including:

[0024] The throughput reward with retransmission discount is calculated for the HARQ process i, and the calculation formula is as follows:

[0025] in , Calculate the discount factor for the reward, represents the number of retransmissions of process i, and at time t, HARQ IDi is compared with the observed state vector when making the MCS decision. Save it in the dictionary as a key-value pair, the key is HARQIDi, and the value is the observed state vector at this time , expressed as:

[0026]

[0027] When the HARQ process ends after n time slots, the observed state vector at the time of decision is retrieved using the HARQ IDi dictionary. , and the reward calculated after n time slots and the MCS index of the first decision , current observation status Form a group of aligned experience points , and store it in the experience pool;

[0028] When the triggered HARQ process ends, the observed state vector is recorded. To update the key-value pair information corresponding to the HARQ process i.

[0029] Specifically, a number of experience data are extracted from the experience pool, and the reinforcement learning neural network is trained by the gradient descent method according to the number of experience data:

[0030] Randomly extract N pieces of experience data from the experience pool to form a data set M, and use the Target network to calculate each piece of experience data in the data set M. TDTarget value , the calculation formula is as follows:

[0031]

[0032] represents the immediate reward of the currently selected MCS and the discounted return of future rewards, Used to estimate the true value of a state-action pair, where ρ is the discount factor, Indicates the next state, Indicates the next action. Represents the network default weight of the Target Actor network;

[0033] Use the Actor network to calculate the action value of each piece of experience data :

[0034]

[0035] Use the gradient descent method to construct the loss function of the reinforcement learning neural network :

[0036] .

[0037] Specifically, monitoring whether the channel environment is switched according to the information includes:

[0038] Performing filtering on the channel state information to obtain filtered channel state information;

[0039] It is determined whether the channel environment should be switched based on the filtered channel status information and a preset condition.

[0040] Specifically, monitoring whether the channel environment is switched according to the information includes:

[0041] The channel status information is filtered, wherein the filtering process is formulated as follows:

[0042]

[0043]

[0044] Where CQI is the CQI value reported by the user, SRS is the uplink SRS signal receiving power, is the filtered value of CQI at time t, is the value of the uplink SRS signal received power after filtering at time t, , is the filter coefficient;

[0045] when or When , the NLOS scene to LOS scene switching indication is triggered;

[0046] when or When the LOS scene switches to the NLOS scene, the indication is triggered;

[0047] in, , , , is the switching control parameter.

[0048] Specifically, reconstructing the experience pool of the reinforcement learning neural network includes:

[0049] Clear the algorithm experience pool, increase the action selection strategy exploration rate according to the preset action selection strategy exploration rate change formula, or extract the experience in the recent time period to reconstruct the experience pool, and then increase the action selection strategy exploration rate according to the preset action selection strategy exploration rate change formula.

[0050] Specifically, after outputting the MCS corresponding to the current HARQ process and before transmitting a data packet containing the MCS to the user, the action selection strategy exploration rate is reduced according to a preset action selection strategy exploration rate change formula.

[0051] Specifically, the preset action selection strategy exploration rate change formula is as follows:

[0052]

[0053] in, represents the initial exploration rate, t is the current time, is the switching triggering time indicated by the LOS / NLOS switching, is the attenuation factor.

[0054] The beneficial effects of the present invention are:

[0055] The present invention proposes a retransmission system and a method for link adaptation in a mobile environment, which improves link throughput by adaptively selecting MCS. The present invention takes into account the complex situation of retransmission mechanism and processing delay existing in the actual 5G communication network and LOS / NLOS switching in the actual environment, combines deep reinforcement learning technology, and adaptively selects MCS, solving the problem that the traditional OLLA algorithm or DRLLA algorithm cannot adapt to the actual 5G network and complex and changing environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 A flowchart of a retransmission system and a method for performing link adaptation in a mobile environment of the present application;

[0057] Figure 2 A schematic diagram of the HARQ retransmission mechanism and delay feedback in the 5G network used in this application;

[0058] Figure 3 Schematic diagrams of several channel switching situations in the 5G mobile scenario of the present application ((a) represents a full LOS channel environment, (b) represents a full NLOS channel environment, (c) represents a LOS to NLOS channel environment, and (d) represents a NLOS to LOS channel environment). DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0062] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use, or are directions or positional relationships commonly understood by those skilled in the art. These directions or positional relationships are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.

[0063] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.

[0064] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0065] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0066] like Figure 1As shown, the present invention provides a retransmission system and a method for link adaptation in a mobile environment, including the following stages:

[0067] Initialization phase to build a neural network for selecting MCS and initialize the neural network structure;

[0068] In the information collection and processing phase, the base station collects the CQI and HARQ information fed back by the user, integrates them and constructs the observation state vector;

[0069] During the monitoring environment switching phase, the base station filters the periodic CQI fed back by the user, and when the filtered CQI meets the given conditions, the LOS / NLOS switching indication is triggered and related operations are performed;

[0070] In the MCS selection phase, the observed state vector is input into the neural network, and the neural network outputs the MCS prepared for the current HARQ process, then reduces the algorithm exploration rate and transmits data packets to the user;

[0071] In the user unpacking phase, the user decodes the received packet and uses CRC to check the data, generates HARQ feedback information including ACK / NACK, number of retransmissions, etc., and feeds it back to the base station;

[0072] During the HARQ process monitoring phase, monitor whether the HARQ process is completed. If not, continue the HARQ retransmission process. If completed, proceed to the next step.

[0073] In the reward calculation and state storage phase, the retransmission discount throughput reward of the HARQ process is calculated, and the state, action, and reward are aligned and integrated and then stored in the experience pool;

[0074] During the neural network training phase, batches of experience data are extracted from the experience pool, and the neural network is trained using the gradient descent method.

[0075] Specifically, the method includes the following steps:

[0076] Step 1.1: Build two neural networks with gated recurrent units (GRU) and several fully connected layers, as the Actor network and the TargetActor network. The Actor network is used to output the MCS selected by the current HARQ process of the base station, and the TargetActor network is used to periodically replicate the weights of the Actor network. In the initialization phase, the two networks are given default weights. as well as ;

[0077] Step 2.1 The base station collects the channel status information fed back by the user at time t , HAQR process i contains: MCS index used in the last transmission of this HARQ process 、Number of times the HARQ process data packet has been retransmitted And the latest data packet of the HARQ process instruct The observation information finally constitutes the observation state vector based on which the algorithm makes the MCS decision at time t. ;

[0078] Step 2.2 combines the most recent L observed state vectors and loops them into the GRU layer for temporal correlation learning. Specifically, ;

[0079] Step 3.1: Filter the periodic CQI or uplink SRS signal receiving power that changes dramatically to smooth its jitter to increase the accuracy of handover recognition. The filtering formula is as follows:

[0080]

[0081]

[0082] Where CQI is the CQI value reported by the user, SRS is the uplink SRS signal receiving power, is the filtered value of CQI at time t, is the value of the uplink SRS signal received power after filtering at time t, , is the filter coefficient;

[0083] Step 3.2: or When , the NLOS scene to LOS scene switching indication is triggered; when or When the LOS scene to NLOS scene switching indication is triggered, the time when the scene switching indication is triggered is recorded at the same time .in , , , is the switching control parameter. When the above switching indication is triggered, the experience pool is reconstructed to increase the exploration rate;

[0084] When the switching indication in step 3.2 is triggered, the reconstructed experience pool performs one of the following operations:

[0085] (1) Clear the algorithm experience pool and increase the exploration rate to fully adapt to the new environment and prevent convergence to a suboptimal solution. Otherwise, proceed directly to the subsequent steps to select MCS;

[0086] (2) Since the scene switching indication recorded in step 2.2 of claim 4 has a certain delay, that is, The switch actually occurs at the moment indicates the average delay for scene switching. The experience pool already has a small amount of experience in the new environment during the time period, so the recent The experience in the time period is used to reconstruct the experience pool. Then the exploration rate is increased to fully adapt to the new environment and prevent convergence to a suboptimal solution. Otherwise, proceed directly to the subsequent steps to select MCS;

[0087] When the switching indication in step 3.2 is triggered, the exploration rate is increased by performing the following operations:

[0088] Since the neural network no longer needs too much exploration of the environment after convergence, after continuous training The exploration rate of a strategy changes according to the following formula:

[0089]

[0090] in, represents the initial exploration rate, t is the current time, The switching triggering moment in the monitoring environment switching in claim 4, is the decay factor, where the initial exploration rate is generally 1 or 0.9.

[0091] Step 4.1: If the HARQ process is in retransmission, the MCS of the retransmitted data packet is kept consistent with the MCS used for the first time.

[0092] Step 4.2: If the HARQ process is new, the agent performs neural network reasoning to output MCS, specifically, After the input is processed by the GRU layer, several fully connected layers output the value of each MCS, and then combined The strategy makes an MCS decision for HARQ process i at time t :

[0093]

[0094] in This value allows the agent to choose actions in a random manner to explore the environment, so as to prevent the algorithm from falling into a suboptimal solution.

[0095] Step 5.1: The user uses the MCS indication in the DCI information to perform MCS demodulation. Then, the CRC cyclic redundancy check code will judge the demodulation result. If the CRC result is that the demodulation is successful, an ACK feedback message is generated. Otherwise, a NACK message is generated and the number of retransmissions is increased. Add 1. Finally, the above information is packaged into HARQ information and waits for the uplink control channel opportunity to be fed back to the base station.

[0096] Step 6.1: When the HARQ information received from the user contains ACK or the number of retransmissions is equal to 3 and is still NACK, the HARQ process is judged to be over, which also means that the HARQ process is about to transmit new data, that is, the MCS output by the agent is used to trigger the reward calculation and state storage phase. If the condition is not met, the HARQ process is judged to be not over, and the MCS used when the data packet was first transmitted is used to retransmit the data packet.

[0097] Step 7.1, calculate reward: Calculate the throughput reward with retransmission discount for a completed HARQ process i as follows:

[0098]

[0099] in is the action taken for HARQ process i at time t, Represents the effective throughput brought by this action decision.

[0100] Step 7.2, record status information: Due to the processing power of the user end, HARQ feedback generally has a delay. In order to avoid confusion between status, reward, and action, the HARQID in the HARQ information is used for discrimination. For example, at time t, HARQIDi is compared with the observed state vector when making the MCS decision. Save it in the dictionary as a key-value pair, the key is HARQIDi, and the value is the observed state vector at this time , expressed as:

[0101]

[0102] According to claim 7, when the HARQ process is identified to have ended after n time slots, the observed state vector at the time of the decision is retrieved by using the HARQ ID to look up the dictionary. , and the reward calculated in step (1) will be used and the MCS index of the first decision , current observation status Form a group of aligned experience points , and stored in the experience pool.

[0103] Step 7.3, update key-value pair: According to claim 7, when the triggered HARQ process ends, record the observed state vector at this time To update the key-value pair information corresponding to the HARQ process i.

[0104] Step 8.1: Randomly extract N pieces of experience data from the experience pool to form a data set M, and use the Target network to calculate each piece of experience data TDTarget value :

[0105]

[0106] Represents the discounted return of the immediate reward and future rewards of the currently selected MCS, which is used to estimate the true value of a state-action pair, where ρ is the discount factor, Indicates the next state, Indicates the next action. Represents the network default weight of the TargetActor network;

[0107] Step 8.2: Use the Actor network to calculate the action value of each piece of experience data :

[0108]

[0109] Step 8.3: Use gradient descent to minimize the loss function of the deep neural network :

[0110]

[0111] like Figure 2 As shown, Figure 1 A schematic diagram of the HARQ retransmission mechanism and delay feedback in the 5G network involved in the present invention is shown. The figure shows that the HAQR feedback has a delay of 4 time slots, and the retransmission mechanism is that the retransmitted data uses the same MCS as the initial transmission.

[0112] like Figure 3 As shown: Figure 3 Several channel switching situations in 5G mobile scenarios are demonstrated, namely: full LOS, full NLOS, LOS switching to NLOS, and NLOS switching to LOS. Figure 3 It represents the movement of users on this two-dimensional plane. The unit is meter. gNB represents 5G base station and UE represents user.

[0113] The reinforcement learning parameters of the reinforcement learning neural network in the embodiment of the present invention are shown in Table 1 below:

[0114] Table 1 Parameter Type Parameter Value Discount Factor 0.99 Learning Rate 0.005 Batch size 64 Hidden layer dimensions 128

[0115] The invention proposes a retransmission system and a method for link adaptation in a mobile environment, which improves link throughput by adaptively selecting MCS. The present invention takes into account the complex situation of the retransmission mechanism and processing delay existing in the actual 5G communication network and the LOS / NLOS switching in the actual environment, combines deep reinforcement learning technology, and adaptively selects MCS, solving the problem that the traditional OLLA algorithm or DRLLA algorithm cannot adapt to the actual 5G network and complex and changing environments. The present invention adaptively selects MCS through mechanisms such as retransmission reward conversion, experience alignment, and switching indication operation, successfully solving the problem that the traditional algorithm cannot adapt to the system caused by the retransmission mechanism, processing feedback delay, and mobility channel switching.

[0116] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for link adaptation in a retransmission system and a mobile environment, characterized in that: include: Build and initialize a reinforcement learning neural network; Acquiring information, wherein the information includes channel status information and HARQ process information fed back by the current user; Monitoring whether the channel environment is switched according to the information, if yes, reconstructing the experience pool of the reinforcement learning neural network and then proceeding to the subsequent steps, if no, directly proceeding to the subsequent steps; constructing an observation state vector based on the information; Determine whether the current HARQ process is in retransmission or a new process, if the current HARQ process is in retransmission, the MCS of the current HARQ process is consistent with the MCS used by the first HARQ process, if the current HARQ process is a new process, input the observed state vector into the reinforcement learning neural network, and output the MCS corresponding to the current HARQ process; Transmitting a data packet containing the MCS to the user and receiving feedback information from the user; Determine whether the current HARQ process is ended according to the feedback information. If so, calculate the retransmission discount throughput reward of the current HARQ process, and perform state, action, and reward alignment and integration according to the throughput reward, store them in the experience pool, and then proceed to the subsequent steps. If not, continue the current HARQ process and directly proceed to the subsequent steps; A plurality of pieces of experience data are extracted from the experience pool, and the reinforcement learning neural network is trained based on the gradient descent method according to the plurality of pieces of experience data.

2. A method for link adaptation in a retransmission system and a mobile environment according to claim 1, characterized in that: Build and initialize the neural network, including: Constructing an Actor network and a Target Actor network, wherein the Actor network is used to output the MCS selected by the current HARQ process, and the Target Actor network is used to periodically replicate the weight of the Actor network, and both the Actor network and the Target Actor network include a gated recurrent unit and a neural network of several fully connected layers; Initialize the Actor network and the Target Actor network, and assign default weights to the Actor network and the Target Actor network respectively.

3. A method for link adaptation in a retransmission system and a mobile environment according to claim 2, characterized in that: Inputting the observed state vector into the reinforcement learning neural network comprises: After the observed state vector is input into the gated recurrent unit for processing, the value of each MCS is output by a plurality of the fully connected layers; Combined according to the value The strategy makes an MCS decision for HARQ process i at time t The formula is: , in, the action selection strategy exploration rate for the reinforcement learning neural network, is the value of MCS, is the network default weight of the Actor network. is the current state vector, and a is the action.

4. A method for link adaptation in a retransmission system and a mobile environment according to claim 2, characterized in that: Calculating the retransmission discounted throughput reward of the HARQ process, and aligning and integrating the state, action, and reward according to the throughput reward and storing them in the experience pool, including: The throughput reward with retransmission discount is calculated for the HARQ process i, and the calculation formula is as follows: , in , Calculate the discount factor for the reward, represents the number of retransmissions of process i, and at time t, HARQ ID i is combined with the observed state vector when making the MCS decision Save it in the dictionary as a key-value pair, the key is HARQ ID i, and the value is the observed state vector at this time , expressed as: , When the HARQ process ends after n time slots, the observed state vector at the time of decision is retrieved by using the HARQ ID i to look up the dictionary , and the reward calculated after n time slots and the MCS index of the first decision , current observation status Form a group of aligned experience points , and store it in the experience pool; When the triggered HARQ process ends, the observed state vector is recorded. To update the key-value pair information corresponding to the HARQ process i.

5. A method for link adaptation in a retransmission system and a mobile environment according to claim 4, characterized in that: Extract a number of experience data from the experience pool, and train the reinforcement learning neural network by gradient descent method according to the number of experience data: Randomly extract N pieces of experience data from the experience pool to form a data set M, and use the Target network to calculate each piece of experience data in the data set M. TD Target value , the calculation formula is as follows: , represents the immediate reward of the currently selected MCS and the discounted return of future rewards, Used to estimate the true value of a state-action pair, where ρ is the discount factor, Indicates the next state, Indicates the next action. Represents the network default weight of the Target Actor network; Use the Actor network to calculate the action value of each piece of experience data : , Use the gradient descent method to construct the loss function of the reinforcement learning neural network : 。 6. A method for link adaptation in a retransmission system and a mobile environment according to claim 1, characterized in that: Monitoring whether the channel environment is switched according to the information includes: Performing filtering on the channel state information to obtain filtered channel state information; It is determined whether the channel environment should be switched based on the filtered channel status information and a preset condition.

7. A method for link adaptation in a retransmission system and a mobile environment according to claim 6, characterized in that: Monitoring whether the channel environment is switched according to the information includes: The channel state information is filtered, where the channel state information includes a CQI or an SRS, wherein a formula for the filtering process is as follows: , , Where CQI is the CQI value reported by the user, SRS is the uplink SRS signal receiving power, is the filtered value of CQI at time t, is the value of the uplink SRS signal received power after filtering at time t, , is the filter coefficient; when or When , the NLOS scene to LOS scene switching indication is triggered; when or When the LOS scene switches to the NLOS scene, the indication is triggered; in, , , , is the switching control parameter.

8. The method for performing link adaptation in a retransmission system and a mobile environment according to claim 1, characterized in that: Reconstructing the experience pool of the reinforcement learning neural network includes: Clear the algorithm experience pool, increase the action selection strategy exploration rate according to the preset action selection strategy exploration rate change formula, or extract the experience in the recent time period to reconstruct the experience pool, and then increase the action selection strategy exploration rate according to the preset action selection strategy exploration rate change formula.

9. A method for link adaptation in a retransmission system and a mobile environment according to claim 1, characterized in that: After outputting the MCS corresponding to the current HARQ process and before transmitting a data packet containing the MCS to the user, the action selection strategy exploration rate is reduced according to a preset action selection strategy exploration rate change formula.

10. A method for link adaptation in a retransmission system and a mobile environment according to claim 8 or 9, characterized in that: The preset action selection strategy exploration rate change formula is as follows: , in, represents the initial exploration rate, t is the current time, is the switching triggering time indicated by the LOS / NLOS switching, is the attenuation factor.

Citation Information

Patent Citations

  • Method for improving transmission performance of wireless communication downlink

    CN114362888A

  • Link adaptation

    CN117581493A

  • Adaptive modulation coding method based on LSTM-DQN

    CN119324762A

  • Downlink resource scheduling method and device for 5G differentiated scene, and storage medium

    CN119584307A

Cited By

  • Beam hopping resource allocation method and device based on deep reinforcement learning

    CN120956320A