Information processing device, information processing method, reception device, reception method, and program
Reinforcement learning is employed to create a learning model for channel estimation, addressing data preparation challenges and improving accuracy in wireless communication systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-17
- Publication Date
- 2026-03-26
AI Technical Summary
Existing channel estimation methods in wireless communication systems face challenges in preparing learning data for all frequency domains, making it difficult to perform accurate channel estimation using machine learning.
A technique utilizing reinforcement learning to generate a learning model for channel estimation by training on channel estimates of multiple wireless resource regions, using channel estimates as state values and performance values as rewards, enabling improved channel estimation through machine learning.
Facilitates easy and accurate channel estimation by leveraging reinforcement learning, enhancing the performance of wireless communication systems.
Smart Images

Figure JP2024033134_26032026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, Receiver, Reception Method, and Program
[0001] The present invention relates to an information processing apparatus, an information processing method, a receiver, a reception method, and a program.
[0002] In a wireless communication system such as LTE (Long Term Evolution) and NR (New Radio), in order for a receiver to correctly receive a wireless signal transmitted from a transmitter, the receiver performs channel estimation using a reference signal to receive the wireless signal. For example, Patent Document 1 discloses a technique for suppressing interference between reference signals and improving channel estimation accuracy.
[0003] Japanese Patent Application Laid-Open No. 2022-174239
[0004] The receiver generally performs channel estimation using a reference signal for channel estimation in the frequency domain where the reference signal is transmitted, and also performs interpolation using linear interpolation, minimum mean square error (MMSE), etc. for the frequency domain around the reference signal to perform channel estimation. In recent years, research has been conducted on methods using supervised learning instead of linear interpolation and MMSE, etc., but there is a problem that it is difficult to prepare learning data based on actual data for all frequency domains.
[0005] Therefore, an object of the present invention is to provide a technique that enables easy channel estimation using machine learning.
[0006] An information processing device according to one aspect of the present invention includes: an acquisition unit that acquires a channel estimate of a reference signal region to which a reference signal is transmitted, from among a plurality of wireless resource regions included in a predetermined range determined by the frequency and time directions of a wireless signal transmitted from a transmitting device; and a learning processing unit that generates a learning model for channel estimation of a wireless signal by training a learning model using reinforcement learning, in which the channel estimates of each of the plurality of wireless resource regions included in the predetermined range, including the channel estimate of the reference signal region, are used as state values, and a value determined based on the performance value of the wireless signal obtained based on the wireless signals of the plurality of wireless resource regions and the channel estimates of each of the plurality of wireless resource regions is used as a reward.
[0007] According to the present invention, channel estimation using machine learning can be easily performed.
[0008] This figure shows an example configuration of a wireless communication system according to this embodiment. This figure illustrates the wireless frame configuration used in the wireless communication system. This figure illustrates the details of the wireless frame configuration. This figure shows an example hardware configuration of a communication device and an information processing device. This figure shows an example of the functional block configuration of a communication device. This figure shows an example of the functional block configuration of an information processing device. This figure shows an example of a learning model used for reinforcement learning. This figure illustrates an action. This figure illustrates the processing procedure when updating the state by applying an action. This flowchart shows an example of reinforcement learning processing performed by an information processing device. This figure illustrates an example of BLER calculation. This flowchart shows an example of a processing procedure in which a communication device performs channel estimation using a learned learning model.
[0009] Embodiments of the present invention will be described with reference to the attached drawings. In each drawing, components denoted by the same reference numerals have the same or similar configurations.
[0010] <System Configuration> Figure 1 shows an example of the configuration of the wireless communication system 1 according to this embodiment. The wireless communication system 1 includes a communication device 10a, a communication device 10b, and an information processing device 20. The communication devices 10a and 10b can communicate with each other using radio signals. In this embodiment, when the communication devices 10a and 10b are not distinguished, they are referred to as communication device 10. When the communication device 10 receives radio signals, it may be called a receiving device. When the communication device 10 transmits radio signals, it may be called a transmitting device.
[0011] Communication devices 10a and 10b may be of the same type or of different types. For example, both communication devices 10a and 10b may be terminals (e.g., smartphones, tablet devices, personal computers, etc.), or one may be a terminal and the other a base station (or access point). The radio signal transmitted from the base station to the terminal is also called the "downlink radio signal." The radio signal transmitted from the terminal to the base station is also called the "uplink radio signal."
[0012] Furthermore, the wireless communication system 1 may be a wireless communication system compliant with 3GPP (Third Generation Partnership Project) (for example, LTE (Long Term evolution), 5G and 6G, etc.) or a wireless communication system compliant with the WLAN (Wireless Local Area Network) standard. In the following description, the communication devices 10a and 10b will be described on the premise that they are a terminal and a base station in a 5G wireless communication system, but this embodiment is not limited to this.
[0013] The information processing device 20 uses reinforcement learning to train the learning model used in the communication device 10. Alternatively, in the wireless communication system 1, the communication device 10 may perform the learning of the learning model itself. In this case, the communication device 10 may also be called an information processing device.
[0014] Figure 2 is a diagram illustrating the radio signals used in wireless communication system 1. Wireless communication system 1 utilizes OFDM (Orthogonal Frequency Division Multiplexing). OFDM is a technology that transmits data using multiple subcarriers. In LTE and 5G, to visually represent the radio resources used for communication, a matrix is often used, as shown in Figure 2, where the vertical direction is the frequency direction and the horizontal direction is the time direction, to represent the radio resources. In 5G, a unit formed by grouping 12 consecutive subcarriers is called a Resource Block (RB). Multiple resource blocks exist within the frequency band used for wireless communication between communication devices 10. In addition, in LTE and 5G, radio resources in the time direction are represented using time units called slots.
[0015] Figure 3 is a diagram illustrating the details of a wireless signal. As shown in Figure 3, one resource block consists of 12 subcarriers. Also, one slot consists of 14 OFDM symbols. The wireless resources in the area represented by one subcarrier and one OFDM symbol are called resource elements (RE). In other words, the area represented by one resource block and one slot shown in Figure 3 (hereinafter referred to as the "unit area") contains 12 × 14 = 168 resource elements. Note that the size of the unit area shown in Figure 3 is just an example. In this embodiment, the unit area may be any size as long as it is an area represented in the frequency direction and the time direction. For example, the unit area may be the same size as one or more CRC blocks, which will be described later.
[0016] Of the 168 resource elements contained within a unit area, some resource elements transmit a reference signal. This reference signal is also called a pilot signal or reference signal. In this embodiment, the description assumes that the reference signal is a DMRS (Demodulation Reference Signal), but this embodiment is not limited to this. In the following description, resource elements to which a reference signal is transmitted are referred to as "DMRS-REs". In the example in Figure 3, there are 20 DMRS-REs in the unit area, but the position and number of DMRS-REs contained within the unit area are arbitrary.
[0017] The communication device 10 performs channel estimation for each resource element in order to demodulate the radio signal transmitted by each resource element. Channel estimation means estimating how much the amplitude and phase of the radio signal transmitted from the transmitting device have changed by the time it is received by the receiving device. Based on the channel estimation result, the receiving device can correctly demodulate the radio signal by correcting the amplitude and phase of the received radio signal (for example, by restoring it to the original amplitude and phase at the time of transmission by the transmitting device).
[0018] Here, the communication device 10 uses DMRS to estimate the channel of DMRS-RE. Furthermore, for resource elements where DMRS does not exist, the communication device 10 performs channel estimation by channel interpolation using the estimated channel values of DMRS-RE. Conventional communication devices 10 performed channel estimation for resource elements where DMRS did not exist by linear interpolation of the estimated channel values of DMRS-RE or by using MMSE. On the other hand, in this embodiment, the communication device 10 performs channel estimation for each resource element in a unit domain using a model generated using reinforcement learning.
[0019] Furthermore, reinforcement learning is performed by training a learning model using the channel estimates of each of the multiple resource elements contained in a unit domain, including the channel estimates of DMRS-RE, as state values, and a value determined based on the performance value of the radio signal (e.g., throughput or error rate) as the reward. The performance value of the radio signal may be obtained by actually demodulating and decoding the radio signal using the radio signals of multiple radio resource domains contained within a predetermined range and the channel estimates of each of the multiple resource elements. Alternatively, the performance value of the radio signal may be calculated according to a predetermined calculation logic based on the radio signals of multiple radio resource domains contained within a predetermined range and the channel estimates of each of the multiple resource elements.
[0020] <Hardware Configuration> Figure 4 shows an example of the hardware configuration of the communication device 10 and the information processing device 20. The communication device 10 and the information processing device 20 include a processor 11 such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), a storage device 12 such as memory (e.g., RAM (Random Access Memory) or ROM (Read Only Memory)), an HDD (Hard Disk Drive) and / or an SSD (Solid State Drive), a communication device 13 that performs wired or wireless communication, and an input / output device 14 that receives input operations and outputs information. The communication device 13 includes an antenna, an RF (Radio Frequency) circuit and a BB (Base Band) circuit, etc.
[0021] <Functional Block Configuration> (Communication Device) Figure 5 is a diagram showing an example of the functional block configuration of the communication device 10. Figure 5 shows the configuration blocks necessary for explaining this embodiment among the functional blocks necessary for the operation of the communication device 10. The communication device 10 includes a storage unit 100, a receiving unit 101, a first estimation unit 102, a second estimation unit 103, and a demodulation unit 104. The storage unit 100 can be realized using a storage device 12 provided by the communication device 10. The receiving unit 101, the first estimation unit 102, the second estimation unit 103, and the demodulation unit 104 can be realized by the processor 11 of the communication device 10 executing a program stored in the storage device 12. The program can be stored in a storage medium. The storage medium in which the program is stored may be a non-transitory computer-readable medium. The non-temporary storage medium is not particularly limited, but may be a storage medium such as a USB (Universal Serial Bus) memory or a CD-ROM (Compact Disc Read-Only Memory). Furthermore, some or all of the receiving unit 101, the first estimation unit 102, the second estimation unit 103, and the demodulation unit 104 may be implemented by hardware such as an ASIC (Application Specific Integrated Circuit) and an FPGA (Field Programmable Gate Array).
[0022] The storage unit 100 stores the model DB (Database) 100a. The model DB 100a is a database that stores data (parameter sets, etc.) of the learning model generated by the information processing device 20.
[0023] The receiving unit 101 receives the radio signal transmitted from the communication device 10 (transmitting device). The receiving unit 101 generates demodulated symbols for each resource element by performing OFDM demodulation processing (FFT: Fast Fourier Transform) on the radio signal (OFDM signal) received from the antenna. The receiving unit 101 also multiplies the channel estimate value for each resource element estimated by the first estimation unit and the second estimation unit by the demodulated symbol for each resource element, and then passes it to the demodulation unit 104.
[0024] The first estimation unit 102 performs channel estimation of a reference signal region to which a reference signal is transmitted, from among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of the radio signal. Here, the unit region is an example of the predetermined range. The resource element is an example of a plurality of radio resource regions included in the predetermined range. The DMRS-RE is an example of a reference signal region.
[0025] The second estimation unit 103 performs channel estimation for each of the multiple wireless resource areas included in a predetermined range, based on the channel estimates of the reference signal area estimated by the first estimation unit 102. More specifically, the second estimation unit 103 performs channel estimation by inputting the channel estimates of the reference signal area into a learning model generated by reinforcement learning, and obtaining the channel estimates for each of the multiple wireless resource areas output from the learning model. The learning model is a model generated by reinforcement learning in the information processing device 20.
[0026] The demodulation unit 104 demodulates the radio signal using the channel estimation results for each of the multiple radio resource areas. More specifically, the demodulation unit 104 extracts data from the demodulated symbols by demodulating the demodulated symbols acquired from the receiving unit 101 based on the modulation scheme.
[0027] (Information Processing Device) Figure 6 shows an example of the functional block configuration of the information processing device 20. The information processing device 20 includes a storage unit 200, an acquisition unit 201, a learning processing unit 202, and a calculation unit 203. The storage unit 200 can be realized using a storage device 12 provided by the information processing device 20. The acquisition unit 201, the learning processing unit 202, and the calculation unit 203 can be realized by the processor 11 of the information processing device 20 executing a program stored in the storage device 12. The program can be stored in a storage medium. The storage medium in which the program is stored may be a computer-readable non-temporary storage medium. The non-temporary storage medium is not particularly limited, but may be a USB memory or a CD-ROM, for example.
[0028] The memory unit 200 stores the model DB 200a. The model DB 200a is a database that stores data (parameter sets, etc.) of the learning model generated by the learning processing unit 202.
[0029] The acquisition unit 201 acquires from the communication device 10 (transmitter) the channel estimate of the reference signal region to which the reference signal is transmitted, from among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of the radio signal transmitted from the communication device 10. Here, the unit region is an example of the predetermined range. The resource element is an example of a plurality of radio resource regions included in the predetermined range. The DMRS-RE is an example of the reference signal region.
[0030] The learning processing unit 202 generates a learning model for channel estimation of wireless signals by training the learning model through reinforcement learning. Reinforcement learning is performed by changing state values according to actions and providing the learning model with rewards corresponding to the state values. The state values are the channel estimates of each of the multiple wireless resource regions included in a predetermined range, including the channel estimate of the reference signal region. The action is to update the channel estimate of each of the multiple wireless resource regions using one of the multiple filtering processes. The reward is a value determined based on the performance value of the wireless signal, which is obtained based on the wireless signals of the multiple wireless resource regions and the channel estimates of each of the multiple wireless resource regions. The performance value of the wireless signal may be, for example, throughput or error rate.
[0031] The calculation unit 203 calculates the performance value of a wireless signal based on the wireless signals of multiple wireless resource areas included in a predetermined range and the estimated channel values of each of the multiple wireless resource areas. The performance value of the wireless signal used by the learning processing unit 202 for reinforcement learning may be the value calculated by the calculation unit 203, or it may be the actual value acquired by the acquisition unit 201 from the communication device 10.
[0032] Furthermore, if the communication device 10 performs the learning of the learning model itself, the communication device 10 may be equipped with the learning processing unit 202 and calculation unit 203 described above.
[0033] <Processing Procedure> (Reinforcement Learning) Figure 7 shows an example of a learning model used for reinforcement learning. The learning model shown in Figure 7 is a model that incorporates an algorithm using a neural network called A3C (Asynchronous Advantage Actor-Critic). As shown in Figure 7, the learning model includes a policy network and a value network. In this embodiment, the algorithm of the learning model is not limited to A3C. An algorithm called A2C (Advantage Actor-Critic) may be used, or other algorithms may be used.
[0034] Here, we will explain the "state," "action," "policy," "state-value," and "reward" used in reinforcement learning. In the following explanation, i and j represent the resource element numbers. The resource elements can be numbered in any way, but for example, in Figure 3, the resource element in the lower left may be number 0, the resource element in the lower right may be number 13, and the resource element in the upper right may be number 167. t represents the number of times the state s has been updated. The initial value of t is 0.
[0035] The "state" is the current value of the channel estimate for each resource element contained in the unit domain. The state is represented by the following equation (1). For convenience, the states represented by equation (1) will be referred to as "state s" and "state s(t)" in the following explanation. The channel estimate for each resource element is a value on the IQ plane, and values on the IQ plane are represented using a complex number (x + yi). That is, the state s of each resource element is represented using two numbers: a real part value x and an imaginary part value y.
[0036] "Action" refers to the process used to change the channel estimates of each resource element contained in a unit domain. "Action" is expressed by the following equation (2). In the following explanation, the actions expressed by equation (2) will be referred to as "action a" and "action a(t)" for convenience. The specific processing method for action a is shown in Figure 8. Figure 8 is a diagram illustrating the actions. As shown in Figure 8, the possible actions in this embodiment are one of nine options: a box filter, two types of bilateral filters, a median filter, two types of Gaussian filters, adding 1 to the value, subtracting 1 from the value, or doing nothing. Since each resource element's state s has two values (x and y), the action is performed on each of the values of x and y. For example, adding 1 to the value means adding 1 to both x and y.
[0037] Figure 9 is a diagram illustrating the processing procedure when updating state s by applying an action. For example, consider the case where a box filter of size 5x5 is applied to resource element RS1 at the positions of subcarrier 9 and OFDM symbol 2. The box filter sets the state s of the resource element to be updated to the average value of the states s of all resource elements contained within a square area centered on the resource element to be updated. In other words, the learning processing unit 202 sets the updated state s of resource element RS1 to the average value of the states s (i.e., current channel estimates) of the 25 resource elements enclosed in the black frame in Figure 9.
[0038] Note that if the resource element is located at the edge of the unit area, the filter area will extend beyond the unit area. In this case, how the updated state s should be calculated may be predetermined for each filter. For example, rules may be established such as considering the value of the part where the resource element does not exist as zero, or considering the value of the part where the resource element does not exist as the average value of the part where the resource element exists.
[0039] "Policy" indicates which action should be applied to each of the plurality of resource elements included in the unit area. "Policy" is represented by the following formula (3). In the following description, the policy represented by formula (3) may be described as "policy π" and "policy π(t)" for convenience. Policy π represents the probability of applying actions 1 to 9 to the state s of each resource element. That is, in policy π, for each of the 168 resource elements, there is a probability value to be applied for each action (for example, for the 0th resource element, the probability of action 1 is 0.01, action 2 is 0.7, action 3 is 0.1,..., action 9 is 0.02; for the 1st resource element, the probability of action 1 is 0.02,...). Also, it means that the action with the highest probability value among actions 1 to 9 is the action to be applied to the state s of each resource element at the current time.
[0040] "State value" is the total value of the rewards that can be obtained when actions are repeated according to the policy in the case of state s, and it means how valuable state s is. The state value is represented by the following formula (4). In the following description, the state value represented by formula (4) may be described as "state value V(s)" for convenience. "Reward" is a value representing the goodness of state s after applying an action. The reward is represented by the following formula (5). In the following description, the reward represented by formula (5) may be described as "reward R" and "reward R(t)" for convenience. r is the immediate reward. The immediate reward means the reward obtained immediately after applying an action. γ is a decay coefficient set between 0 and 1. w i-j is a weight value set so that the value becomes smaller as the distance between resource elements is farther. N is the number of resource elements. In this embodiment, instead of using formula (5) itself, formula (7) which approximates formula (5) is used to determine the reward. Formula (7) will be described later.
[0041] FIG. 10 is a flowchart showing an example of the reinforcement learning process performed by the information processing apparatus 20. The reinforcement learning shown in FIG. 10 is performed using the radio signal received by the communication apparatus 10b.
[0042] In step S100, the acquisition unit 201 acquires, from the communication device 10, the channel estimation value of the DMRS-RE included in the unit area and the radio signal (demodulated symbol) of each resource element included in the unit area. Subsequently, the learning processing unit 202 generates an initial state (state s(0)) of the state s to be input to the learning model. Here, the state s(0) is shown in Equation (6). When the resource element i is a DMRS-RE, the state s(0) is h. h is the actual channel estimation value of the DMRS-RE estimated by the first estimation unit 102 of the communication device 10. When the resource element i is other than the DMRS-RE, the state s(0) is 0. As described above, the state s is expressed using two numbers, the real part value x and the imaginary part value y. Therefore, h includes both the real part value and the imaginary part value.
[0043] In step S101, the learning processing unit 202 inputs the state s to the learning model.
[0044] In step S102, the learning processing unit 202 acquires the policy π output from the policy network of the learning model and the state value V(s) output from the value network.
[0045] In step S103, the learning processing unit 202 updates the state s according to the policy π acquired in step S102. The updated state s can be expressed as state s(t + 1). The learning processing unit 202 applies any one of actions 1 to 9 based on the policy π for each resource element included in the unit area. Note that the learning processing unit 202 does not necessarily need to apply the action with the highest probability among the policies π. The action with the highest probability among the policies π may be applied, or an action randomly selected from actions 1 to 9 may be applied. By applying both the action with the highest probability and actions other than the action with the highest probability, learning can be performed in various action patterns without being biased toward the action with the highest probability, and a better model can be generated.
[0046] The learning processing unit 202 may update the state s of multiple resource elements in order of resource element number, or it may update them simultaneously (in parallel).
[0047] In step S104, if the learning processing unit 202 has updated the state s a predetermined number of times (i.e., "t = predetermined number of times"), it proceeds to step S105. If the learning processing unit 202 has not updated the state s a predetermined number of times, it returns to step S101 and continues learning. The predetermined number of times is arbitrary, but may be set to, for example, 5 times or 10 times.
[0048] In step S105, the calculation unit 203 calculates a performance value of the radio signal from the radio signal (demodulated symbol) of each of the multiple resource elements contained in the unit area and the state s of each of the multiple resource elements (i.e., the current channel estimate). This performance value may be the block error rate (BLER). Since BLER is an error rate, a smaller value means fewer errors.
[0049] Here, an example of a method for calculating BLER will be described. For example, the calculation unit 203 may demodulate and decode the radio signal used in the processing procedure of step S105 by passing the radio signal through an equalizer (also called an equalizer) and a decoder, and calculate the block error rate by checking whether the CRC (Cyclic Redundancy Check) contained in the obtained data is correct or not. For example, assume that a unit area is composed of one CRC block. In this case, the calculation unit 203 may set the BLER of the unit area to 0% if the one CRC block is correctly decoded (the CRC check is successful), and set the BLER of the unit area to 100% if the one CRC block is not correctly decoded (the CRC check fails).
[0050] A CRC block refers to a predetermined area consisting of multiple resource elements used for wirelessly transmitting a bit sequence containing data and a single CRC for checking that data. In other words, demodulating and decoding the wireless signals of each resource element in the area corresponding to a single CRC block yields data and a single CRC for checking that data.
[0051] Furthermore, it is assumed that the unit region is composed of multiple CRC blocks. In this case, the BLER of the unit region may be calculated by summing the values obtained by multiplying the BLER (100% or 0%) of each CRC block by the proportion of each CRC block within the unit region as a coefficient (i.e., a weighted average of the block error rates).
[0052] A specific example will be explained using Figure 11. Figure 11 corresponds to the case where a unit region is composed of multiple CRC blocks. In Figure 11, the unit region is assumed to be a region represented by 12 subcarriers and 14 OFDM symbols (i.e., 168 resource elements). Furthermore, the unit region is assumed to consist of all of CRC block 1, part of CRC block 2, part of CRC block 3, and part of CRC block 4. For the sake of explanation, although not shown in the illustration, CRC blocks 2 to 4 include not only the resource elements of the unit region shown in Figure 11, but also the resource elements of other unit regions adjacent to that unit region.
[0053] Note that Figure 11 is merely an example illustrating the method for calculating BLER, and it is not intended to imply that the size of the CRC blocks or the size of the unit regions in this embodiment are limited to the example shown in Figure 11. When a unit region is composed of multiple CRC blocks, the size of the CRC blocks may be the same as the size of an integer number (i.e., one or more) of unit regions. For example, the size of a CRC block may be the size of two combined unit regions, four combined unit regions, or eight combined CRC blocks.
[0054] Here, the number of resource elements occupied by CRC block 1 within the unit area is 9 × 8 = 72. Also, the number of resource elements occupied by CRC block 2 within the unit area is 9 × 6 = 54. Also, the number of resource elements occupied by CRC block 3 within the unit area is 3 × 8 = 24. The number of resource elements occupied by CRC block 4 within the unit area is 3 × 6 = 18.
[0055] For example, suppose CRC blocks 1 and 3 are correctly decoded, but CRC blocks 2 and 4 are not. In this case, the BLER of the unit region can be calculated as 100% × (72 / 168) + 0% × (54 / 168) + 100% × (24 / 168) + 0% × (18 / 168) = 43% + 0% + 14% + 0% = 57%. Similarly, suppose CRC blocks 1 to 3 are correctly decoded, but CRC block 4 is not. In this case, the BLER of the unit region can be calculated as 100% × (72 / 168) + 100% × (54 / 168) + 100% × (24 / 168) + 0% × (18 / 168) = 43% + 32% + 14% + 0% = 89%.
[0056] In step S106, the learning processing unit 202 calculates the reward R using the performance value of the wireless signal. When the performance value of the wireless signal is BLER, the reward R is determined by the following formula (7). α is a hyperparameter that is set to an arbitrary value before starting model training. h is the same as in equation (6). α is set to prevent the state s of DMRS-RE from deviating significantly from the actual channel estimate h estimated by the first estimation unit 102 of the communication device 10 as training progresses. In other words, in equation (7), the reward R of DMRS-RE is largest when the state s is the same as the channel estimate h. In reinforcement learning, the learning model is trained to output a policy π that yields a larger reward. Therefore, by defining DMRS-RE as in equation (7), it is expected that training will proceed so that action 9 (doing nothing) is applied to DMRS-RE.
[0057] On the other hand, for resource elements other than DMRS-RE, the reward R increases as the block error rate decreases. Therefore, by defining it as in equation (7), it is expected that learning will proceed so that the most appropriate action from actions 1 to 8 is applied to resource elements other than DMRS-RE.
[0058] In addition, in the processing procedure of step S105, instead of BLER, a value such as throughput, which indicates higher wireless signal performance when the value is larger, may be used as the wireless signal performance value. In this case, the expression for resource elements other than DMRS-RE in equation (7) may be, for example, "(wireless signal performance value) 2 It can also be replaced with ".
[0059] In step S107, the learning processing unit 202 trains the model using the advantage function A and the loss function. The advantage function A is determined by the following equation (8). Furthermore, the loss function is determined by the following equations (9) and (10). v represents the value network, and N is the number of resource elements. In other words, the parameters of the value network are primarily updated by the loss function in equation (9). p represents the policy network, and N is the number of resource elements. In other words, the loss function in equation (10) primarily updates the parameters of the policy network.
[0060] In step S108, if the performance value of the wireless signal exceeds a predetermined value, or if the number of times the learning model has been trained reaches a predetermined number, the learning processing unit 202 proceeds to the processing procedure of step S109 to terminate reinforcement learning of the learning model. On the other hand, if the performance value of the wireless signal does not exceed a predetermined value and the number of times the learning model has been trained has not reached a predetermined number, the processing procedure of step S101 continues reinforcement learning of the learning model. If the performance value of the wireless signal is BLER, the learning processing unit 202 proceeds to the processing procedure of step S109 to terminate model learning if the BLER calculated in step S105 is less than the threshold, or if the model has been trained a predetermined number of times. If BLER is greater than or equal to the threshold and the model has not been trained a predetermined number of times, the learning processing unit 202 proceeds to the processing procedure of step S101, returns state s to the initial state (state of equation (6)), and then repeats the processing procedures of steps S101 to S107. The predetermined number of times is arbitrary, but may be, for example, 5,000 times or 10,000 times.
[0061] In step S109, the learning processing unit 202 stores the various parameter values that constitute the learning model in the model DB 100a and terminates the learning process.
[0062] Regarding the reinforcement learning described above, it is conceivable that there are multiple variations in the position of the DMRS-RE within a unit domain. For example, the position of the DMRS-RE as defined in LTE and 5G has multiple predefined variations depending on whether it is a downlink or uplink wireless signal, the state of wireless quality, etc. Therefore, the learning processing unit 202 may prepare a separate learning model for each variation of the DMRS-RE position and train the learning model for each variation of the DMRS-RE position.
[0063] Furthermore, even if the channel estimates for DMRS-RE are the same, it is assumed that the channel estimates for resource elements other than DMRS-RE will not be the same if the locations where the wireless signals are transmitted and received are different. Therefore, the state s (state value) may include location information indicating the location where the wireless signal is transmitted. The learning processing unit 202 may also train the learning model using the state s (state value) containing the location information and the reward. As a result, the learning model will output a channel estimate corresponding to the location where the wireless signal is transmitted, making it possible to perform more accurate channel estimation.
[0064] (Channel Estimation Using a Learned Model) Figure 12 is a flowchart showing an example of a processing procedure in which the communication device 10 performs channel estimation using a learned model. The model DB 100a of the communication device 10 is assumed to store various parameters of the learned model learned by the learning process shown in Figure 9. The second estimation unit 103 is also assumed to be able to perform channel estimation using the learned model by reading various parameters of the learned model from the model DB 100a.
[0065] In step S200, the first estimation unit 102 estimates the channel of the DMRS-RE from the radio signal received by the receiving unit 101. The second estimation unit 103 generates the initial state of state s shown in equation (6).
[0066] In step S201, the second estimation unit 103 inputs the state s into the learning model.
[0067] In step S202, the second estimation unit 103 obtains the policy π output from the policy network of the learning model and the state value V(s) output from the value network.
[0068] In step S203, the second estimation unit 103 updates the state s according to the policy π obtained in step S102. At this time, for each resource element included in the unit domain, the second estimation unit 103 applies the action with the highest probability among actions 1 to 9 based on the policy π. This is because, since the learning model has already completed learning, the action with the highest probability among the policies π is considered to be the most appropriate action.
[0069] In step S204, if the second estimation unit 103 has performed the update of state s a predetermined number of times (i.e., "t = predetermined number of times"), it proceeds to step S205. If the learning processing unit 202 has not performed the update of state s a predetermined number of times, it returns to step S201 and continues updating state s.
[0070] In step S205, the second estimation unit 103 outputs the state s updated in the processing procedures of steps S201 to S204 as the channel estimate value of each resource element in the unit domain.
[0071] In the processing procedure described above, the learning model may be trained using a state s (state value) that includes the location information and a reward. In other words, the state s (state value) may include location information indicating the location where the wireless signal is transmitted. In this case, the second estimation unit may input the state s (state s including the channel estimate of the reference signal domain) and the location information indicating the location where the wireless signal is transmitted to the learning model in the processing procedure of step S201. As a result, the learning model will output a channel estimate corresponding to the location where the wireless signal is transmitted, making it possible to perform channel estimation with higher accuracy.
[0072] <Summary> According to the embodiment described above, a learning model capable of channel estimation is generated by training the model using reinforcement learning. This makes it possible to easily perform channel estimation using machine learning. Therefore, the technology according to this embodiment can contribute to achieving Sustainable Development Goal (SDG) 9, "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."
[0073] The embodiments described above are provided to facilitate understanding of the present invention and are not intended to limit its interpretation. The flowcharts, sequences, elements, and their arrangement, materials, conditions, shapes, and sizes described in the embodiments are not limited to those exemplified and can be modified as appropriate. Furthermore, configurations shown in different embodiments can be partially substituted or combined.
[0074] 1 Wireless communication system, 10a Communication device, 10b Communication device, 11 Processor, 12 Storage device, 13 Communication device, 14 Input / output device, 20 Information processing device, 100 Storage unit, 100a Model DB, 101 Receiving unit, 102 First estimation unit, 103 Second estimation unit, 104 Demodulation unit, 200 Storage unit, 200a Model DB, 201 Acquisition unit, 202 Learning processing unit, 203 Calculation unit
Claims
1. An information processing device comprising: an acquisition unit that acquires a channel estimate of a reference signal region to which a reference signal is transmitted, from among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of a radio signal transmitted from a transmitting device; and a learning processing unit that generates a learning model for channel estimation of a radio signal by training a learning model using reinforcement learning, in which the channel estimates of each of the plurality of radio resource regions included in the predetermined range, including the channel estimate of the reference signal region, are used as state values, and a value determined based on the performance value of the radio signal obtained based on the radio signals of the plurality of radio resource regions and the channel estimates of each of the plurality of radio resource regions is used as a reward.
2. The information processing apparatus according to claim 1, wherein the learning processing unit terminates reinforcement learning of the learning model when the performance value exceeds a predetermined value, or when the number of times the learning model is trained reaches a predetermined number of times.
3. The information processing apparatus according to claim 1, wherein the state value includes location information indicating the location where the wireless signal is transmitted, and the learning processing unit uses the state value including the location information and the reward to train the learning model.
4. The information processing apparatus according to claim 1, wherein the learning processing unit trains the learning model by updating the channel estimates of each of the plurality of wireless resource areas using one of the plurality of filtering processes.
5. An information processing method performed by an information processing device, comprising: a step of obtaining a channel estimate of a reference signal region to which a reference signal is transmitted, from among a plurality of radio resource regions included in a predetermined range determined by the frequency direction and time direction of a radio signal transmitted from a transmitting device; and a step of generating a learning model for channel estimation of a radio signal by training a learning model using reinforcement learning, in which the channel estimates of each of the plurality of radio resource regions included in the predetermined range, including the channel estimate of the reference signal region, are used as state values, and a value determined based on the performance value of the radio signal obtained based on the radio signals of the plurality of radio resource regions and the channel estimates of each of the plurality of radio resource regions is used as the reward.
6. A program to cause a computer to perform the following steps:
1. Obtain a channel estimate of a reference signal region to which a reference signal is transmitted, from among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of a radio signal transmitted from a transmitting device; and 2. Generate a learning model for channel estimation of a radio signal by training a learning model using reinforcement learning, in which the channel estimates of each of the plurality of radio resource regions included in the predetermined range, including the channel estimate of the reference signal region, are used as state values, and a value determined based on the performance value of the radio signal obtained based on the radio signals of the plurality of radio resource regions and the channel estimates of each of the plurality of radio resource regions is used as the reward.
7. A receiving device comprising: a receiving unit that receives a radio signal transmitted from a transmitting device; a first estimation unit that performs channel estimation of a reference signal region to which a reference signal is transmitted, among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of the radio signal; a second estimation unit that performs channel estimation by inputting the channel estimation of the reference signal region into a learning model generated by reinforcement learning, in which the channel estimation values of each of the plurality of radio resource regions, including the channel estimation value of the reference signal region, are used as state values, and the reward is a value determined based on the performance value of the radio signal obtained based on the radio signals of the plurality of radio resource regions and the channel estimation values of each of the plurality of radio resource regions, and obtaining the channel estimation values of each of the plurality of radio resource regions output from the learning model; and a demodulation unit that demodulates the radio signal using the channel estimation results of each of the plurality of radio resource regions.
8. The receiving device according to claim 7, wherein the state value includes location information indicating the location where the wireless signal is transmitted, and the second estimation unit inputs the channel estimate of the reference signal region and the location information to the learning model.
9. An information processing method performed by an information processing device, comprising: receiving a radio signal transmitted from a transmitting device; performing channel estimation of a reference signal region to which a reference signal is transmitted, among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of the radio signal; performing channel estimation by inputting the channel estimation of the reference signal region into a learning model generated by reinforcement learning, in which the channel estimation of each of the plurality of radio resource regions, including the channel estimation of the reference signal region, is used as a state value, and the reward is a value determined based on the performance value of the radio signal obtained based on the radio signals of the plurality of radio resource regions and the channel estimation of each of the plurality of radio resource regions, and obtaining the channel estimation of each of the plurality of radio resource regions output from the learning model; and demodulating the radio signal using the channel estimation results of each of the plurality of radio resource regions.
10. A program to cause a computer to perform the following steps: receiving a radio signal transmitted from a transmitting device; performing channel estimation of a reference signal region to which a reference signal is transmitted, among a plurality of radio resource regions included in a predetermined range determined by the frequency and time directions of the radio signal; performing channel estimation by inputting the channel estimation of the reference signal region into a learning model generated by reinforcement learning, in which the channel estimation values of each of the plurality of radio resource regions, including the channel estimation value of the reference signal region, are used as state values, and the reward is a value determined based on the performance value of the radio signal obtained based on the radio signals of the plurality of radio resource regions and the channel estimation values of each of the plurality of radio resource regions, and obtaining the channel estimation values of each of the plurality of radio resource regions output from the learning model; and demodulating the radio signal using the channel estimation results of each of the plurality of radio resource regions.