Communication system and communication method

The communication system uses Q-value calculation and control units to manage beam patterns dynamically, addressing throughput issues in stratospheric platforms by maintaining optimal performance across varying user densities.

JP2025118218APending Publication Date: 2025-08-13KEIO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024013410
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Communication throughput decreases in areas with a high concentration of user terminals due to varying beam patterns in stratospheric platform communication systems.

Method used

A communication system utilizing a first Q-value calculation unit and a second Q-value calculation unit, along with an update unit and control information calculation unit, to dynamically control beams based on current and future information to maintain optimal communication throughput.

Benefits of technology

The system effectively suppresses decreases in communication throughput by adaptively managing beam patterns, ensuring consistent performance even with varying user distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025118218000001_ABST
    Figure 2025118218000001_ABST
Patent Text Reader

Abstract

To provide a communication system in which reduction in throughput of communication can be inhibited.SOLUTION: A communication system comprises a first Q value calculation unit, a second Q value calculation unit, an update unit, and a control information calculation unit. The first Q value calculation unit includes a first model. The first model outputs a first Q value according to input of current information on a beam to be used for communication between a stratosphere platform and a user terminal. The second Q value calculation unit includes a second model. The second model outputs a second Q value according to input of future information on the beam. The update unit updates the first model on the basis of the first Q value output from the first model and the second Q value output from the second model. The control information calculation unit calculates control information for controlling the beam on the basis of the first Q value output from the first model updated by the update unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a communication system and a communication method. [Background technology]

[0002] A communication system using a stratospheric platform (HAPS: High Altitude Platform Station) is known (see, for example, Non-Patent Document 1). For example, an unmanned aerial vehicle is used as the HAPS. In this communication system, the flying HAPS searches for a user terminal, and communication is carried out between the HAPS and the user terminal. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Yuta NAKAMOTO, Naoki HASEGAWA, Yoshichika OHTA and Naoki SHINOHARA, “A Study on Microwave Power Transfer System to High Altitude Platform Station Considering Rectification Efficiency,” Space Solar Power Generation, vol.6 (2021), pp. 38-41 Summary of the Invention [Problem to be solved by the invention]

[0004] HAPS is required to cover the area where user terminals are located. However, for example, when multiple ranges corresponding to the beam pattern are determined using k-means, the communication throughput for each range varies depending on the distribution of users. In areas with many users, the communication throughput decreases.

[0005] An object of one aspect of the present invention is to provide a communication system and a communication method that can suppress a decrease in communication throughput. [Means for solving the problem]

[0006] A communication system according to one embodiment of the present invention includes a first Q-value calculation unit, a second Q-value calculation unit, an update unit, and a control information calculation unit. The first Q-value calculation unit includes a first model. The first model outputs a first Q-value in response to input of current information related to a beam used for communication between a stratospheric platform and a user terminal. The second Q-value calculation unit includes a second model. The second model outputs a second Q-value in response to input of future information related to the beam. The update unit updates the first model based on the first Q-value output from the first model and the second Q-value output from the second model. The control information calculation unit calculates control information for controlling the beam based on the first Q-value output from the first model updated by the update unit.

[0007] A communication method according to another aspect of the present invention includes: outputting a first Q value from a first model; outputting a second Q value from a second model; updating the first model based on the first Q value output from the first model and the second Q value output from the second model; and calculating control information for controlling a beam based on the updated first Q value output from the first model. The first model outputs the first Q value in response to input of current information regarding a beam used for communication between a stratospheric platform and a user terminal. The second model outputs the second Q value in response to input of future information regarding the beam. [Effects of the Invention]

[0008] One aspect of the present invention provides a communication system and a communication method that can suppress a decrease in communication throughput. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram of a communication system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a schematic diagram illustrating a communication system. [Figure 3] 10A shows the strength of received signal power in the embodiment, and FIG. 10B shows the strength of received signal power in a modified example of the embodiment. [Figure 4] 10(a) and 10(b) are diagrams for explaining the influence of HPAS movement. [Figure 5] 10(a) to 10(d) are diagrams for explaining determination of cells in a coverage area. [Figure 6] FIG. 10 is a diagram for explaining cell determination in a modified example of the present embodiment. [Figure 7] 10(a) and 10(b) are diagrams for explaining parameters relating to a beam pattern. [Figure 8] FIG. 1 is a schematic diagram for explaining a beam control fish in a communication system. [Figure 9] FIG. 1 is a schematic diagram illustrating a first model of an online network. [Figure 10] FIG. 2 illustrates an example of a hardware configuration. [Figure 11] 10(a) and 10(b) are diagrams showing user distribution datasets. [Figure 12] Graphs (a) and (b) show simulation results of throughput when HAPS is stationary. [Figure 13] (a) to (c) show the SINR analysis for different cell configurations. [Figure 14] 2(a) and 2(b) are graphs showing simulation results of throughput when the HAPS is moving. [Figure 15] A comparative analysis of the performance of three reinforcement learning algorithms is presented. DETAILED DESCRIPTION OF THE INVENTION

[0010] [Description of the embodiments of the present disclosure]

[0011] First, embodiments of the present disclosure will be listed and described.

[0012] [1] A communication system according to an embodiment of the present disclosure includes a first Q-value calculation unit, a second Q-value calculation unit, an update unit, and a control information calculation unit. The first Q-value calculation unit includes a first model. The first model outputs a first Q-value in response to input of current information related to a beam used for communication between a stratospheric platform and a user terminal. The second Q-value calculation unit includes a second model. The second model outputs a second Q-value in response to input of future information related to the beam. The update unit updates the first model based on the first Q-value output from the first model and the second Q-value output from the second model. The control information calculation unit calculates control information for controlling the beam based on the first Q-value output from the first model updated by the update unit.

[0013] In the communication system in [1] above, the first model is updated based on both the first Q value and the second Q value, and the beam control information is calculated based on the first Q value output from the updated first model. In this case, the beam can be controlled so as to suppress a decrease in communication throughput.

[0014] [2] The communication system of [1] above may further include a reward calculation unit. The reward calculation unit may calculate a reward based on the first Q value output from the first model and the position of the stratospheric platform. The update unit may update the first model based on the reward calculated by the reward calculation unit and the second Q value. In this case, the beam may be controlled so as to further suppress a decrease in communication throughput.

[0015] [3] The communication system of [2] above may further include a terminal distribution acquisition unit. The terminal distribution acquisition unit may acquire distribution information of user terminals. The reward calculation unit may calculate the reward based on the first Q value output from the first model, the position of the stratospheric platform, and position information of the user terminals. In this case, the beam may be controlled so as to further suppress a decrease in communication throughput.

[0016] [4] The communication system according to any one of [1] to [3] above may further include a target Q value calculation unit. The target Q value calculation unit may calculate a target Q value based on a first Q value output from the first model and a second Q value output from the second model. The update unit may update the first model based on a difference between the first Q value and the target Q value. In this case, the beam may be controlled so as to further suppress a decrease in communication throughput.

[0017] [5] The communication system of [4] above may further include a progress information acquisition unit. The progress information acquisition unit may acquire learning progress information of the first model resulting from updating the first model. The update unit may calculate a target Q value based on the learning progress information acquired by the progress information acquisition unit. In this case, the beam may be controlled so as to further suppress a decrease in communication throughput.

[0018] [6] The communication system according to any one of [1] to [5] above may further include a second model setting unit. The second model setting unit may set the second model to be the same as the first model.

[0019] [7] A communication method according to another aspect of the present disclosure includes: outputting a first Q value from a first model; outputting a second Q value from a second model; updating the first model based on the first Q value output from the first model and the second Q value output from the second model; and calculating control information for controlling a beam based on the updated first Q value output from the first model. The first model outputs the first Q value in response to input of current information regarding a beam used for communication between a stratospheric platform and a user terminal. The second model outputs the second Q value in response to input of future information regarding the beam. [Details of the embodiments of the present disclosure]

[0020] Hereinafter, an embodiment of a communication system according to the present invention will be described in detail with reference to the drawings. In the description of the drawings, the same or corresponding parts will be denoted by the same reference numerals, and duplicated explanations will be omitted.

[0021] First, a schematic configuration of a communication system according to an embodiment of the present disclosure will be described with reference to Fig. 1 to Fig. 10. Fig. 1 is a block diagram of a communication system 1 according to this embodiment. Fig. 2 is a schematic diagram of the communication system 1.

[0022] The communication system 1 is a radio station that uses a stratospheric platform. Hereinafter, the stratospheric platform is referred to as a HAPS (High Altitude Platform Station). The communication system 1 includes a base station 2 and a HAPS 3. The communication system 1 transmits or receives information to a user terminal U, or transmits and receives information corresponding to the user terminal U, via the HAPS 3. In this specification, "communication" includes only transmitting information, only receiving information, and all of transmitting and receiving information. The communication system 1 provides a strong line-of-sight link with one or more user terminals U. The user terminal U is, for example, a mobile communication terminal.

[0023] The base station 2 communicates with the HAPS 3. The base station 2 is located on the ground, for example, as shown in Fig. 2. The base station 2 is, for example, a gateway station.

[0024] The HAPS3 relays signals received from the base station 2. The HAPS3 is, for example, an unmanned aerial vehicle located in the stratosphere as shown in FIG. 2. The altitude at which the HAPS3 is located is, for example, in the range of 20 km to 50 km. The unmanned aerial vehicle used as the HAPS3 is, for example, an airplane. The unmanned aerial vehicle used as the HAPS3 may also be a stratospheric airship. The airplane used as the HAPS3 may also be a solar plane.

[0025] The HAPS3 is equipped with multiple antennas 10 and communicates with a user terminal U via the multiple antennas 10. The HAPS3 communicates with a user terminal U located in a coverage area AR. The HAPS3 receives signals from the user terminal U. The coverage area AR corresponds to a communication area in which communication is possible using the HAPS3. For example, the coverage area AR is a ground area covered by the HAPS3. The diameter of the coverage area AR ranges, for example, from 40 km to 200 km. The coverage area AR includes multiple regions BR. The regions BR correspond to clusters. Setting multiple regions BR in the coverage area AR is called "clustering." The HAPS3 sets a beam pattern corresponding to each of the multiple regions BR. A unit area of the beam pattern on the ground is called a cell CL.

[0026] 3(a) and 3(b) show the strength of received signal power in a coverage area AR. In FIGS. 3(a) and 3(b), the strength of received signal power is shown by SINR (Signal to Interference and Noise Ratio). In FIGS. 3(a) and 3(b), the range with a higher SINR is shown in black, and the range with a lower SINR is shown in white. FIG. 3(a) shows a state in which three cells CL are uniformly configured. FIG. 3(b) shows a state in which nine cells CL are uniformly configured. The nine cells CL shown in FIG. 3(b) are configured in two layers, with three cells CL arranged on the inside and six cells CL arranged on the outside.

[0027] The HAPS 3 moves due to unexpected factors, for example, as shown in Figures 4(a) and 4(b). Figures 4(a) and 4(b) are diagrams for explaining the influence of HPAS movement. For example, the HAPS 3 moves due to wind. In Figure 4(a), the HAPS 3 moves horizontally relative to the ground. In Figure 4(b), the HAPS 3 moves rotationally around a vertical axis. The movement of the HAPS 3 may also move the coverage area AR. As a result, there is a risk that the received signal power from the user terminal U at the HAPS 3 will decrease. The decrease in the strength of the received signal power will increase the number of user terminals U whose communication throughput will decrease.

[0028] Regarding the throughput at the user terminal U, the equation (1)T j k =(b j / k j )*log2(1+γ k ) holds. "J" is the number of antenna arrays and also the number of cells. "j" means the jth cell. "K j " is the number of user terminals U in the j-th cell. "T j k " is the throughput of the kth user terminal U in the jth cell. "b j ” is the bandwidth of the user terminal in the j-th cell. k" is the SINR of the kth user terminal. The smaller the bandwidth, the more gradual the increase in throughput with an increase in SINR. Therefore, if bandwidth is secured, throughput can also be secured. The bandwidth allocated equally to each user terminal U is B / K.

[0029] 1, the communication system 1 includes a search unit 11, a terminal distribution acquisition unit 12, a density calculation unit 13, a division position setting unit 14, a region determination unit 15, a determination unit 16, an antenna control unit 17, a beam control unit 18, and a storage unit 19. In the example shown in this embodiment, the HAPS 3 includes the search unit 11 and the antenna control unit 17. The base station 2 includes the terminal distribution acquisition unit 12, the density calculation unit 13, the division position setting unit 14, the region determination unit 15, the determination unit 16, and the beam control unit 18. As a variation of this embodiment, the HAPS 3 may include at least one of the terminal distribution acquisition unit 12, the density calculation unit 13, the division position setting unit 14, the region determination unit 15, the determination unit 16, and the beam control unit 18, in addition to the search unit 11 and the antenna control unit 17.

[0030] The search unit 11 searches for user terminals U located in the coverage area AR of the HAPS 3. For example, the search unit 11 searches for user terminals U to obtain position information of each user terminal U in the coverage area AR.

[0031] The terminal distribution acquisition unit 12 acquires distribution information of user terminals U in the coverage area AR of HAPS 3. The terminal distribution acquisition unit 12 acquires distribution information indicating the position of each user terminal U. For example, the terminal distribution acquisition unit 12 acquires the search results of user terminals located in the coverage area AR of HAPS 3 from the search unit 11. The terminal distribution acquisition unit 12 may acquire information output from each functional unit in the communication system 1 or from the user terminal U.

[0032] The density calculation unit 13 calculates the density of the number of users in each of multiple search ranges based on the search results acquired by the terminal distribution acquisition unit 12. For example, the density calculation unit 13 calculates the density of the number of users in the search range using equation (3): D=K / S, where "D" is the density of the number of users, "K" is the number of users in the search range, and "S" is the area of the search range.

[0033] The division position setting unit 14 uses the search results to set division positions for dividing the coverage area AR of the stratospheric platform. The division position setting unit 14 sets division positions for dividing multiple search ranges based on the calculation results of the density calculation unit 13. The division position setting unit 14 sets division positions for dividing search ranges that have a lower user density than the search range as division positions for dividing the coverage area AR of the stratospheric platform. The division position setting unit 14 sets division positions at positions in the search range with the lowest user density. For example, the division position setting unit 14 sets division positions for dividing the search range with the lowest user density among the multiple search ranges based on the user density in each of the multiple search ranges.

[0034] For example, the division position setting unit 14 sets a line that divides the search range as the division position. The division line is, for example, a line that divides the area of the search range into two equal parts. For example, the division position setting unit 14 sets a line that divides the search range with the smallest user density as the division position.

[0035] For example, as shown in FIG. 5(a), the search unit 11 scans the coverage area AR in a sector V having an angle ω1 in a counterclockwise direction. When viewed vertically, the sector V has a fan shape with a center C. For example, the center C is the position of the HAPS 3. As a modification of this embodiment, the center C may be a predetermined position other than the position of the HAPS 3. Based on the calculation result of the density calculation unit 13, the search unit 11 finds a search range L1 with the smallest density of users.

[0036] Next, as shown in Figure 5(b), the search unit 11 scans the search range L1 counterclockwise in a sector V having an angle ω2. The search unit 11 finds the search range L2 with the smallest user density based on the calculation result of the density calculation unit 13. The division position setting unit 14 sets the midline of the search range L2 as the division position P.

[0037] The region determination unit 15 determines a plurality of regions BR that are included in the coverage area AR and correspond to the beam pattern from the stratospheric platform, based on the division positions set by the division position setting unit 14. In the example shown in this embodiment, the plurality of regions BR have the same size.

[0038] As shown in FIG. 2, the multiple regions BR are regions that do not overlap one another. Each of the multiple regions BR is located adjacent to an adjacent region BR. As shown in FIG. 5(c), the region determination unit 15 determines the positions of the multiple regions BR so that the division position P set by the division position setting unit 14 is the boundary between two adjacent regions BR. For example, the region determination unit 15 uses the division position P as the boundary between the multiple regions BR, and scans the coverage area AR counterclockwise, sequentially determining J regions BR1, ..., BR2. i ,…,BR j Determine the position of

[0039] In the example shown in this embodiment, the region determination unit 15 calculates the average number of users by dividing the total number of users in the coverage area AR by the number of multiple regions BRs to be determined. The region determination unit 15 determines the positions of the multiple region BRs based on the division positions set by the division position setting unit 14 and the average number of users. The region determination unit 15 determines the positions of the multiple region BRs so that the number of users in each of the multiple region BRs to be determined is equal to the average number of users. As a result, multiple region BRs having the same number of users can be set. The J region BRs are, for example, regions obtained by dividing the coverage area AR by the same number of user terminals U. In this case, the number of user terminals U in each region BR is K / J. As shown in FIG. 5(d), a cell CL is determined so that the maximum half-power beam width covers the user terminals U in the region BR.

[0040] As a modification of this embodiment, the region determination unit 15 may determine multiple regions BR so that each region BR has the same size. Furthermore, as another modification, the region determination unit 15 may determine the positions of multiple regions BR based on the throughput of a user terminal U located in each region BR. The region determination unit 15 may determine the positions of multiple regions BR so that, for example, a predetermined percentile value of the throughput of a user terminal U located in each region BR is equal. For example, the region determination unit 15 may determine the positions of multiple regions BR so that the 5th percentile value of the throughput of a user terminal U located in each region BR is equal. For example, the region determination unit 15 may determine the positions of multiple regions BR so that the 50th percentile value of the throughput of a user terminal U located in each region BR is equal.

[0041] As a modification of this embodiment, the coverage area AR may be divided into a plurality of regions BR1 to BR7. In Fig. 6, the coverage area AR is divided into a coverage area AR1 and a coverage area AR2. The coverage area AR1 is provided along the outer periphery of the coverage area AR2. A division position P1 is set in the coverage area AR1, and four regions BR1, BR2, BR3, and BR4 are determined in counterclockwise order from the division position P1. A division position P2 is set in the coverage area AR2, and three regions BR5, BR6, and BR7 are determined in counterclockwise order from the division position P2.

[0042] The determination unit 16 determines characteristics when a beam pattern is irradiated onto the multiple regions BR determined by the region determination unit 15. The characteristics determined by the determination unit 16 are, for example, communication throughput values. The communication throughput is, for example, the throughput of a user terminal U located in each region BR. For example, the throughput of a user terminal U located in each region BR is calculated by the determination unit 16 based on the SINR of the user terminal U, the number of users, and the allocated bandwidth. The region determination unit 15 determines the positions of the multiple regions BR based on, for example, the communication throughput values determined by the determination unit 16.

[0043] For example, the setting of the division position P by the division position setting unit 14 and the determination by the determination unit 16 may be repeated. For example, a flow in which the setting of the division position P by the division position setting unit 14, the determination of multiple regions BR by the region determination unit 15, and the determination by the determination unit 16 are performed in this order may be repeated. In this case, the division position setting unit 14 sets the division position based on, for example, the determination result of the determination unit 16. The region determination unit 15 determines the multiple regions BR based on, for example, the division position set based on the determination result of the determination unit 16.

[0044] The antenna control unit 17 controls the antenna 10 provided in the HAPS 3. The antenna control unit 17 sets the antenna 10 based on information set by the beam control unit 18. For example, the antenna control unit 17 sets the antenna 10 according to parameters set by the beam control unit 18. The parameters of the antenna 10 are set, for example, by a geometric model. For example, the antenna control unit 17 controls the physical state of the antenna, the tilt of the antenna 10, etc., according to the parameters set by the beam control unit 18.

[0045] The beam control unit 18 sets parameters related to the beam pattern from the HAPS3. FIGS. 7(a) and 7(b) are diagrams for explaining parameters related to the beam pattern. The parameters related to the beam pattern are, for example, antenna parameters. The beam control unit 18 sequentially resets the parameters related to the beam pattern from the HAPS3. The resetting frequency is, for example, every few milliseconds to tens of seconds. The beam control unit 18 determines multiple cells CL that are included in the coverage area AR and correspond to the beam pattern from the HAPS3. In the example shown in this embodiment, the multiple cells CL have the same size. The multiple cells CL are areas that do not overlap with each other. Each of the multiple cells CL is located adjacent to an adjacent cell CL.

[0046] The beam control unit 18 determines the half-power bandwidth “θ” in the vertical direction based on the radius of the cell CL. 3dB " and the horizontal half-power bandwidth "φ 3dB " and the vertical half-power band "θ 3dB " is expressed by the following equation (4). The horizontal half-power bandwidth "φ 3dB " is expressed by the following equation (5).

number

number

[0047] "h" is the altitude of HAPS3. "g" is the distance between the center of cell CL and HAPS3 when viewed vertically. In other words, "g" is the distance between the center of cell CL and the projected position of HAPS3 on the ground. "r" is the radius of cell CL.

[0048] The beam control unit 18 calculates the change in vertical tilt "Δθ" based on the center position of the cell CL before the movement and the center position of the cell CL after the movement. tilt " and the change in horizontal tilt "Δφ tilt " is obtained. Here, cell CL A From Cell CL B The change in vertical tilt “Δθ tilt " is expressed by the following equation (6). The change in horizontal tilt "Δφ tilt " is expressed by the following equation (7). "g A " is Cell CL A is the distance between the center of the map and the projected position of HAPS3 on the ground. B " is Cell CL B is the distance between the center of the map and the projected position of HAPS3 on the ground. a " and "y a " is Cell CL A It corresponds to the center coordinate of "x b " and "y b " is Cell CL B corresponds to the coordinates of the center of

number

number

[0049] After clustering is performed by the division position setting unit 14 and the region determination unit 15 and the antenna parameters are designed using the geometric coverage model, the beam control unit 18 fine-tunes the antenna parameters so as to maximize the throughput of the user terminal U in each cell CL. The antenna parameters are the vertical half-power width “θ 3dB" and the horizontal half-power width "φ 3dB " and the change in vertical tilt "Δθ tilt " and the change in horizontal tilt "Δφ tilt The beam control unit 18 fine-tunes the antenna parameters by, for example, DQN (deep Q-Network) so as to maximize the throughput of the user terminal U in each cell CL.

[0050] The beam control unit 18 fine-tunes the antenna parameters by reinforcement learning. As shown in FIG. 8 , the beam control unit 18 includes an online network 21, a target network 22, and a calculation information acquisition unit 23. The beam control unit 18 includes a first Q value calculation unit 31, a second Q value calculation unit 32, a behavior calculation unit 33, a reward calculation unit 34, a target Q value calculation unit 35, an update unit 36, a progress information acquisition unit 37, a control information calculation unit 38, and a second model setting unit 39. For example, the online network 21 includes the first Q value calculation unit 31, the behavior calculation unit 33, the update unit 36, and the progress information acquisition unit 37. The target network 22 includes the second Q value calculation unit 32, the reward calculation unit 34, and the target Q value calculation unit 35.

[0051] The calculation information acquisition unit 23 acquires information input to the beam control unit 18. The calculation information acquisition unit 23 acquires current state information as current information related to the beam used for communication between the stratospheric platform and the user terminal. For example, the calculation information acquisition unit 23 acquires the current state information from the storage unit 19. For example, the calculation information acquisition unit 23 acquires distribution information of the user terminal U from the terminal distribution acquisition unit 12. The storage unit 19 stores current state information and future state information. The storage unit 19 stores and periodically updates the distribution information of the user terminal U acquired from the terminal distribution acquisition unit 12, the current antenna parameters, the action set immediately before, and the reward calculated by the reward calculation unit 34.

[0052] The current state information S acquired by the calculation information acquisition unit 23 t For example, the current distribution information Z tand the current antenna parameters X t and the reward in the current environment, r t and the immediately preceding action a t―1 "t" indicates the time step. The antenna parameter X t is a function of the number of parameters of the antenna arrays, "J", and the number of antenna arrays, "M". When the current time step is t, the current state information S t is S t =[Z t ,X t ,r t ,a t―1 ] matrix. The distribution information Z of user terminal U t is a discretization of the area where users are distributed into a lattice matrix.

[0053] action a t are discrete values that represent the change in antenna parameters, and a t =[Δφ tilt ,Δθ tilt ,Δφ 3dB ,Δθ 3dB ] matrix. The reward r t For example, r t =R(S t ,a t )=|Ku t | / K. "R" is the reward function. "K" is the total number of user terminals U. "u t ” is the number of user terminals U whose throughput is below the threshold τ at time step t. This definition of the reward function aims to minimize the number of users with low throughput.

[0054] The first Q value calculation unit 31 includes a first model. The second Q value calculation unit 32 includes a second model. The first Q value calculation unit 31 calculates a first Q value using the first model. The second Q value calculation unit 32 calculates a second Q value using the second model.

[0055] The first model outputs a first Q value in response to input of current information regarding a beam used for communication between the stratospheric platform and the user terminal. The second model outputs a second Q value in response to input of future information regarding the beam. The first and second models, for example, have the same structure. As shown in FIG. 9, the first and second models include a positional encoding 40, an embedding layer 41, an encoder 42, a decoder 43, and a softmax layer 44. In the first and second models, the positional encoding 40, the embedding layer 41, the encoder 42, the decoder 43, and the softmax layer 44 are arranged in this order.

[0056] The first Q value calculation unit 31 calculates the state information S t The first Q value is output in response to the input. The first Q value output from the first Q value calculation unit 31 corresponds to the predicted current Q value and is expressed by equation (8). "p" indicates the value for each action. "A" indicates the number of types of actions.

number

[0057] The behavior calculation unit 33 sets a predicted behavior based on the predicted Q value output from the first Q value calculation unit 31. For example, the behavior calculation unit 33 sets the behavior with the highest value in the predicted Q value as the predicted behavior. The behavior calculation unit 33 predicts an appropriate behavior at the current time step t. The predicted behavior is stored in the storage unit 19. The predicted behavior is expressed by equation (9).

number

[0058] The online network 21 is trained for "T" time steps per training epoch. Every epoch takes "E" time steps. The training trains the first model. At the start of each epoch, the environment is initialized and the weighting in the first model is updated. In each training, the state information St acquired by the calculation information acquisition unit 23 is input to the first model, a predicted Q value is output from the first Q value calculation unit 31, and a predicted behavior is set by the behavior calculation unit 33.

[0059] For example, the first model is trained using the ε-greedy algorithm. The ε-greedy algorithm is used to balance exploring new actions and utilizing current knowledge to identify the optimal policy in reinforcement learning tasks. With probability ε, the predicted action is randomly set from the action space. For example, agent-essential random search is performed to identify new states and actions that are important for learning in unknown environments. Meanwhile, with probability 1-ε, the predicted Q-value is output and the optimal predicted action is set according to the learned policy of the first model.

[0060] The reward calculation unit 34 calculates a reward based on the first Q value output from the first model and the position of the stratospheric platform. The reward calculation unit 34 calculates a reward based on the first Q value output from the first model, the position of the stratospheric platform, and distribution information of user terminals U. The reward calculation unit 34 simulates the environment based on the action calculated by the action calculation unit 33, and calculates a feedback reward r obtained in the environment by executing the predicted action and next state information S t+1 The second Q value calculation unit 32 calculates the state information S output from the reward calculation unit 34. t+1 The second Q value calculation unit 32 outputs a second Q value in response to the input of the second Q value calculation unit 32. The second Q value output from the second Q value calculation unit 32 corresponds to a predicted future Q value and is expressed by equation (10). The future Q value corresponds to an optimal Q value.

number

[0061] Here, "A" is the number of actions. The target Q value calculation unit 35 calculates a target Q value based on the first Q value output from the first model and the second Q value output from the second model. The target Q value calculation unit 35 calculates a target Q value based on the predicted Q value output from the second Q value calculation unit 32 and the reward r calculated in the reward calculation unit 34. The target Q value is calculated by equation (11) in the proposed TFRL (Transformer reinforcement learning).

number

[0062] Here, "e" indicates the current training epoch number, "E" is the number of training epochs, and "α" indicates the learning rate. The fitness function is calculated by Equation (12).

number

[0063] "X best ” is the antenna parameter corresponding to the maximum reward in the current history. The fitness metric, introduced to allow the model to gradually approach the current best solution when no superior solution can be identified, prevents the model from failing to converge. In equation (11), as the training epochs increase, the influence of the reward r decreases, and the influence of the future Q-value Q t+1 and the effect of fitness increases gradually.

[0064] The update unit 36 updates the first model based on the first Q value output from the first model and the second Q value output from the second model. The update unit 36 updates the first model based on the difference between the first Q value and the target Q value. For example, the update unit 36 reflects the square error between the first Q value and the target Q value in the first model.

[0065] Compared to traditional formulations that employ fixed discount factors, this dynamic adaptive design allows the online network 21 to rely on rewards in the current environment for learning when the target network 22 is trained. Once the target network 22 is stabilized and can reliably predict future rewards, the online network 21 relies more on future rewards and fitness for learning.

[0066] The progress information acquisition unit 37 acquires learning progress information of the first model resulting from the update of the first model. For example, the update unit 36 calculates a target Q value based on the learning progress information acquired by the progress information acquisition unit 37.

[0067] The control information calculation unit 38 calculates control information for controlling the beam based on the first Q value output from the first model updated by the update unit 36. For example, the control information calculation unit 38 calculates antenna parameters as the control information for controlling the beam.

[0068] The second model setting unit 39 sets the second model based on the first model. The second model setting unit 39 may set the second model to the same model as the first model. For example, the second model setting unit 39 sets parameters in the first model to parameters in the second model. These parameters are, for example, parameters in at least one of the embedding layer 41, the encoder 42, the decoder 43, and the softmax layer 44. The second model setting unit 39 periodically updates the second model to ensure stability during training. For example, the second model setting unit 39 updates the weights in the second model at regular intervals. After the second model is updated, the second Q value calculation unit 32 calculates a second Q value using the second model.

[0069] Next, the hardware configuration of the base station 2 and the HAPS 3 will be described with reference to Fig. 10. Fig. 10 is a diagram illustrating an example of the hardware configuration of the base station 2 and the HAPS 3.

[0070] In the communication system 1, the base station 2 and the HAPS 3 each include a processor 101, a main memory device 102, an auxiliary memory device 103, a communication device 104, an input device 105, and an output device 106. The base station 2 and the HAPS 3 each include one or more computers configured with this hardware and software such as programs. The search unit 11, the terminal distribution acquisition unit 12, the density calculation unit 13, the division position setting unit 14, the region determination unit 15, the determination unit 16, the antenna control unit 17, and the beam control unit 18 may each be configured with one computer or multiple computers. The base station 2 and the HAPS 3 are realized in cooperation with hardware.

[0071] When the search unit 11, terminal distribution acquisition unit 12, density calculation unit 13, division position setting unit 14, region determination unit 15, judgment unit 16, antenna control unit 17, and beam control unit 18 are configured by multiple computers, these computers may be connected locally or via a communication network such as the Internet or an intranet. This connection logically constructs a single search unit 11, terminal distribution acquisition unit 12, density calculation unit 13, division position setting unit 14, region determination unit 15, judgment unit 16, antenna control unit 17, and beam control unit 18.

[0072] The processor 101 executes an operating system, application programs, etc. The main memory device 102 is composed of a read-only memory (ROM) and a random access memory (RAM). For example, at least some of the various functional units of the base station 2 and the HAPS 3 can be realized by the processor 101 and the main memory device 102.

[0073] The auxiliary storage device 103 is a storage medium configured with a hard disk, a flash memory, etc. The auxiliary storage device 103 generally stores a larger amount of data than the main storage device 102. For example, at least a part of the terminal distribution acquisition unit 12 can be realized by the auxiliary storage device 103.

[0074] The communication device 104 is configured by a network card or a wireless communication module. For example, at least a part of the terminal distribution acquisition unit 12 can be realized by the communication device 104. The input device 105 is configured by an input port, a keyboard, a mouse, a touch panel, etc. For example, at least a part of the terminal distribution acquisition unit 12 can be realized by the input device 105. The output device 106 is configured by an output port, a display, a projection device such as a projector, etc.

[0075] The auxiliary storage device 103 stores in advance a program and data necessary for processing. This program causes a computer to execute each functional element of the base station 2 and the HAPS 3. This program causes the computer to execute, for example, each process performed in a communication method described below. As the communication method, for example, the process performed in the communication system described above is performed. This program may be provided after being recorded on a tangible recording medium such as a CD-ROM, a DVD-ROM, or a semiconductor memory. This program may also be provided as a data signal via a communication network.

[0076] Next, the effects of the communication system and communication method according to the above-described embodiment will be described.

[0077] In the communication system 1, the first model is updated based on both the first Q value and the second Q value, and beam control information is calculated based on the first Q value output from the updated first model. The first Q value is output in response to input of current information about the beam, and the second Q value is output in response to input of future information about the beam. In this case, the beam can be controlled so as to suppress a decrease in communication throughput.

[0078] In the example shown in this embodiment, the communication system 1 further includes a reward calculation unit 34. The reward calculation unit 34 calculates a reward based on the first Q value output from the first model and the position of the stratospheric platform. The update unit 36 updates the first model based on the reward calculated by the reward calculation unit 34 and the second Q value. In this case, the beam can be controlled so as to further suppress a decrease in communication throughput.

[0079] In the example shown in this embodiment, the communication system 1 further includes a terminal distribution acquisition unit 12. The terminal distribution acquisition unit 12 may acquire distribution information of user terminals U. The reward calculation unit 34 calculates the reward based on the first Q value output from the first model, the position of the stratospheric platform, and position information of the user terminals. In this case, the beam can be controlled so as to further suppress a decrease in communication throughput.

[0080] In the example shown in this embodiment, the communication system 1 further includes a target Q value calculation unit 35. The target Q value calculation unit 35 calculates a target Q value based on the first Q value output from the first model and the second Q value output from the second model. The update unit 36 updates the first model based on the difference between the first Q value and the target Q value. In this case, the beam can be controlled so as to further suppress a decrease in communication throughput.

[0081] In the example shown in this embodiment, the communication system 1 may further include a progress information acquisition unit 37. The progress information acquisition unit 37 acquires learning progress information of the first model resulting from updating the first model. The update unit 36 may calculate a target Q value based on the learning progress information acquired by the progress information acquisition unit 37. In this case, the beam can be controlled so as to further suppress a decrease in communication throughput.

[0082] In the example shown in this embodiment, the communication system 1 further includes a second model setting unit 39. The second model setting unit 39 sets the second model to the same model as the first model.

[0083] Next, the simulation results of the throughput performance of the user terminal U in the communication system 1 and the comparative example will be described with reference to FIGS.

[0084] In this simulation, the altitude of the HAPS3 was set to 20 km, and the coverage radius of each HAPS3 was set to 20 km. The transmit power of each HAPS3 was 43 dBm, and the transmit frequency was 2 GHz. The bandwidth of each antenna array was 20 MHz. Furthermore, it was assumed that 18 surrounding HAPS3s would cause interference to each HAPS3. Figures 11(a) and 11(b) show the user distribution datasets in Sendai and Nagoya, respectively. These user distribution datasets are non-uniform. In these user distribution data, when cells CL are set uniformly, the number of users per cell varies significantly. The bandwidth of each user terminal U located in cell CL with a large number of users is small. The HAPS system environment was simulated using Python, and the proposed DQN method was implemented using TensorFlow.

[0085] Figures 12(a) and 12(b) are graphs showing simulation results of throughput when HAPS is stationary. Figure 12(a) shows the CDF (cumulative distribution function) of the throughput of a user terminal U in Sendai. Figure 12(b) shows the CDF of the throughput of a user terminal U in Nagoya. The throughput threshold is defined as the median of the throughput in a uniform cell configuration. In Figures 12(a) and 12(b), data DA1 is data when cells CL are uniformly arranged in a coverage area AR, and data DA2 is data using only equivalent clustering. Data DA3 is data based on K-means clustering, data D4 is data using MFDDQN, and data DA5 is data using a deep reinforcement learning evolutionary algorithm (DRLEA). Data DA6 is data using TFRL, and data DA7 is data using EC-TFRL.

[0086] EC-TFRL is a communication method using the equivalent clustering and reinforcement learning shown in this embodiment, which uses a search unit 11, a terminal distribution acquisition unit 12, a density calculation unit 13, a division position setting unit 14, a region determination unit 15, a determination unit 16, an antenna control unit 17, and a beam control unit 18. TFRL is a communication method using only the reinforcement learning shown in this embodiment, which uses the antenna control unit 17 and the beam control unit 18, without using equivalent clustering.

[0087] As can be seen from Figures 12(a) and 12(b), the proposed EC-TFRL exhibited better throughput performance than the other methods. The proposed TFRL achieved better throughput performance near the threshold TH in Sendai than the other methods except for the proposed EC-TFRL.

[0088] 13(a) to 13(c) show the SINR distributions of different cell configurations. FIG. 13(a) shows the SINR distribution of a uniform cell configuration. FIG. 13(b) shows the SINR distribution of the proposed TFRL. FIG. 13(c) shows the SINR distribution of the proposed EC-TFRL. In TFRL and EC-TFRL, user terminals U within the coverage area AR are accurately covered in each cell CL. In EC-TFRL, the cell CL located in the center is C is the cell CL located at the center of the TFRL. C It is thought that the smaller equivalent clustering reduces the action search space, thereby reducing the probability that reinforcement learning will fall into a suboptimal solution.

[0089] Figures 14(a) and 14(b) show the CDF of the throughput of the user terminal U when HAPS3 moves. The simulation was performed for HAPS3 moving 5 km to the left and rotating 30°. In Figures 14(a) and 14(b), data DA11 is data when cells CL are uniformly arranged in the coverage area AR, and data DA12 is data using only equivalent clustering. Data DA13 is data based on K-means clustering, data D14 is data using TFRL, and data DA15 is data using EC-TFRL.

[0090] In this case, the proposed EC-TFRL and TFRL showed better throughput performance than other methods. The number of user terminals U near the throughput threshold was significantly smaller in the proposed EC-TFRL and TFRL, and high throughput was observed.

[0091] Fig. 15 shows a comparative analysis of the performance of three reinforcement learning algorithms, TFRL, MFDQN, and DRLEA. In Fig. 15, data D21 shows data when MFDQN is used. Data D22 shows data when DRLEA is used. In Fig. 15, each algorithm is evaluated based on the total reward per epoch.

[0092] The TFRL learning curve shows the best average sum of reward performance throughout training. TFRL consistently outperforms other methods, with significant improvement over the past 20 epochs. This suggests that TFRL not only converges effectively, but also achieves higher levels of cumulative reward on average. Compared to MFDQN, TFRL explores the action space more comprehensively, reducing the tendency to fall into suboptimal states.

[0093] The above describes embodiments and modifications of the present invention, but the present invention is not necessarily limited to the above-described embodiments and modifications, and various modifications are possible without departing from the spirit of the present invention.

[0094] For example, the communication system 1 may include a plurality of base stations 2. The communication system 1 may include a plurality of HAPSs 3. [Explanation of symbols]

[0095] 1...communication system, 3...HAPS, 12...terminal distribution acquisition unit, 31...first Q value calculation unit, 32...second Q value calculation unit, 34...reward calculation unit, 35...target Q value calculation unit, 36...update unit, 37...progress information acquisition unit, 38...control information calculation unit, 39...second model setting unit, r...reward, U...user terminal, Zt...distribution information.

Claims

1. a first Q-value calculation unit including a first model that outputs a first Q-value in response to input of current information related to a beam used for communication between the stratospheric platform and the user terminal, the first Q-value calculation unit calculating the first Q-value using the first model; a second Q value calculation unit including a second model that outputs a second Q value in response to input of future information regarding the beam, and that calculates the second Q value using the second model; an updating unit that updates the first model based on a first Q value output from the first model and a second Q value output from the second model; A communication system comprising: a control information calculation unit that calculates control information for controlling the beam based on a first Q value output from the first model updated by the update unit.

2. a reward calculation unit that calculates a reward based on a first Q value output from the first model and a position of the stratospheric platform; The communication system according to claim 1 , wherein the update unit updates the first model based on the reward calculated by the reward calculation unit and the second Q value.

3. Further, a terminal distribution acquisition unit that acquires distribution information of the user terminals is provided, The communication system according to claim 2 , wherein the reward calculation unit calculates the reward based on a first Q value output from the first model, the position of the stratospheric platform, and distribution information of the user terminals.

4. a target Q value calculation unit that calculates a target Q value based on a first Q value output from the first model and a second Q value output from the second model, The communication system according to claim 1 , wherein the update unit updates the first model based on a difference between the first Q value and the target Q value.

5. a progress information acquisition unit that acquires learning progress information of the first model due to the update of the first model, The communication system according to claim 4 , wherein the update unit calculates the target Q value based on the learning progress information acquired by the progress information acquisition unit.

6. The communication system according to claim 1 , further comprising a second model setting unit that sets the second model to the same model as the first model.

7. outputting a first Q-value from a first model that outputs a first Q-value in response to input of current information regarding a beam used for communication between the stratospheric platform and the user terminal; outputting a second Q value from a second model that outputs a second Q value in response to input of future information about the beam; updating the first model based on a first Q value output from the first model and a second Q value output from the second model; and calculating control information for controlling the beam based on a first Q value output from the updated first model.