Method and device for on-demand service coverage of ocean communication based on low-orbit satellite beam hopping

By using low-orbit satellite beam skipping technology and hierarchical heterogeneous multi-agent deep reinforcement learning, the allocation of wave positions and resources is optimized, solving the problems of low spectrum utilization and severe beam interference in marine communications, and achieving on-demand coverage and efficient resource utilization.

CN119232236BActive Publication Date: 2025-11-28BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411346374.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-11-28
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Existing satellite communication systems suffer from low spectrum utilization, severe beam interference, high system complexity, and inflexible resource allocation in marine communication scenarios, making it difficult to meet the needs of uneven user distribution and rapidly changing business requirements.

Method used

By employing low-Earth orbit (LEO) satellite-based hopping beam technology and combining it with a hierarchical heterogeneous multi-agent deep reinforcement learning method, the location of user terminals is determined through TDOA positioning. The number of beam positions and the center point location are optimized, and an optimization model for the joint hopping beam of multiple LEO satellites and the transmit power of user terminals is constructed to achieve on-demand coverage and resource allocation.

Benefits of technology

It improved the coverage performance and resource utilization efficiency of marine communication networks, suppressed beam interference, and enhanced network resource utilization efficiency and system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119232236B_ABST
    Figure CN119232236B_ABST
Patent Text Reader

Abstract

The application discloses a low-orbit satellite beam hopping-based marine communication on-demand service coverage method and device, relates to the technical field of satellite communication, and comprises the following steps: determining the positions of user terminals on the sea through a low-orbit satellite; determining the number of wave positions and the positions of candidate wave position center points according to the positions of the user terminals; taking the low-orbit satellite as a relay for communication between each user terminal and a ground station based on the positions of the user terminals and the positions of the candidate wave position center points, and constructing an optimization model of multi-low-orbit satellite joint beam hopping and user terminal transmission power; and in each time slot, a hierarchical heterogeneous multi-agent deep reinforcement learning method is adopted to solve the optimization variables of the optimization model in a single time slot, wherein the optimization variables comprise the uplink transmission power of the user terminal, the beam hopping pattern of the low-orbit satellite and subchannel allocation. The application can improve the service coverage performance and wireless resource utilization rate of marine communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of satellite communication technology, and in particular to a method and apparatus for on-demand marine communication service coverage based on low-orbit satellite hopping beams. Background Technology

[0002] Fifth-generation mobile communication technology (5G) provides powerful access capabilities for densely populated areas. However, it faces challenges in network deployment and operation and maintenance costs in remote, sparsely populated areas, polar regions, and oceans. Satellite communication networks offer advantages such as wide coverage, flexible and highly directional beams, and communication unrestricted by geographical conditions. They can provide effective services in remote mountainous areas, the air, deserts, and oceans. Furthermore, satellite communication can be applied in special scenarios such as military applications and emergency communications.

[0003] Currently, satellite communication has become an important research direction for the sixth-generation mobile communication standard (6G). Satellites can be categorized by application field into communication satellites, remote sensing satellites, and navigation satellites; and by orbital altitude into low Earth orbit (LEO) satellites, medium Earth orbit (MEO) satellites, and geostationary orbit (GEO) satellites. In recent years, the number of LEO satellite launches has increased significantly, and they are gradually becoming a major component of satellite communication. Operating at altitudes of 500-2000 kilometers, they offer shorter propagation delays, higher data rates, and lower launch costs compared to GEO satellites, making them attractive.

[0004] For marine communication scenarios, the spatial and temporal distribution of service demands is uneven due to the uneven distribution of users on the sea surface and the potential for sparseness, as well as the rapid movement of terminals within the coverage area. In traditional static multi-beam systems, the service position of the beam does not change with time relative to the satellite or changes periodically in a specific pattern. There is still considerable room for improvement in network coverage performance and the utilization of spatial, temporal, and spectrum resources. Summary of the Invention

[0005] The purpose of this application is to provide a method and apparatus for on-demand marine communication coverage based on low-orbit satellite hopping beams, which can improve the coverage performance and resource utilization efficiency of marine communication networks.

[0006] To achieve the above objectives, this application provides the following solution:

[0007] Firstly, this application provides a method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams, including:

[0008] determining positions of user terminals on the ocean via low earth orbit satellites;

[0009] determining a number of wave positions and positions of candidate wave position center points according to the positions of the user terminals;

[0010] constructing an optimization model of joint low earth orbit satellite hopping wave beams and user terminal transmission power based on the positions of the user terminals and the positions of the candidate wave position center points, taking the low earth orbit satellites as relays for communication between the user terminals and ground stations;

[0011] solving, in each time slot, optimization variables of the optimization model in the single time slot by using a hierarchical heterogeneous multi-agent deep reinforcement learning method, the optimization variables including uplink transmission power of a corresponding user terminal output by a user terminal agent, a hopping wave beam pattern of a corresponding low earth orbit satellite output by a low earth orbit satellite agent, and subchannel allocation of the user terminal.

[0012] In a second aspect, the present application provides a low earth orbit satellite hopping wave beam based on-demand service coverage device for ocean communication, the low earth orbit satellite hopping wave beam based on-demand service coverage device for ocean communication comprising:

[0013] a user terminal position determination module configured to determine positions of user terminals on the ocean via low earth orbit satellites;

[0014] a wave position determination module configured to determine a number of wave positions and positions of candidate wave position center points according to the positions of the user terminals;

[0015] an optimization model construction module configured to construct an optimization model of joint low earth orbit satellite hopping wave beams and user terminal transmission power based on the positions of the user terminals and the positions of the candidate wave position center points, taking the low earth orbit satellites as relays for communication between the user terminals and ground stations;

[0016] a solving module configured to solve, in each time slot, optimization variables of the optimization model in the single time slot by using a hierarchical heterogeneous multi-agent deep reinforcement learning method, the optimization variables including uplink transmission power of a corresponding user terminal output by a user terminal agent, a hopping wave beam pattern of a corresponding low earth orbit satellite output by a low earth orbit satellite agent, and subchannel allocation of the user terminal.

[0017] According to the specific embodiments provided by the present application, the following technical effects are disclosed:

[0018] The application provides a low-orbit satellite beam hopping based marine communication on-demand service coverage method and device, the number of wave positions and the positions of candidate wave position center points are determined according to the positions of each user terminal, the distribution of multiple low-orbit satellite coverage tasks can be realized, meanwhile, excessive coverage overlap is avoided, and the on-demand coverage capability and resource utilization efficiency are improved; the low-orbit satellite is used as a relay for communication between each user terminal and a ground station based on the positions of each user terminal and the positions of each candidate wave position center point, an optimization model of joint low-orbit satellite beam hopping and user terminal transmission power is constructed, in each time slot, a hierarchical heterogeneous multi-agent deep reinforcement learning method is used to solve the optimization variables of the optimization model in the single time slot, the optimization variables include the uplink transmission power of the corresponding user terminal output by the user terminal agent, the beam hopping pattern of the corresponding low-orbit satellite output by the low-orbit satellite agent and the subchannel allocation of the user terminal, the low-orbit satellite and the user terminal are optimized according to the obtained beam hopping pattern scheme and the uplink transmission power, the user service on-demand coverage is realized, and the beam interference is further suppressed and the network resource utilization efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0020] Figure 1 A flowchart of a low-orbit satellite beam hopping based marine communication on-demand service coverage method provided by an embodiment of the present application is shown in the figure.

[0021] Figure 2 A flowchart of candidate beam center point determination provided by an embodiment of the present application is shown in the figure.

[0022] Figure 3 A marine communication on-demand coverage scenario provided by an embodiment of the present application is shown in the figure.

[0023] Figure 4 A hierarchical heterogeneous multi-agent optimization algorithm framework diagram provided by an embodiment of the present application is shown in the figure.

[0024] Figure 5 A local network and hybrid network composition diagram provided by an embodiment of the present application is shown in the figure.

[0025] Figure 6 A functional module schematic diagram of a low-orbit satellite beam hopping based marine communication on-demand service coverage device provided by another embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of the present application.

[0027] The above purposes, features and advantages of the present application will be more apparent and understandable. The present application will be further described in detail below with reference to the drawings and specific embodiments.

[0028] For the marine communication scenario, using the traditional multi-beam technology will bring a series of deficiencies to the communication system: first, in terms of spectrum utilization, due to the need to maintain a certain distance between different beams, i.e. a certain frequency interval is needed between beams, which will result in the waste of part of the spectrum resources; multi-beam technology is difficult in interference management. When multiple beams transmit on the same spectrum resource at the same time, the adjacent beams are too close to each other, which will cause inter-beam interference, and interference management strategies such as frequency reuse and interference cancellation technology need to be taken, and the traditional multi-beam technology needs complex antenna arrays and signal processing algorithms to realize the transmission and reception of multiple beams, which will increase the complexity and cost of the system, and may require higher requirements for hardware and software; on the other hand, in order to reasonably save resources and avoid interference between multiple users, the user uplink transmission power should be reduced as much as possible, so as to improve the system throughput of the communication system and make efficient use of system resources.

[0029] Based on the actual background, someone has proposed a beam hopping (BH) technology. Beam hopping is a time-slicing-based technology. In a time slot, only part of the beams are activated driven by business needs, and the non-uniform and time-varying needs are matched with the satellite system resources. LEO satellites realize efficient use of on-board resources through beam hopping technology, and provide services for marine user terminal systems in the coverage area, and return the information to be transmitted to the ground station or data center. Through the beam hopping technology, the flexibility of resource configuration in the spatial dimension of the system is increased, and by adjusting the number of time slot allocation in different wave positions, communication resources matching the business needs are provided, and more communication areas can be covered with fewer hopping beams; and in terms of interference avoidance, reasonable activation of wave positions can effectively avoid interference, meet user needs, and improve network efficiency.

[0030] However, with the continuous development of satellite communication systems, the hopping beams can meet the user's uneven temporal and spatial distribution of service requirements to some extent through time slicing and on-demand activation hopping, and improve the utilization rate of on-board resources. However, the current application of the hopping beam technology in satellite communication still faces some challenges, mainly in the following aspects: (1) Although the hopping beam technology can meet the service requirements of each beam by changing the residence time, the shape and relative position of the beam are usually fixed, and most existing researches consider the resource allocation of a single LEO satellite, that is, the resource allocation scheme of a single satellite serving a specific area on the ground. In practice, due to the low orbit height, the coverage area of LEO satellites is limited, and there are usually multiple satellites in each orbit, and multiple orbits around the earth. Therefore, the coverage of LEO satellites is cooperatively covered by multiple satellites. For these cooperatively covered satellites, the sidelobe electromagnetic wave signals of their beams may interfere with the main lobe of the service beam of other satellites. Considering only the resource allocation of a single satellite will cause beam interference between satellites, resulting in a decline in communication quality and a decrease in system throughput; (2) To realize user-centered on-demand access, LEO satellite networks need to adjust the hopping beam pattern according to the dynamic changes of user distribution, service performance requirements, and network channel environment, so new hopping beam pattern design and optimization methods are needed. (3) The existing resource allocation algorithm usually has two steps. First, allocate the corresponding time slots for each beam, and then design the hopping pattern. However, this two-step design increases the complexity of resource allocation and is not suitable for scenarios with rapidly dynamic changes in service demand. In the face of real-time changes in service demand, this method may not be able to adapt in time, resulting in a decline in resource utilization efficiency. Therefore, more flexible hopping beam and resource allocation algorithms need to be researched to meet the changing service requirements.

[0031] The application provides a low-orbit satellite hopping beam-based ocean communication on-demand service coverage method, as shown in the accompanying drawings. Figure 1 The low-orbit satellite hopping beam-based ocean communication on-demand service coverage method includes the following steps 101 to 104.

[0032] Step 101: Determine the positions of each user terminal on the ocean through a low-orbit satellite.

[0033] Step 102: Determine the number of beams and the positions of each candidate beam center point according to the positions of each user terminal.

[0034] Step 103: Based on the positions of each user terminal and the positions of each candidate beam center point, use the low-orbit satellite as a relay for communication between each user terminal and the ground station, and construct an optimization model of multi-low-orbit satellite joint hopping beam and user terminal transmit power.

[0035] Step 104: In each time slot, the optimization variables of the optimization model in a single time slot are solved by using a hierarchical heterogeneous multi-agent deep reinforcement learning method, and the optimization variables include the uplink transmit power of the corresponding user terminal output by the user terminal agent, the jump beam pattern of the corresponding low earth orbit satellite output by the low earth orbit satellite agent, and the subchannel allocation of the user terminal.

[0036] The present application designs an on-demand communication coverage method based on low earth orbit satellite jump beam technology according to the characteristics and needs of communication in the marine scene, that is, the communication method of the present application meets the user service needs while improving the system efficiency of the LEO satellite network. Specifically, the number of wave positions and the positions of each candidate wave position center point are determined according to the positions of each user terminal, which can realize the allocation of multiple low earth orbit satellite coverage tasks, while avoiding excessive coverage overlap, improving the on-demand coverage capability and resource utilization efficiency; based on the positions of each user terminal and the positions of each candidate wave position center point, the low earth orbit satellite is used as a relay for communication between each user terminal and the ground station, an optimization model of joint jump beam of multiple low earth orbit satellites and transmit power of user terminals is constructed, in each time slot, the optimization variables of the optimization model in a single time slot are solved by using a hierarchical heterogeneous multi-agent deep reinforcement learning method, and the optimization variables include the uplink transmit power of the corresponding user terminal output by the user terminal agent, the jump beam pattern of the corresponding low earth orbit satellite output by the low earth orbit satellite agent, and the subchannel allocation of the user terminal, which realizes the optimization of the beam pattern scheme and the uplink transmit power of the low earth orbit satellite and the user terminal, further suppresses beam interference and improves the network resource utilization efficiency.

[0037] With the development of future 6G space-ground-sea integrated communication systems, the number of low earth orbit Internet satellites increases, and the beam coverage area shrinks. The position information of the satellite communication terminal needs to be determined during the registration, access and switching process. Due to the particularity of the marine user, especially the Internet of Things device, it is impossible to calculate and solve at the user end, so the user terminal position determination can only be performed at the LEO satellite side.

[0038] Among them, step 101 specifically includes: determining the positions of each user terminal on the ocean by using a Time Difference Of Arrival (TDOA) positioning method based on a low earth orbit satellite. At the same time, the marine environment information of each user terminal can also be obtained through the low earth orbit satellite.

[0039] The TDOA positioning algorithm is based on the time of arrival (TOA) and determines the position of the receiving end through the intersection of three double hyperboloids on the same side. The measurement value is derived from the distance difference between the two transmitting ends and the receiving end. By substituting the satellite positions as one focus of the hyperboloid and the position of another satellite as the other focus into the distance difference expression of the TDOA measurement value, the analytical expression of the receiving end position can be obtained. Due to the short communication time between satellites, real-time positioning can be achieved. In order to comprehensively grasp the ocean environment information, LEO satellites also need to be equipped with various sensors for monitoring ocean temperature, salinity, ocean current, sea wave height and other parameters. Through real-time collection and analysis of these data, accurate environmental information support can be provided for marine end users.

[0040] Wherein, through the low earth orbit satellite, the position of each user terminal on the sea is determined by using the time difference of arrival positioning method, specifically including: through at least 4 low earth orbit satellites, the time difference of arrival of signals between at least 3 low earth orbit satellites and the same user terminal is obtained; based on at least 3 said time difference of arrival, the coordinates of the user terminal, i.e. the position of the user terminal, are solved by using the weighted least squares (WLS) algorithm.

[0041] The TDOA positioning method determines the position of each user terminal on the sea, more specifically, including: in a spatial rectangular coordinate system, taking the center of the earth as the origin, setting the coordinates of the user terminal to be positioned P as (x, y, z), and setting the coordinates of the low earth orbit satellite as (x i ,y i ,z i ), the receiving end obtains the satellite time difference by transmitting signals to the satellite for observation positioning. The TDOA of the signal can be obtained by the signal difference received between the satellites, the user terminal on the sea transmits a radio signal with a specific code and time mark, which contains a time stamp marking the precise time when the signal leaves the ground station, and the satellite receiver records the exact time when the signal is received. In order to accurately position, multiple satellites are needed for TDOA measurement.

[0042] Each satellite will provide an independent time difference data, which is combined through complex algorithm processing, and the equation of TDOA is: Wherein, R i is the distance between the target point P and the i-th satellite, and the above formula is squared. The above formula is rewritten as: R i,j =cΔt i,j =R i -R j 、 From the signals transmitted by the satellites, the corresponding TDOA values can be obtained. c is the signal propagation speed, taking the distance between one low-orbit satellite and the terminal as a reference, and taking this satellite as a reference station, R i,j , Δt i,j are the distance and transmission time difference values of the ith satellite and the jth satellite, respectively.

[0043] Substituting i = 1 into the above formula, we have: R1 2 = K1-2x1x-2y1y-2z1z+x 2 +y 2 +z 2 , x i,1 = x i -x1; y i,1 = y i -y1; z i,1 = z i -z1、

[0044] When solving the above formula, at least 4 equations are needed, and when the number of satellites n ≥ 4, at least 3 TDOA values can be obtained. The TDOA algorithm usually uses the least square algorithm to obtain the smallest sum of squares of errors to obtain the optimal solution of the system. The weighted least square algorithm adds a weighting matrix on the basis of the least square algorithm, solving the heteroscedasticity problem existing in the original algorithm. The TDOA measurement value contains errors, and the WLS algorithm performs weighted processing on the TDOA measurement value of each satellite, thereby obtaining more accurate positioning accuracy.

[0045] According to the TDOA measurement equation, the target equation Y = AX is constructed, it is assumed that δ is the residual, δ = Y - AX, the weight matrix is H, and the WLS algorithm derives the minimum weighted value of the residual, that is, the estimation value of X. The specific expression is: Let the partial derivative of the function f(x) be 0, and the coordinates of the user terminal can be obtained: -2A T H(Y-AX) = 0, X = (A T HA) -1 A T HY. Y is an observation value matrix, which contains the time difference information received by multiple receivers. A is a design matrix, which includes the positions of various receivers, the position of the user terminal, and possibly other related parameters.

[0046] After the position of the user terminals on the sea surface is determined, the satellite needs to optimize and divide the wave positions, including determining the number of wave positions and the position of the center point. The goal is to cover all user terminals with the least number of wave positions and ensure a certain interval between wave positions, while making the user terminals as close to the center point as possible to reduce interference between adjacent wave positions and reduce user transmission power. In this application, the system delay is positively correlated with the number of wave positions, and the more wave positions served by the same beam, the greater the delay. Therefore, the goal is to cover all user terminals with the least number of wave positions, and this problem is to solve the P-center problem in plane geometry. In this way, the wave position distribution can be optimized, the system delay and interference can be reduced, and the communication efficiency can be improved.

[0047] First, the number of wave positions is determined. A number of circular wave positions with the same radius (preset value) are used to cover all user terminals. As shown in Figure 2 The flowchart for determining the number of wave positions is shown. Given the positions of the user terminals on the sea surface, the coverage radius of the satellite beam is set to ensure that the users within this radius can obtain satisfactory service quality, and the optimization goal is established: to minimize the number of wave positions required while ensuring that each user is covered by at least one beam; constraint condition: the coverage radius of each beam is fixed and all users must be covered by at least one beam.

[0048] Step 102 specifically includes the following steps:

[0049] Step 1021: Place the positions of the user terminals in the position list, and initialize the number of wave positions to 1.

[0050] Step 1022: Take the center user position of all user terminals in the current position list as the initial wave position center point (x1, y1), and place the initial wave position center point in the initial wave position center point list. The initial wave position center point placed in the initial wave position center point list is taken as the current wave position center point, and the geometric center point (x1, y1) is taken as the center with a preset value as the radius to construct the first wave position (circular region).

[0051] Step 1023: Calculate the Euclidean distance of all user terminals to the initial wave position center point and determine which user terminals are covered, i.e. the user terminals with a distance less than or equal to the set radius. Remove the user terminals within the coverage range of the circle with the current wave position center point as the center and a preset value as the radius from the current position list to obtain the updated position list.

[0052] Step 1024: Determine whether there are still user terminals in the updated position list, i.e. whether there are still user terminals that have not been covered in the updated user position list.

[0053] If a user terminal exists, increase the number of waveforms and execute step 1025.

[0054] Step 1025: Update the location of the user terminal that is more than twice the preset value away from the current wave position center point to the current wave position center point, and add the current wave position center point to the initial wave position center point list.

[0055] Step 1026: Use the updated location list as the current location list, and return to the step of removing user terminals within the circle covered by a preset value centered on the current wave position center point from the current location list to obtain the updated location list.

[0056] If no user terminal exists, proceed to step 1027.

[0057] Step 1027: Output the initial list of wave position center points. One wave position center point corresponds to one wave position. Number all wave positions.

[0058] Step 1028: To obtain the optimal position of each wave position center point and ensure that the wave position can optimally cover all user points, the positions of the wave position center points need to be further optimized. The Weiszfeld algorithm is used to update the positions of each wave position center point in the initial wave position center point list. After the update, each wave position center point is used as a candidate wave position center point. The Weiszfeld algorithm is an iterative algorithm used to solve for minimizing the maximum distance of a single facility. It is particularly suitable for processing the coordinates of points in continuous space, such as finding a point on a plane such that the sum of the distances from this point to a given set of points is minimized.

[0059] The specific implementation process of the Weiszfeld algorithm includes: Step 81: Select an initial point as the starting position of the algorithm. In this application, the starting point is the wave position center point of the initial wave position center point list obtained in step 1027.

[0060] Step 82: Iteratively update the position of the wave position center point using the following formula: Where d(x) i ,x (k) ) represents the user terminal within the wave position and the current iteration point x. (k) The weighted Euclidean distance between them, w i As the weight, x (k+1) x is the center point of the updated wave position. (k) To update the center point of the preceding wave position, x i Let n be the location of the user terminal and n be the number of user terminals.

[0061] Step 83: After each iteration, check whether the positional difference between the center points of the new and old wave positions is less than a predefined threshold ε, i.e.: ||x (k+1) -x(k) || < ε. If this condition is met, the algorithm converges and iteration can be stopped. Otherwise, use x (k+1) as the input point for the next iteration, and repeat step 82.

[0062] After traversing all the wave positions in the initial wave position center point list, when the algorithm converges, the final iteration point x (k +1) is output as the final wave position center point.

[0063] Considering that the user terminal cannot directly communicate with the ground station, for this case, by introducing LEO satellites as relays, the process of establishing uplink connection between the sea terminal and LEO satellites and satellite backhaul to the ground station is established, as shown in Figure 3 . Given the user terminal location and network information such as channel state, a multi-LEO satellite joint hop-beam and terminal transmission power optimization model is established. The optimization goal is to maximize network throughput, while considering multiple factors, including the number of users scheduled per time slot, power constraints, wave position center point and satellite association, wave position activation state, number of sub-channels occupied by users, and binary variable constraints. These constraints ensure efficient allocation of resources and help solve resource allocation problems in satellite communication networks. By optimizing beam hopping patterns and frequency reuse, resource utilization and interference avoidance are improved.

[0064] The signal model is described. In the uplink resource scheduling scenario of LEO satellite communication system, it is assumed that there are M LEO satellites covering the entire area, The satellite is equipped with a planar phased array antenna, which divides the area covered by the above LEO satellite into N wave positions, It is particularly pointed out that a satellite can cover multiple wave positions, which is a major feature of Non-Geostationary Satellite Orbit (NGSO) satellite communication systems. It is assumed that there are K Marine User Equipment (MUE) in total, and N << K. The satellite position is represented as d m = (x m , y m , h m ), the MUE position can be represented as d k = (x k , y k , h k ), and the coordinates of the wave position center point are d n = (x n , y n , h n). To fully exploit the spectrum resources and improve the spectrum utilization, it is assumed that all beams occupy the entire satellite bandwidth B, i.e., the multi-beam satellite system adopts full frequency reuse (FFR), there is co-channel interference between beams, and users within each beam use orthogonal multiple access (OMA), which means that users within a beam use different sub-channels and there is no interference between users within a beam. It is idealized that the network topology remains unchanged within the performance observation period.

[0065] In such a scenario, the communication link works on the ka band (26.5-40 GHz), and the satellite m and user terminal k channel follows the following standard model: δ mk = ω mk + ρ mk 10 log 10 d mk + ψ, where δ mk represents the path loss of all users related to the satellite, ω mk represents the intercept parameter, i.e., the path loss (dB) at a distance of 1 meter, ρ mk is the path loss exponent, d mk represents the distance between satellite m and user k, which can be expressed as: ψ represents the fitting deviation (dB), which is a Gaussian random variable, and for a distance of 1 meter, the mean is zero and the variance is Then, the channel gain experienced by the signal received by the user terminal k on the beam n is: where G(θ) represents the gain of the transmit antenna of the user terminal m pointing to the beam n, β mk is the Rician fading channel gain coefficient, and represents the small-scale fading between satellite m and user terminal k. The receive antenna gain can be expressed as: where θ is the angle between the line connecting user k and satellite m and the center line of the beam pointing to the beam, i.e., the off-axis angle. Let d m,k , d m,n , and d k,n represent the distance from satellite m to user terminal k, the distance from the satellite to the beam center position, and the distance from user terminal k to the beam center position of beam n, respectively, then the off-axis angle can be given by the cosine law: where G0 is the peak gain of the beam, defined as η is the antenna efficiency, generally equal to 0.65, N' represents the phased array antenna, θ 3dB is the 3dB gain angle of the antenna, J1(·) and J3(·) represent the first and third Bessel functions, respectively, and u(θ) = 2.07123sinθ / sinθ3dB .

[0066] Let be the set of time slots, in a time slot t, let denote the beam position matrix associated with the satellite, denote that beam position n can be covered by satellite m, be a binary variable, denote that beam n is activated, denote that it is not activated, the activated binary variable is denoted as Let denote the user terminal association matrix, denote that user terminal k resides in beam position n, since there can be overlapping between beam positions, user terminal k can be covered by multiple beam positions. Considering that there are multiple users inside each beam, and there can be differentiated service requirements among users; in order to facilitate reducing interference level and meeting differentiated performance requirements, it is assumed that beams are fully spectrum multiplexed between beams, let the total bandwidth be B, and the bandwidth of each subchannel be ΔB; it satisfies B = S x ΔB, that is, the number of subchannels is The noise power spectral density is N0. Considering that the distribution of ocean communication users is sparse, and the differentiated service requirements of users, it is assumed that different numbers of subchannels can be allocated to users inside the beam for uplink transmission; define to indicate that in time slot t, user terminal k in beam position n covered by satellite m is allocated subchannel s for uplink transmission, the corresponding set can be denoted as Correspondingly, in time slot t, the uplink transmission power of user terminal k in subchannel s under the beam n covered by satellite m is For all users, define the transmission power set in time slot t as Correspondingly, the uplink received signal-to-noise ratio of user terminal k on subchannel s under the beam n covered by satellite m is which can be denoted as:

[0067]

[0068] where, denotes the channel gain of user terminal k in subchannel s under the beam n covered by satellite m in time slot t.

[0069] Therefore, the corresponding transmission rate can be written as:

[0070] When user terminal k needs to allocate multiple subchannels due to service requirements and network benefits, the uplink transmission rate can be denoted as:

[0071] The optimization objective of the optimization model proposed in the present application is to jointly optimize satellite association wave positions, wave position activation and terminal transmission power to maximize the overall system throughput. Therefore, in T time slots, the average system throughput maximization corresponds to the optimization problem, i.e., the optimization model can be represented as:

[0072]

[0073] wherein, represents the average value of T time slots, K represents the number of user terminals, w k represents the weight of user terminal k, which can be dynamically adjusted according to the performance requirements of marine users, represents the uplink transmission rate of user terminal k in time slot t, C1 represents the minimum user number constraint of each time slot scheduling, C2 represents the user terminal uplink transmission power constraint, C3 represents the maximum number of satellites associated with wave position n in time slot t, C4 represents the constraint that wave position can only be activated or not activated by one satellite, C5 represents the constraint that user terminal belongs to at most one wave position, M represents the number of low-orbit satellites, and N represents the number of wave positions, represents that wave position n is covered by low-orbit satellite m in time slot t, represents that beam n is activated in time slot t, represents that beam n is not activated in time slot t, represents that user terminal k resides in wave position n in time slot t, K min represents the minimum user number; represents that user terminal k in wave position n covered by low-orbit satellite m in time slot t is allocated subchannel s for uplink transmission; represents the uplink transmission power of user terminal k on subchannel s in wave position n covered by low-orbit satellite m in time slot t; P max,k represents the upper limit of the uplink transmission power of user terminal k; is a set of time slots, is a set of wave positions, is a set of user terminals.

[0074] The above optimization problem is a mixed integer nonlinear programming problem, involving continuous variables and discrete variables These variables are coupled with each other, resulting in a huge combination of solution space. In addition, the non-convexity of the objective function and the constraints increases the complexity of the algorithm. In order to solve this problem, the problem is redefined as a cooperative distributed partially observable Markov decision process (Dec-POMDP), and a multi-agent deep reinforcement learning method is used to solve it.

[0075] The interference management in uplink transmission depends not only on user terminals but also on LEO satellites. To solve the optimization problem in a single time slot, a hierarchical heterogeneous multi-agent deep reinforcement learning (HH-MADRL) method is adopted. Considering the real-time and distributed requirements, the LEO satellites and sea surface terminals are modeled as independent agents, and the problem is re-modeled as a Dec-POMDP. Specifically, the task of LEO satellites is to optimize their associated wave positions and activation problems. It should be noted that the same satellite can cover multiple wave positions, but the same wave position can only be activated by one satellite. The terminal agent is responsible for determining the appropriate transmission power level. The deep reinforcement learning (DRL) method uses a deep neural network (DNN) to optimize the policy, and through continuous interaction with the environment, the agent can learn and optimize its policy to maximize the reward.

[0076] In the hierarchical multi-agent deep reinforcement learning process, each user terminal agent and each low-orbit satellite agent is modeled by a distributed partially observable Markov decision process, and a shared hybrid network and a local network are used for collaborative training and decision-making, respectively outputting the uplink transmission power of the user terminal and the hop beam pattern of the low-orbit satellite; the local network is a local Q network.

[0077] To learn the policy of the agent, the application adopts a novel heterogeneous coordination algorithm, and constructs a shared hybrid network to train a centralized Q function, which can be decomposed into a single Q function of the LEO satellite and the terminal agent. After the local Q function is trained, the LEO satellite and the terminal can make their own decisions.

[0078] In the hierarchical multi-agent deep reinforcement learning process, each user terminal agent and each low-orbit satellite agent is modeled by a distributed partially observable Markov decision process, and a shared hybrid network and a local network are used for collaborative training and decision-making, respectively outputting the uplink transmission power of the user terminal and the hop beam pattern of the low-orbit satellite; the local network is a local Q network.

[0079] Dec-POMDP is an extension of Markov Decision Process, designed to handle the decision problem of distributed multi-agent, where each agent can only observe part of the state of the environment. Dec-POMDP allows these agents to make decisions under limited, local information, and through cooperation to achieve a common goal or optimize the collective benefit. At each time slot, the environment information is described by state s t ∈S, each agent receives a separate observation value which provides partial information about the state s t . Subsequently, each agent selects an action a under the guidance of the strategy The joint action taken by all agents makes the environment transition to the next state s t+1 , and at the same time, all agents obtain a shared reward. Here, the reward refers to the cumulative reward, that is: where γ ∈ [0, 1) is the discount factor. It should be noted that in this application, LEO satellites and sea surface user terminals are considered as two types of heterogeneous agents, each having different observation and action spaces. All agents make decisions at the end of each time slot to ensure that user association or resource allocation related operations can be performed in the subsequent time slot. According to the optimization problem described in step three, the basic components of Dec-POMDP are defined as follows.

[0080] Observation: Due to partial observability, each agent can only perceive environmental information within its designated perception range. The observation includes the observation value of the user terminal agent and the observation value of the low earth orbit satellite agent.

[0081] The observation value of the user terminal agent includes the current coordinates of the user terminal and the transmission power in the last time slot, that is wherein is the observation value of the user terminal k in time slot t, is the coordinates of the user terminal k in time slot t, is the transmission power of the user terminal k in time slot t-1.

[0082] The observation value of the low earth orbit satellite agent includes the current (time slot t) coordinates of the low earth orbit satellite, the positions of the user terminals, the average throughput of multiple time slots before the current time slot, the transmission power of all associated user terminals in the last time slot the set of wave positions associated with the last time slot the set of wave positions activated in the current time slot and the subchannel allocation in the wave positions that have been activated in the current time slot wherein is the observation value of the low earth orbit satellite m in time slot t, is the coordinate of LEO satellite m in time slot t, is the estimated coordinate of user terminal k in time slot t, is the average throughput of multiple time slots before time slot t. and respectively denote the set of beam positions activated by LEO satellite m in time slot (t-1), and denotes that all beam positions are not activated.

[0083] The state includes the position of each LEO satellite, the position of each user terminal, the set of beam positions associated with the last time slot, the set of activated beam positions in the current time slot, the transmit power of each user terminal, the average throughput of multiple time slots before the current time slot, and the subchannel allocation in the activated beam positions in the current time slot.

[0084] The state is:

[0085] The actions include the actions of user terminal agents and the actions of LEO satellite agents.

[0086] Since the trajectory of LEO satellites is fixed and the running speed is basically stable, in each time slot, LEO satellites must adjust the antenna pointing to cover the beam positions to provide fair service. The actions of LEO satellite agents include the set of beam positions associated with the current time slot, the set of activated beam positions in the current time slot, and the subchannel allocation in the activated beam positions in the current time slot; the actions of user terminal agents include the transmit power in the current time slot.

[0087] The action of a LEO satellite agent is where denotes the selection of the associated beam position indicator vector in time slot t, is a binary vector whose elements take values of 0 or 1, indicating whether LEO satellites activate their associated beam positions in time slot t, denotes the subchannel allocation in the activated beam position numbered n by LEO satellites.

[0088] The action of a user terminal agent is

[0089] The action space of satellite agents is discrete, while the action space of user terminal agents is continuous. Since searching for the optimal action in continuous space is still challenging for MADRL configuration, it is considered to discretize all continuous actions. Specifically, the transmit power of the terminal is processed as follows: the range of transmit power [0, P max ] is quantized into b power levels, denoted as: Such discrete power levels constitute the action space of user terminal agents.

[0090] The goal of the reward optimization problem is to maximize the system average throughput over the entire service period. To achieve this, a greedy policy is adopted, and the reward is defined as a function of the system average throughput per time slot. The goal of the MADRL is to find an optimal joint policy such that the expected return can be maximized, where R0is the reward at the initial time instant. Typically, the optimal Q-function is used to describe the expected return from state s t starting, taking action a t and then following the optimal policy * : where R t is the reward at time slot t. Given Q * (s t , a t ), the optimal policy * can be obtained by choosing the greedy policy:

[0091] Learning the optimal policy in a multi-agent setting is much more complex than in the single-agent case. One approach is to train a centralized Q-function that depends on the global state and joint actions by extending DQN to the multi-agent setting. However, training this function becomes challenging as the number of agents increases, and it lacks the ability to generate local policies for distributed execution. Another approach is to train each agent independently, considering other agents as part of the environment. This approach often fails to converge due to the non-stationarity issue, as the environment of each agent evolves with the policy updates of other agents. This problem is particularly challenging in the marine communication scenario of the present application due to the different capabilities, roles, and contributions of LEO satellites and sea surface terminal agents. The present application employs a heterogeneous MADRL framework to address the above challenges.

[0092] The present application employs a Heterogeneous Coordination (HC-QMIX) algorithm, which is extended based on the QMIX framework, to learn local policies of LEO satellites and terminal agents coordinately. The basic concept of HC-QMIX includes estimating the joint Q-value as a complex nonlinear combination of individual Q-values of heterogeneous agents conditioned only on local observations, a process referred to as heterogeneous value decomposition. By guaranteeing the monotonicity of this combination function, each agent selects a greedy behavior according to the local Q-value to achieve global optimality. To implement this approach, HC-QMIX includes a set of local Q-networks for learning individual Q-values of LEO satellites and sea surface terminal agents, and a shared mixing network to combine these Q-values into a joint Q-value, as shown in Figure 4 .

[0093] Each of the local Q networks includes a first fully connected layer, a gated loop unit, and a second fully connected layer, such as... Figure 5 As shown in (a), a single local network (local Q-network) represents a single Q-function for each agent. It takes the agent's observations as input and generates Q-values ​​for all actions in the action space. Based on these Q-values, an arbitrary policy is used to determine the optimal behavior. The first fully connected layer inputs the observations o. t Processed into embedded features z t The gate loop unit maintains a hidden state h. t It is composed of the current embedded feature z t The hidden state h of the previous time slot t-1 Decision. The second fully connected layer uses the current hidden state h. t As input, generate the Q-values ​​for all actions, i.e. Finally, these Q values ​​are used to determine the optimal action a. t and its Q t .

[0094] Shared hybrid network representation joint Q tot (s t ,a t ) value and a single Q t The relationship between values. Specifically, it takes the outputs of all local Q-networks as input to generate a joint Q-network. tot (s t ,a t The shared hybrid network includes an input layer and an output layer, such as... Figure 5 As shown in (b), the input layer is represented as follows: The output layer is represented as:

[0095] Among them, Q t W is the set of Q-values ​​output by each local Q-network. in and b in These are the weight matrix and bias vector of the input layer, respectively. For the embedded features output by the input layer, σ in and σ out Both are activation functions, W out and b out These are the weight matrix and bias vector of the output layer, respectively, Q. tot (s t ,a t ) represents the joint Q-value output by the output layer, s t For the state within time slot t during deep reinforcement learning, a t For actions within time slot t during deep reinforcement learning.

[0096] The present application adopts a framework of centralized training and decentralized execution, which ensures the utilization of global information in the training process while allowing independent decision-making in the execution process. Specifically, in the centralized training process, the agents interact with the environment using a greedy policy derived from the local Q-network, thereby collecting experiences. These interactions are organized into uniform-length episodes. At the beginning of each episode, the environment state is reset, and the hidden state of each agent's local Q-network is initialized to an all-zero vector. The transition tuples {(s t ,{o t},{a t},s t+1 ,{o t+1},r t ,{h t}) t=1:T are stored in a replay buffer. Periodically, a batch of episodes is drawn from the buffer, and the network parameters are trained using the stochastic gradient ascent method. Specifically, the parameters θ = {θ LEO ,θ MUE ,θ hyper} are updated by minimizing the loss function: where θ - denotes the target network parameters that are updated periodically.

[0097] After reaching convergence, the trained local Q-networks are deployed to the corresponding user terminals or low earth orbit satellites, and the user terminal agents and the low earth orbit satellite agents both select the greedy action to output the optimal policy in a distributed manner according to the formula .

[0098] where, is the optimal policy output by agent i at time slot t, if ξ = LEO, agent i is a low earth orbit satellite agent, if ξ = MUE, agent i is a user terminal agent, is the set of user terminals, is the set of low earth orbit satellites, Q ξ,i (o ξ,i ,a ξ,i ) is the sum of the Q-values of the action, o ξ,i is the observation value of agent i, and a ξ,i is the action of agent i.

[0099] The communication process between the LEO satellite network system and the ocean user terminals in each time slot follows the beam pattern scheme, user scheduling strategy, and uplink transmission power configuration determined in step 104, aiming to optimize communication efficiency and resource utilization. Specifically, first, the LEO satellite adjusts its antenna array according to the optimal beam pattern calculated in step 104, forming a highly directional beam to ensure that the signal can cover the corresponding wave position. At the same time, based on the user scheduling strategy, the system dynamically allocates communication resources to determine which ocean terminals can transmit data in the current time slot, thereby avoiding unnecessary competition and interference. Ocean terminal users adjust the power level of their transmitters according to the uplink transmission power settings indicated in step 104 to ensure that the signal is neither too strong to waste nor too weak to be overwhelmed by noise when it reaches the LEO satellite. This power control helps to prolong the battery life of the ocean terminal and maintain good communication quality. The communication process of the current time slot ends, and the system automatically transitions to the next time interval. In this process, the LEO satellite network system will recalculate and update the beam pattern, user scheduling, and uplink power parameters based on the communication effect of the previous time slot and environmental changes. This closed-loop dynamic adjustment mechanism ensures that the system can continuously adapt to changing communication demands and external conditions. The entire process constitutes a cycle from the initial parameter setting to the final strategy optimization, and the communication between the LEO satellite and the ocean terminal users continues until the end of the entire operation period or the satisfaction of specific stopping conditions.

[0100] The application has the following technical effects: (1) The approximate position of the ocean terminal user and the ocean environment information are estimated based on the TDOA positioning method on the LEO satellite; the position information and network information of the ocean user terminal can be estimated, which facilitates the design of beam hopping and on-demand coverage. (2) A method for determining the number of wave positions is proposed according to the position information of the sea surface user, and the position of the candidate wave position center point is further determined; the distribution of the multi-LEO satellite coverage task can be realized, while avoiding excessive coverage overlap, improving the on-demand coverage capability and resource utilization efficiency. (3) Based on the obtained position and network information, a joint optimization model of multi-LEO satellite joint beam hopping and terminal transmission power is established; by optimizing the beam hopping and uplink transmission power, the interference can be further suppressed and the network resource utilization efficiency can be improved. (4) Based on the hierarchical heterogeneous multi-agent deep reinforcement learning method, the optimization problem in a single time slot is solved, and the ocean user terminal and the LEO satellite make decisions on their uplink transmission power and beam pattern scheme according to the algorithm output; centralized training and distributed execution online deployment can be realized; the LEO satellite and the ocean terminal user complete the communication process of the corresponding time slot according to the obtained beam pattern scheme, user multiplicity strategy and uplink transmission power. The LEO satellite and the ocean terminal user complete the communication process of the corresponding time slot according to the beam pattern scheme, user multiplicity strategy and uplink transmission power obtained in the last time slot, and the network system enters the next time slot and executes the above steps in turn, realizing the resource optimization of the whole communication process.

[0101] The application can meet the communication coverage demand of unevenly distributed sea surface users and differentiated service performance, and realize joint optimization of beam hopping and terminal transmission power.

[0102] Based on the same inventive concept, the embodiments of the application also provide a low-orbit satellite beam hopping based ocean communication on-demand service coverage device for implementing the low-orbit satellite beam hopping based ocean communication on-demand service coverage method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more low-orbit satellite beam hopping based ocean communication on-demand service coverage device embodiments provided below can be referred to the limitations of the low-orbit satellite beam hopping based ocean communication on-demand service coverage method in the foregoing, which will not be described here again.

[0103] In one exemplary embodiment, as shown in Figure 6 A low-orbit satellite beam hopping based ocean communication on-demand service coverage device is provided, which includes:

[0104] A user terminal position determination module is configured to determine the positions of the user terminals on the ocean by the low-orbit satellite.

[0105] A wave position determination module is configured to determine the number of wave positions and the positions of candidate wave position center points according to the positions of the user terminals.

[0106] An optimization model construction module is configured to construct an optimization model of multi-low-orbit satellite joint hop-beam and user terminal transmission power based on the positions of the user terminals and the positions of the candidate wave position center points, with the low-orbit satellites serving as relays for communication between the user terminals and the ground station.

[0107] A solving module is configured to solve the optimization variables of the optimization model in each time slot by using a hierarchical heterogeneous multi-agent deep reinforcement learning method, the optimization variables including the uplink transmission power of the corresponding user terminal output by the user terminal agent, the hop-beam pattern of the corresponding low-orbit satellite output by the low-orbit satellite agent, and the sub-channel allocation of the user terminal.

[0108] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered as falling within the scope of the present disclosure.

[0109] The principles and implementation manners of the present application are described by using specific examples herein, and the above embodiments are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, the specific implementation manners and application range can be changed according to the idea of the present application. In conclusion, the content of the present specification should not be understood as limiting the present application.

Claims

1. A method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams, characterized in that, The method for providing on-demand marine communication services based on low-Earth orbit satellite hopping beams includes: The location of each user terminal on the ocean is determined by low-orbit satellites; The number of wave positions and the location of the center point of each candidate wave position are determined based on the location of each user terminal. Based on the locations of each user terminal and the center points of each candidate beam position, low-Earth orbit (LEO) satellites are used as relays for communication between each user terminal and the ground station. An optimization model is constructed to optimize the combined beam hopping of multiple LEO satellites and the transmit power of the user terminals. The optimization objective of this model is to maximize... The weighted average rate of all users within a time slot, expressed in the optimization model, is as follows: ; in, express T The average value of each time slot, K Indicates the number of user terminals. Indicates user terminal k The weight, Indicates time slot Internal user terminal k Allocate uplink transmission rates for multiple sub-channels; C 1 indicates the minimum number of users required for scheduling per time slot. C 2 indicates the uplink transmit power constraint of the user terminal. C 3 indicates the wave position in the time slot. The system can be constrained to associate with at most one satellite. C 4 indicates that the wave position can only be activated or deactivated by one satellite. C 5 indicates that the user terminal belongs to at most one wave position constraint. Indicates the number of low-orbit satellites. The number of wave positions, Indicates time slot Inner wave position By low-orbit satellites cover, Indicates time slot Beam Activated Indicates time slot Beam Not activated Indicates time slot Internal user terminal Residing at wave position middle, Indicates the minimum number of users; Indicates in time slot Inside, by low-orbit satellites Coverage wave position User terminals within Sub-channels assigned Perform uplink transmission; Indicates in time slot Inside, by low-orbit satellites Coverage wave position User terminals within In sub-channel Uplink transmit power; Indicates user terminal The upper limit of uplink transmit power; For time slot set, For wave position set, A collection of user terminals; Within each time slot, a hierarchical heterogeneous multi-agent deep reinforcement learning method is used to solve the optimization variables of the optimization model within a single time slot. These optimization variables include the uplink transmit power of the corresponding user terminal output by the user terminal agent, the hopping beam pattern of the corresponding low-Earth orbit satellite output by the low-Earth orbit satellite agent, and the sub-channel allocation of the user terminal. During the heterogeneous multi-agent deep reinforcement learning process, a distributed partially observable Markov decision process is used to model each user terminal agent and each low-Earth orbit satellite agent. Training and decision-making are conducted collaboratively through a shared hybrid network and local networks. The local network is a local... network.

2. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 1, characterized in that, The location of user terminals on the ocean is determined using low-Earth orbit satellites, specifically including: The location of each user terminal on the ocean is determined by using low-orbit satellites and a positioning method based on time difference of arrival.

3. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 2, characterized in that, The location of user terminals at sea is determined using low-Earth orbit satellites and a time-difference-of-arrival (TDOA) positioning method, specifically including: By using at least four low-Earth orbit satellites, obtain the signal arrival time difference between at least three low-Earth orbit satellites and the same user terminal; The coordinates of the user terminal are solved using a weighted least squares algorithm based on at least three arrival time differences.

4. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 1, characterized in that, The number of wave positions and the location of the center point of each candidate wave position are determined based on the location of each user terminal. Specifically, this includes: Add the location of each user terminal to the location list; Take the center user position of all user terminals in the current position list as the initial wave position center point, put the initial wave position center point into the initial wave position center point list, and take the initial wave position center point put into the initial wave position center point list as the current wave position center point; Remove user terminals from the current location list that are within the radius of a circle centered on the current wave position center point, and obtain the updated location list; Determine if the user terminal still exists in the updated location list; If a user terminal exists, the location of the user terminal that is more than twice the preset value away from the current wave position center point is updated to the current wave position center point, and the current wave position center point is added to the initial wave position center point list. The updated location list is used as the current location list. The steps are: to remove user terminals within the circle covered by a preset value as the center point of the current wave position from the current location list, and to obtain the updated location list. If no user terminal exists, output the initial wave position center point list; The position of each wave position center point in the initial wave position center point list is updated using the Weiszfeld algorithm, and after the update, each wave position center point is used as a candidate wave position center point.

5. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 1, characterized in that, In the heterogeneous multi-agent deep reinforcement learning process, centralized training is performed on each of the user terminal agents and each of the low-Earth orbit satellite agents. Specifically, each of the user terminal agents uses a local... Online learning Each of the aforementioned low-orbit satellite agents employs a local value. Online learning Values, each local Network output All values ​​are input to a shared hybrid network, and the shared hybrid network outputs a joint... The value will be the value of each trained local Network deployment to corresponding user terminals or low-Earth orbit satellites; each local The network's input consists of the observations of the corresponding agent, and its output consists of actions. Value, and then according to The value is selected to determine the corresponding action decision.

6. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 5, characterized in that, In the heterogeneous multi-agent deep reinforcement learning process, a distributed partially observable Markov decision process is adopted for modeling, where the observations include the observations of the user terminal agent and the observations of the low-orbit satellite agent. The observations of the user terminal intelligent agent include the current coordinates of the user terminal and the transmission power of the previous time slot; The observations of the low-Earth orbit satellite agent include the current coordinates of the low-Earth orbit satellite, the position of each user terminal, the average throughput of multiple time slots before the current time slot, the transmit power of all associated user terminals in the previous time slot, the set of wavelets associated with the previous time slot, the set of wavelets activated in the current time slot, and the sub-channel allocation of the wavelets already activated in the current time slot. The status includes the position of each low-Earth orbit satellite, the position of each user terminal, the set of wavelets associated with the previous time slot, the set of wavelets activated in the current time slot, the transmit power of each user terminal, the average throughput of multiple time slots before the current time slot, and the allocation of sub-channels in the wavelets activated in the current time slot. Actions include actions of user terminal intelligent agents and actions of low-Earth orbit satellite intelligent agents; The actions of the user terminal intelligent agent include the transmit power of the current time slot; The actions of the low-Earth orbit satellite agent include the set of associated wavelengths in the current time slot, the set of active wavelengths in the current time slot, and the allocation of sub-channels within the active wavelengths in the current time slot. The reward is a function of the system's average throughput in the current time slot.

7. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 5, characterized in that, Each of the local descriptions The networks all include a first fully connected layer, a gated loop unit, and a second fully connected layer; The shared hybrid network includes an input layer and an output layer, wherein the input layer is represented as follows: The output layer is represented as: ; in, For each local Network output A set of values and These are the weight matrix and bias vector of the input layer, respectively. The embedded features are the output of the input layer. and All are activation functions. and These are the weight matrix and bias vector of the output layer, respectively. The joint output of the output layer value, For time slots in deep reinforcement learning Internal state, For time slots in deep reinforcement learning Internal action.

8. The method for on-demand marine communication service coverage based on low-Earth orbit satellite hopping beams according to claim 5, characterized in that, The trained local After the network is deployed to the corresponding user terminals or low-Earth orbit satellites, the user terminal intelligent agents and the low-Earth orbit satellite intelligent agents all communicate in a distributed manner via formulas. , Choose a greedy action to output the optimal strategy; in, For time slots Inner agent The optimal strategy output, if Then intelligent agent For low-orbit satellite intelligent agents, if Then intelligent agent For user terminal intelligent agents, For user terminal collection, For low-Earth orbit satellites, For action The sum of values For intelligent agents The observed values, For intelligent agents The action.

9. A marine communication on-demand service coverage device based on low-orbit satellite hopping beams, characterized in that, The marine communication on-demand service coverage device based on low-Earth orbit satellite hopping beams includes: The user terminal location determination module is used to determine the location of each user terminal on the ocean via low-orbit satellites. The wave position determination module is used to determine the number of wave positions and the position of the center point of each candidate wave position based on the location of each user terminal. The optimization model construction module is used to construct an optimization model of the joint beam hopping of multiple low-Earth orbit satellites and the transmit power of user terminals, based on the location of each user terminal and the location of the center point of each candidate beam position, using low-Earth orbit satellites as relays for communication between each user terminal and the ground station; the optimization objective of the optimization model is to maximize... The weighted average rate of all users within a time slot, expressed in the optimization model, is as follows: ; in, express T The average value of each time slot, K Indicates the number of user terminals. Indicates user terminal k The weight, Indicates time slot Internal user terminal k Allocate uplink transmission rates for multiple sub-channels; C 1 indicates the minimum number of users required for scheduling per time slot. C 2 indicates the uplink transmit power constraint of the user terminal. C 3 indicates the wave position in the time slot. The system can be constrained to associate with at most one satellite. C 4 indicates that the wave position can only be activated or deactivated by one satellite. C 5 indicates that the user terminal belongs to at most one wave position constraint. Indicates the number of low-orbit satellites. The number of wave positions, Indicates time slot Inner wave position By low-orbit satellites cover, Indicates time slot Beam Activated Indicates time slot Beam Not activated Indicates time slot Internal user terminal Residing at wave position middle, Indicates the minimum number of users; Indicates in time slot Inside, by low-orbit satellites Coverage wave position User terminals within Sub-channels assigned Perform uplink transmission; Indicates in time slot Inside, by low-orbit satellites Coverage wave position User terminals within In sub-channel Uplink transmit power; Indicates user terminal The upper limit of uplink transmit power; For time slot set, For wave position set, A collection of user terminals; The solution module is used to solve the optimization variables of the optimization model within each time slot using a hierarchical heterogeneous multi-agent deep reinforcement learning method. These optimization variables include the uplink transmit power of the corresponding user terminal output by the user terminal agent, the hopping beam pattern of the corresponding low-Earth orbit satellite output by the low-Earth orbit satellite agent, and the sub-channel allocation of the user terminal. During the heterogeneous multi-agent deep reinforcement learning process, a distributed partially observable Markov decision process is used to model each user terminal agent and each low-Earth orbit satellite agent, and training and decision-making are conducted collaboratively through a shared hybrid network and local networks. The local network is a local... network.

Citation Information

Patent Citations

  • Multi-beam satellite communication system resource allocation method based on multi-agent A3C algorithm

    CN116896407A

  • Service-centered seaborne Internet access switching method and system for ultra-dense low-orbit satellites

    CN117856875A