Resource allocation method and device of communication and navigation integrated system, medium and product
By constructing an integrated communication and navigation system for a low-Earth orbit (LEO) satellite constellation, and utilizing Markov decision processes and near-end strategy optimization algorithms, communication and navigation resources are dynamically allocated. This solves the problems of dynamic adaptability and performance coordination in resource allocation within the LEO satellite ICaN system, thereby improving the overall performance and flexibility of the system.
Patent Information
- Application Number
- CN202511234233.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-11
AI Technical Summary
Existing resource allocation methods lack dynamic adaptability in low-Earth orbit satellite communication and navigation integrated scenarios, rely on accurate modeling, and are difficult to simultaneously take into account communication and positioning performance.
To construct an integrated communication and navigation ICaN user system that uses a low-Earth orbit satellite constellation to serve the ground, an optimization objective function is constructed based on channel state and the relative position relationship between satellite users. This objective function is then reconstructed into a Markov Decision Process (MDP) and solved using the Proximity-Based Policy Optimization (PPO) algorithm. Communication and navigation resources are then dynamically allocated.
It achieves dynamic adaptability and performance synergy in resource allocation within the low-Earth orbit satellite ICaN system, improving the overall system performance and configuration flexibility, adapting to dynamic network environments, and balancing communication capacity and navigation accuracy.
Smart Images

Figure CN120935659A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of integrated communication and navigation technology, specifically to a resource allocation method, device, medium, and product for an integrated communication and navigation system. Background Technology
[0002] The sixth-generation (6G) wireless network aims to achieve stronger communication capabilities, higher positioning accuracy, and ubiquitous services. Achieving this goal relies heavily on the deployment of large-scale low-Earth orbit (LEO) satellite constellations. LEO satellites not only possess abundant payload and link resources, but also boast wide coverage and high-performance transmission capabilities. Their dynamic geometry and high ground voltage can enhance the effectiveness of existing satellite navigation systems, thus demonstrating significant potential in providing integrated communication and navigation (ICaN) services. Currently, while Global Navigation Satellite Systems (GNSS) such as GPS, Galileo, and BeiDou can provide high-precision spatiotemporal services globally, they are susceptible to obstruction or interference in complex environments such as indoors, underground, and urban canyons, making stable and continuous high-precision positioning difficult and creating application bottlenecks. Therefore, achieving deep integration of communication and navigation functions within LEO satellite systems has become a key path to overcome existing positioning bottlenecks and build a highly reliable ubiquitous network.
[0003] Integrated Communication and Navigation (ICaN) technology integrates communication and navigation technologies at the signal structure, transceiver equipment, and functional implementation levels, achieving bidirectional empowerment: communication signals can help improve navigation and positioning accuracy and expand coverage; navigation signals can provide high-precision synchronization support for communication systems and simplify signal detection and demodulation. Simultaneously, the coupled use of these two technologies can improve spectrum resource utilization, reduce system deployment and maintenance costs, and optimize radio equipment resource allocation. However, resource allocation is a core challenge in ICaN. Existing resource allocation strategies rely on static allocation based on preset models or average channel conditions, lacking responsiveness to dynamic network changes; rule-driven scheduling relies on fixed weights or priorities, making it difficult to adapt to the high-speed movement, dynamic channels, and complex topology environments of LEO constellations. Although some research has introduced joint allocation methods based on optimization theory, such as convex optimization, limitations still exist.
[0004] Specifically, existing technical solutions have significant shortcomings: Solution 1 optimizes communication and positioning for multi-user multiple-input multiple-output (MIMO) systems based on differential convex optimization, but relies on precise channel modeling and cannot adapt to the rapid channel changes and dynamic interference of low-Earth orbit (LEO) satellites; Solution 2 optimizes resources for ground-based millimeter-wave ad hoc networks by using user clustering and spectrum allocation, but it is geared towards static networks and lacks the ability to adapt to the high-speed movement of LEO satellites and changes in satellite-to-ground links; Solution 3 proposes joint beam and power optimization for LEO satellites, but it has high computational complexity, relies on predefined beam sets and prior information, and only optimizes communication performance without covering navigation requirements, thus failing to meet the cooperative positioning requirements of ICaN.
[0005] In summary, existing technologies are ill-suited to the dynamic environment, multi-task collaboration, and real-time requirements of low-Earth orbit satellite ICaN systems, necessitating a more efficient and flexible resource allocation scheme. Summary of the Invention
[0006] At least one embodiment of this application provides a resource allocation method, apparatus, medium, and product for an integrated communication and navigation system, which addresses the problems of existing resource allocation methods lacking dynamic adaptability in low-Earth orbit satellite integrated communication and navigation scenarios, relying on accurate modeling, and struggling to simultaneously consider communication and positioning performance.
[0007] To solve the above-mentioned technical problems, this application is implemented as follows:
[0008] In a first aspect, embodiments of this application provide a resource allocation method for an integrated communication and navigation system, including:
[0009] To construct a system that integrates communication and navigation for ground-based ICaN users using a low-Earth orbit satellite constellation;
[0010] Based on the channel state of the system and the relative positions of the satellite and the user, an optimization objective function is constructed.
[0011] The resource allocation problem corresponding to the optimization objective function is reconstructed into a Markov Decision Process (MDP).
[0012] The PPO algorithm is optimized using a near-end strategy to solve the MDP and dynamically allocate low-Earth orbit satellite communication and navigation resources.
[0013] Optionally, in the system, low-Earth orbit satellites transmit ICaN signals modulated by orthogonal frequency division multiplexing (OFDM) to users. The ICaN signals include information data carrying communication data and satellite ephemeris information, as well as measurement data consisting of structured sequences used for synchronization and navigation ranging.
[0014] Optionally, the objective function is to maximize the weighted sum of communication capacity and positioning accuracy; wherein, the communication capacity is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the user's received channel, and the positioning accuracy is quantified by the Cramér-Rao lower bound of the standard deviation of the ranging error and the geometrical precision factor (GDOP).
[0015] Optionally, the MDP includes a state space, an action space, state transition probabilities, and a reward function;
[0016] The state space includes the geometric precision factor GDOP at the current time, the radial relative velocity between the user and the satellite, the signal-to-interference-plus-noise ratio (SINR), and the pilot insertion subcarrier spacing, pilot insertion symbol spacing, and communication resource allocation factor selected at the previous time step.
[0017] The action space in time slot t includes the pilot configuration scheme and the weight ratio of communication and navigation functions. The pilot configuration scheme includes the pilot insertion subcarrier interval and the pilot insertion symbol interval.
[0018] The reward function is constructed based on communication capacity and positioning accuracy.
[0019] Optionally, dynamic allocation of low-Earth orbit satellite communication and navigation resources can be performed, including:
[0020] Initialize the environment, action space, and deep neural network parameters; set up the experience replay buffer.
[0021] For each round of training, obtain the initial state, and at each time step execute: select an action based on the current state and policy, calculate the immediate reward, observe the next state, and store the state, action, reward and the transition tuple of the next state into the experience replay buffer.
[0022] Randomly sample small batches of transition tuples from the experience replay buffer, calculate the dominance function value, and update the parameters of the deep neural network through stochastic gradient descent.
[0023] Repeated training continues until the maximum number of training rounds is reached to obtain the optimal resource allocation strategy. This strategy adaptively adjusts the distribution of structured pilots based on the dynamic channel state and relative position changes of low-Earth orbit satellites and users, thereby achieving efficient allocation of communication and navigation resources.
[0024] Secondly, embodiments of this application provide a resource allocation device for an integrated communication and navigation system, comprising:
[0025] The first building block is used to build a system that integrates communication and navigation for ground-based ICaN users using a low-Earth orbit satellite constellation.
[0026] The second construction module is used to construct an optimization objective function based on the channel state of the system and the relative positional relationship between the satellite and the user;
[0027] The first processing module is used to reconstruct the resource allocation problem corresponding to the optimization objective function into a Markov Decision Process (MDP).
[0028] The second processing module is used to optimize the PPO algorithm using a near-end strategy, solve the MDP, and dynamically allocate low-orbit satellite communication and navigation resources.
[0029] Optionally, in the system, low-Earth orbit satellites transmit ICaN signals modulated by orthogonal frequency division multiplexing (OFDM) to users. The ICaN signals include information data carrying communication data and satellite ephemeris information, as well as measurement data consisting of structured sequences used for synchronization and navigation ranging.
[0030] Optionally, the objective function is to maximize the weighted sum of communication capacity and positioning accuracy; wherein, the communication capacity is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the user's received channel, and the positioning accuracy is quantified by the Cramér-Rao lower bound of the standard deviation of the ranging error and the geometrical precision factor (GDOP).
[0031] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method as described in any one of the first aspects.
[0032] Fourthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method as described in any one of the first aspects.
[0033] Compared with existing technologies, the resource allocation method, apparatus, medium, and product of the integrated communication and navigation system provided in this application construct a system for low-Earth orbit (LEO) satellite constellations to serve ground-based integrated communication and navigation ICaN users. Based on the system's channel state and the relative positional relationship between satellites and users, an optimization objective function is constructed. The resource allocation problem corresponding to the optimization objective function is reconstructed as a Markov Decision Process (MDP). A near-end strategy optimization (PPO) algorithm is used to solve the MDP, enabling dynamic allocation of LEO satellite communication and navigation resources. This application's solution, through a technical framework of dynamic state awareness, model-free reinforcement learning, and multi-objective collaborative optimization, specifically overcomes the three major bottlenecks of existing methods in LEO satellite ICaN scenarios, achieving dynamic adaptability, model independence, and performance synergy in resource allocation, providing an efficient solution. Attached Figure Description
[0034] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0035] Figure 1 This is a flowchart illustrating a resource allocation method for an integrated communication and navigation system according to an embodiment of this application.
[0036] Figure 2 A structural diagram of the LEO-ICaN system model provided in the embodiments of this application;
[0037] Figure 3 This is a schematic diagram of the ICaN signal dynamic configuration scheme according to an embodiment of this application;
[0038] Figure 4This is a structural diagram of the resource allocation device of the integrated communication and navigation system according to an embodiment of this application. Detailed Implementation
[0039] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0040] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0041] As described in the background section, existing resource allocation methods lack dynamic adaptability in low-Earth orbit satellite communication and navigation integration (LEO-ICaN) scenarios, rely on accurate modeling, and struggle to simultaneously consider communication and positioning performance. To address these issues, this application provides a resource allocation method, apparatus, medium, and product for an integrated communication and navigation system, which can reduce or avoid the occurrence of the above situations.
[0042] This application provides a resource allocation method, apparatus, medium, and product for an integrated communication and navigation system. The method and apparatus are based on the same concept, and since they solve problems based on similar principles, their implementations can be mutually referenced; repeated details will not be repeated.
[0043] It should be understood that communication and navigation resource management and optimization is a highly interdisciplinary research field, integrating multiple technologies such as wireless communication, satellite navigation, signal processing, and artificial intelligence. Its core objective is to rationally coordinate resource allocation between communication and navigation tasks under limited spectrum and time resources, in order to achieve efficient and reliable information transmission and accurate position awareness. This application, based on deep reinforcement learning methods and oriented towards dynamic network environments, adaptively adjusts the resource allocation scheme through intelligent strategies, balancing communication capacity and navigation accuracy, and significantly improving the overall system performance and configuration flexibility.
[0044] Please refer to Figure 1 This application provides a resource allocation method for an integrated communication and navigation system, comprising:
[0045] Step 11: Construct a system that integrates communication and navigation for ICaN users on the ground via a low-Earth orbit satellite constellation.
[0046] In this embodiment, step 11 is the system foundation construction stage of the method. Its core is to establish a service architecture for ground-based ICaN users of the low-Earth orbit (LEO) satellite constellation, clarifying the signal transmission and functional carrying logic. The LEO satellite constellation of this application is configured with multiple LEO satellites, forming a constellation network covering ground users. The satellites have the capability to transmit and receive integrated communication and navigation signals. Preferably, there are at least four LEO satellites to meet the geometric coverage requirements of the positioning function.
[0047] Low-Earth orbit (LEO) satellites transmit ICaN signals modulated using Orthogonal Frequency Division Multiplexing (OFDM) to ground users. These signals consist of two core data categories: information data and measurement data. Information data carries user communication data, such as voice, video, and text, as well as satellite ephemeris information, which is used by users to calculate satellite positions. Measurement data comprises structured sequences and is used to achieve time synchronization and distance measurement between users and satellites. Ground-based ICaN users receive ICaN signals from multiple satellites, simultaneously acquiring the measurement information required for communication services and positioning, thus providing the foundation for subsequent resource allocation—the service recipients and signal carriers.
[0048] Step 12: Based on the channel state of the system and the relative positional relationship between the satellite and the user, construct an optimization objective function.
[0049] Here, step 12 transforms the need to balance communication and navigation performance into quantifiable mathematical objectives, providing direction for resource allocation.
[0050] Optionally, obtain channel status and the relative positional relationship between satellites and users, including:
[0051] Obtaining channel status includes: obtaining the signal-to-interference-plus-noise ratio (SINR) of the satellite-to-ground link through feedback from the user receiver or monitoring on the satellite side, which reflects the quality of the communication channel. The higher the SINR, the stronger the communication capability.
[0052] Obtaining the relative positional relationship between satellites and users includes: calculating the geometrical factor of precision (GDOP) using satellite ephemeris and preliminary user positioning information; quantifying the impact of satellite constellation topology on positioning accuracy; the smaller the GDOP, the weaker the amplification effect of positioning error.
[0053] In this application, the objective function is constructed with the goal of maximizing the synergistic performance of communication capacity and positioning accuracy. The objective function is constructed in the form of a weighted sum.
[0054] Step 13: Reconstruct the resource allocation problem corresponding to the optimization objective function into a Markov Decision Process (MDP).
[0055] In this embodiment of the application, step 13 is the problem transformation step of the method, which abstracts the complex resource allocation optimization problem into a reinforcement learning-solvable MDP framework to realize the association between "environment-policy-performance".
[0056] The MDP consists of a four-tuple: a state space S, an action space A, state transition probabilities P, and a reward function R. Each part is designed to fit the characteristics of the ICaN (Independent Credential Network) scenario for low-Earth orbit (LEO) satellites. The state space S includes current environmental observations (GDOP, radial relative velocity between the user and the satellite, SINR) and historical decision information (previous pilot insertion interval, communication resource allocation factor), comprehensively reflecting the dynamic factors affecting resource allocation. The action space A defines the resource adjustment operations that the agent can perform, including pilot configuration schemes, such as time-domain or frequency-domain insertion intervals, affecting positioning accuracy, and the weighted ratio of communication and navigation performance. The weighted ratio influences the tilt direction of resource allocation. The state transition probabilities P describe the probabilistic relationship between "current state, action, and next state," requiring no precise modeling and automatically learned through subsequent reinforcement learning and environmental interaction. The reward function R is derived from the optimization objective function and directly reflects the performance effect of resource allocation actions. For example, higher communication capacity and better positioning accuracy result in a larger reward value, and the pilot proportion ρ balances the weights of communication and navigation resources.
[0057] Step 14: The PPO algorithm is optimized using a near-end strategy to solve the MDP and dynamically allocate low-Earth orbit satellite communication and navigation resources.
[0058] In this embodiment, the Near-End Policy Optimization (PPO) algorithm is used to solve the MDP and dynamically allocate low-Earth orbit satellite communication and navigation resources. The optimal resource allocation strategy is learned through the PPO algorithm to achieve dynamic and adaptive resource scheduling.
[0059] This application first models the time-frequency resource structure of the integrated communication and navigation signal to achieve communication and navigation fusion. Considering the rapidly changing characteristics of the low-Earth orbit satellite environment, this application proposes a joint optimization problem. This problem dynamically configures communication and navigation resources to maximize communication capacity and positioning accuracy, thereby achieving joint optimization of communication and navigation performance. In order to solve this non-convex problem and adapt to the rapidly changing environment, this application adopts a deep reinforcement learning algorithm to dynamically adjust the pilot configuration to approach the optimal solution.
[0060] Optionally, in the system, low-Earth orbit satellites transmit ICaN signals modulated by orthogonal frequency division multiplexing (OFDM) to users. The ICaN signals include information data carrying communication data and satellite ephemeris information, as well as measurement data consisting of structured sequences used for synchronization and navigation ranging.
[0061] Specifically, such as Figure 2 As shown, this application constructs a system architecture for a low-Earth orbit (LEO) satellite constellation to serve ground-based ICaN users. To ensure the positioning function, at least four satellites are used for service. The LEO satellite s transmits the ICaN signal to the ground user u at time index t. The signal uses an orthogonal frequency division multiplexing (OFDM) modulation scheme, and its expression is shown below:
[0062]
[0063] The ICaN signal is divided into two parts; The communication data portion includes communication data and satellite ephemeris information; the other portion... The measurement data consists of a structured sequence used for synchronization and navigation ranging. Here, m and q represent the symbol index (time domain) and subcarrier index (frequency domain) occupied in the OFDM system, respectively, used to accurately locate the distribution of information and measurement data on time and frequency resources. s,u (t) represents the time-domain composite signal transmitted by satellite s to user n, integrating communication and navigation functions; it is a superposition of OFDM symbols in the time domain. T represents the OFDM symbol period, which determines the symbol duration and affects the signal's time-domain resolution and multipath immunity. c The index representing the communication data symbol, indicating the m-th... c Each OFDM symbol (distinguishing communication data across different time slots / symbols); represents the communication subcarrier index, corresponding to the subcarrier position in the OFDM frequency domain, q c The value of x determines the subcarrier frequency; s,u,c Modulation symbols, which represent communication data symbols, carry user communication information and are the baseband mapping values of communication data. This represents the time-domain complex exponential modulation of the communication subcarrier, realizing the frequency shift of the OFDM subcarrier, (t-(m c-1)T) represents the m-th... c The effective time domain interval of each symbol. P The index of the navigation pilot symbol indicates the m-th digit. P One navigation OFDM symbol (distinguishing navigation pilots from different time slots / symbols). q P This represents the navigation pilot subcarrier index, corresponding to the subcarrier position allocated to navigation functions in the OFDM frequency domain. PRN represents the pseudo-random noise (PRN) sequence, which is the core structured sequence for navigation ranging, used to achieve time synchronization between the user and the satellite, and distance measurement. This represents the time-domain complex exponential modulation of the navigation pilot subcarrier, similar to the communication section, to achieve frequency shifting of the navigation subcarrier, (t-(m P -1)T) represents the m-th... P The effective time domain range of each navigation symbol.
[0064] This formula achieves integrated communication data transmission and navigation ranging services between low-Earth orbit satellites and ground users through the time-domain superposition of communication subcarriers and navigation pilot subcarriers. The communication function is achieved through x s,u,c Modulated communication information is transmitted in parallel using OFDM subcarriers to meet high data rate requirements; the navigation function achieves accurate ranging through PRN sequences, and utilizes the orthogonality of OFDM subcarriers to avoid conflicts between communication and navigation resources, supporting positioning calculations.
[0065] This application aims to dynamically optimize the resource allocation scheme for communication and measurement data in ICaN signals in response to topology and environmental changes in the LEO satellite constellation. For example... Figure 3 As shown, within the time-frequency resource grid, apart from the portion used to carry communication data, the remaining area is arranged in a structured manner using pilot symbols and synchronization symbols. With increasing positioning accuracy requirements and the intensified Doppler frequency shift effect caused by high-speed motion, the pilot configuration density needs to be gradually increased. However, increasing pilot density will consume some communication data resources, thus raising a trade-off between communication and navigation resource allocation.
[0066] Optionally, the objective function is to maximize the weighted sum of communication capacity and positioning accuracy; wherein, the communication capacity is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the user's received channel, and the positioning accuracy is quantified by the Cramér-Rao lower bound of the standard deviation of the ranging error and the geometrical precision factor (GDOP).
[0067] Specifically, in order to reasonably construct the optimization objective function and achieve joint optimization of communication and navigation performance, the relevant system principles are analyzed and modeled before designing the optimization objective function.
[0068] Channel Model: The ICaN signal generated by the s-th LEO satellite is transmitted to the user through the lower planetary ground channel, and the channel vector is determined by h.s,u This indicates that the channel vector h s,u The expression is as follows:
[0069]
[0070] Among them, h s,u This represents the composite channel vector from the satellite to the user, describing channel characteristics such as signal attenuation and phase during transmission, which affect communication quality and navigation ranging accuracy. s,u This represents the channel gain vector, reflecting the scaling effect of distance, antenna, and environment on signal amplitude, dynamically adapting to channel changes. K s Represents the Rice factor, quantifying the power ratio of line-of-sight (direct) and non-line-of-sight (reflected / scattered) signals, distinguishing channels where direct signal is dominant (K). s (Large) or multipath (K) s Small). and These represent the line-of-sight (LOS) and non-line-of-sight (NOS) components of the Ricean channel, respectively. The LOS channel vector corresponds to the direct path signal, which is fundamental to accurate navigation ranging and high-quality communication transmission. The NOS channel vector corresponds to the reflected / scattered path signal, leading to multipath fading in communication and navigation ranging errors, and exhibits strong randomness. Here, the channel vector h... s,u By weighting and superimposing line-of-sight and non-line-of-sight components, and combining gain and Rice factor, the characteristics of the satellite-user channel are accurately characterized, supporting the dynamic optimization of communication and navigation resource allocation.
[0071] Channel theoretical capacity: One of the important indicators for analyzing communication performance is the channel's theoretical capacity (Shannon capacity). The expression for the ICaN signal received by user u from a satellite is as follows, used to represent the full-frequency multiplexing of the beam of the same low-Earth orbit satellite s, leading to mutual interference between the signals of different users (u and u′). The formula describes the superposition process of the desired signal, intra-satellite multi-user interference, and noise when user u receives the signal.
[0072]
[0073] Where u' represents other users of satellite s; y s,u This represents the baseband signal received by user u from satellite s, which includes the desired signal, interference, and noise. x represents the channel vector from satellite s to user u, describing the propagation characteristics of the signal from the satellite to user u. s,u This represents the desired signal (communication data / navigation pilot) that satellite s sends to user u. Indicates multi-user interference within the satellite; C s,u′U represents the set of users served by satellite s, such as all users covered by the same satellite beam; u′ represents other users served by satellite s, which share satellite frequency resources with user u, causing interference. The channel vector from satellite s to interfering user u′ represents the characteristics of the interfering signal after transmission through the channel; x s,u′ This indicates that the signal transmitted by satellite s to user u′ is co-channel interference for u; n s,u This represents the noise received by user u. This expression accurately models the intra-satellite multi-user interference problem under full frequency reuse of low-Earth orbit satellites, and is the underlying basis for the technical link of this application's method in sensing interference, modeling problems, and optimizing resources.
[0074] Optionally, in the formula for the received signal y s,u Li n s,u This represents the noise introduced when user u receives a signal from satellite s, where, Follows a mean of 0 and a variance of The signal-to-interference-plus-noise ratio (SINR) exhibits a Gaussian distribution because, within the LEO satellite constellation, the beams of the same satellite utilize full frequency reuse, resulting in intra-satellite interference. s,u The expression is:
[0075]
[0076] Therefore, the total capacity of the receiving channel for user u is: C u =T c,u B c,u log2(1+SINR s,u ); where T c,u and B c,u These represent the equivalent bandwidth and equivalent time span used by user u for communication, respectively.
[0077] The Cramér-Rao Lower Bound (CRLB) provides a theoretical lower bound on the variance of unbiased parameter estimates given a measurement model and noise covariance. Based on Fisher information analysis of the time-of-arrival parameter estimation, the Cramér-Rao lower bound for the standard deviation of the ranging error is:
[0078]
[0079] Where, σ ρm σ represents the lower bound of the standard deviation of the ranging error, Cramér-Rao; τ The standard deviation of the TOA measurement error is represented by c; the speed of light is b. n B is the receiver's equivalent noise bandwidth. eff For the ranging equivalent bandwidth, T effThe effective observation time and the time of arrival (TOA) together determine the accuracy of ranging. In OFDM systems, the number of pilot subcarriers and their frequency domain distribution directly affect the effective ranging bandwidth B. eff The more pilot signals there are, and the wider and more uniform their distribution, the greater the mean square bandwidth (B) of the signal power spectrum. eff The larger the value of ), the lower the lower bound of the Cramér-Rao, and thus the higher the ranging accuracy. Simultaneously, the observation duration of the pilot in the time domain (i.e., T) eff The larger the value, the more effective ranging information the receiver accumulates, the more Fisher information is added, and the CRLB is further reduced, which helps to improve ranging accuracy.
[0080] GDOP: The topology of the LEO satellite constellation has a significant impact on navigation accuracy, and this impact is usually quantified by the Geometric Precision Factor (GDOP). Assume the user's three-dimensional coordinates are: p u =[x u ,y u ,z u ] T The three-dimensional coordinates of the satellite serving the user are p s =[x s ,y s ,z s ] T The distance between the satellite and the user is d. u,s =‖p u -p s ‖2, The expression for GDOP is Where G is a geometric matrix, its expression is as follows:
[0081]
[0082] in Let D be the unit vector pointing from the user to the s-th satellite. It is evident that the GDOP value changes with the positional relationship between the satellite and the user. GDOP amplifies positioning errors, and the relationship is as follows: D u =GDOP·σ ρm ; where D u Defined as positioning error, it can be seen that the larger the GDOP, the more significantly the positioning error is amplified, and the positioning accuracy deteriorates significantly.
[0083] This application aims to flexibly adjust the distribution of structured pilots based on changes in the relative positions of the satellite constellation and users in the channel state (i.e., dynamic changes in GDOP and SINR) within a real LEO-ICaN system, thereby maximizing communication capacity and positioning accuracy. Its optimization objective function is as follows:
[0084]
[0085] stSINR s,u ≥SINR min
[0086] k p-sym,s,u ≤k p-sym,max
[0087] k p-sc,s,u ≤k p-sc,max
[0088]
[0089] α+β=1
[0090] Where α and β are the weighting parameters for communication and navigation performance, respectively. p-sym and k p-sub T represents the pilot insertion interval in the time and frequency domains, respectively, and is measured by the number of resource elements (REs) within an OFDM time-frequency resource block (RB). p T represents the actual time interval for time-domain pilot insertion; c T is the channel coherence time. p The design must meet the requirements of counteracting the Doppler frequency shift effect caused by the high-speed motion of the satellite. Due to 1 / D u It can be considered as a representation of positioning accuracy, minimizing the measurement error D u This is equivalent to maximizing positioning accuracy. The problem described above is a mixed-integer nonlinear programming (MINLP) problem, which is nonconvex. Furthermore, in the LEO-ICaN system, the environment changes rapidly, and the parameters within the constraints also change dynamically. Therefore, traditional mathematical optimization methods are difficult to solve this problem effectively. To address this issue, this application employs a deep reinforcement learning (DRL) method.
[0091] Optionally, the MDP includes a state space, an action space, state transition probabilities, and a reward function;
[0092] The state space includes the geometric precision factor GDOP at the current time, the radial relative velocity between the user and the satellite, the signal-to-interference-plus-noise ratio (SINR), and the pilot insertion subcarrier spacing, pilot insertion symbol spacing, and communication resource allocation factor selected at the previous time step.
[0093] The action space in time slot t includes the pilot configuration scheme and the weight ratio of communication and navigation functions. The pilot configuration scheme includes the pilot insertion subcarrier interval and the pilot insertion symbol interval.
[0094] The reward function is constructed based on communication capacity and positioning accuracy.
[0095] In this embodiment, the objective function is reconstructed into a Markov Decision Process (MDP) operated by an agent deployed on a LEO satellite at discrete time steps. This MDP consists of a quadruple (S, A, P, R), where S is the state space, A is the action space, P represents the state transition probability, and R is the reward function, defined as follows.
[0096] 1) The state space S represents the state space St in time slot t. t The expression is as follows:
[0097] S t =(GDOP) t ,v rel,t SINR t ,k p-sc,t-1 ,k p-sym,t-1 ,α t-1 )
[0098] Among them, GDOP t v rel,t SINR t These are the geometric precision factor, the radial relative velocity between the user and the satellite, and the signal-to-interference-plus-noise ratio at the current moment, respectively. p-sc,t-1 k p-sym,t-1 α t-1 These are the pilot insertion subcarrier interval, pilot insertion symbol interval, and communication resource allocation factor selected in the previous time step, respectively.
[0099] 2) Action space A represents the action space that the agent can execute. It mainly includes the pilot configuration scheme and the weight ratio of communication and navigation functions. Its action space in time slot t is as follows: A t =(k p-sc,t ,k p-sym,t ,α t ); where k p-sc,t k p-sym,t α t These are the pilot insertion subcarrier interval, pilot insertion symbol interval, and communication resource allocation factor selected by the agent at the current time step.
[0100] 3) The reward function R plays a crucial role in shaping the learning behavior of the reinforcement learning agent. It consists of the first objective function described below, and the relevant constraints are already integrated into the environment design and do not need to be reflected as penalty terms. The expression for the first objective function of the reward function Rt is as follows:
[0101]
[0102] Where ρ represents the proportion of pilots in the frequency or time domain, and its value is determined by the pilot spacing k. R comm and R navThese represent communication capacity and positioning accuracy, respectively. C shannon Indicates the maximum Shannon capacity. (D) u denoted as positioning error. α and β are the resource allocation factors for communication and navigation, respectively. t a t These represent the state space and action space at the current moment, respectively. Due to the order-of-magnitude difference between communication capacity and positioning accuracy, normalization (Norm) is required to achieve joint optimization of both.
[0103] Proximal Policy Optimization (PPO) is a model-free, policy-based approach that iteratively optimizes a policy using a pruned surrogate objective function to approximate the optimal solution. Compared to other DRL methods, PPO provides more robust updates and exhibits superior stability in high-dimensional and complex scenarios. Let π denote the action selection policy function, which is determined by parameters θ.
[0104] It consists of a deep neural network (DNN). Within the PPO framework, the policy function is executed. After a certain number of steps, the expression for the advantage function is as follows:
[0105]
[0106] Where V(st) is the state value function, used to calculate the variance-reduced advantage function estimate. γ is a discount factor used to determine the importance of future rewards; r t+k This is the immediate reward obtained at time step t+k. To stabilize policy updates, PPO introduces a pruning second objective function to limit the magnitude of policy changes. The expression for this second objective function is:
[0107]
[0108] Among them, s t a t These represent the state space and action space, respectively. π represents the action selection policy function. θ old This is the reference policy network used to compute the dominance function and importance sampling ratio. The hyperparameter ∈ is used to adjust the pruning scale of the pruning range to limit the magnitude of policy updates. The corresponding parameter θ is updated using a mini-batch stochastic gradient descent (SGD) method based on B transition samples (st, at, rt, st+1), and its expression is:
[0109]
[0110] Where ▽ represents the gradient calculation.
[0111] Optionally, dynamic allocation of low-Earth orbit satellite communication and navigation resources can be performed, including:
[0112] Initialize the environment, action space, and deep neural network parameters; set up the experience replay buffer.
[0113] For each round of training, obtain the initial state, and at each time step execute: select an action based on the current state and policy, calculate the immediate reward, observe the next state, and store the state, action, reward and the transition tuple of the next state into the experience replay buffer.
[0114] Randomly sample small batches of transition tuples from the experience replay buffer, calculate the dominance function value, and update the parameters of the deep neural network through stochastic gradient descent.
[0115] Repeated training continues until the maximum number of training rounds is reached to obtain the optimal resource allocation strategy. This strategy adaptively adjusts the distribution of structured pilots based on the dynamic channel state and relative position changes of low-Earth orbit satellites and users, thereby achieving efficient allocation of communication and navigation resources.
[0116] In this embodiment, step 1 is performed first: initialization. This involves initializing the environment, action space, deep neural network parameters θ, and experience replay buffer.
[0117] The environment includes: the actual operating scenario corresponding to low-orbit satellite communication and navigation, including channel status such as channel gain and SINR, the relative positional relationship between the satellite and the user such as geometrical precision factor GDOP, radial relative velocity, interference and other dynamically changing factors, providing an interactive virtual operating scenario for the intelligent agent (resource allocation decision-making subject).
[0118] The action space is used to define the set of resource allocation operations that the agent can perform. In this scenario, it typically includes pilot configuration schemes (such as pilot insertion subcarrier spacing and symbol spacing), communication and navigation performance weighting parameters, etc. These actions directly determine the allocation method of communication and navigation resources.
[0119] Deep neural networks act as the brain of an agent, learning resource allocation strategies. After initialization, the parameters θ are continuously adjusted and optimized through training to enable the agent to output resource allocation strategies adapted to the environment. For example, the network can output probability distributions of different pilot configurations and resource weights based on the input environmental state, allowing the agent to select actions.
[0120] The experience replay buffer is used to store the transition tuples of "state-action-reward-next state" generated by the agent's interaction with the environment. During subsequent training, these historical data are randomly sampled to break the data correlation, stabilize the training process, and improve the efficiency of policy learning.
[0121] Step 2 involves multiple rounds of training iterations. Step 2.1: Obtain the initial state s t (t=0). At the beginning of each training round, the initial state s at time t=0 is obtained from the environment.t This is the starting point for the agent to make decisions. The initial state contains basic information about the current environment, such as the initial channel quality between the satellite and the user (initial SINR), the geometric positional relationship (initial GDOP), and the remaining resource allocation state from the previous time step (if any), providing initial observational basis for decisions in subsequent time steps. Step 2.2: Time Step Iteration (Time-by-Time Interaction and Learning). For each time step t (from 1 to the maximum number of steps), the following operations are performed sequentially:
[0122] a, based on the current state s t And strategy, choose action a t The agent determines the current environmental state based on s. t Using a learned strategy (determined by the parameters θ of the deep neural network), a specific action 'a' is selected from the action space. t For example, given the current state of low channel SINR (poor communication quality) and high GDOP (low positioning accuracy), the strategy might choose a combination of actions: increasing the pilot subcarrier spacing (to optimize positioning accuracy) and adjusting the communication resource allocation factor (to ensure basic communication).
[0123] b, according to the above expression Get instant rewards t Action a t After execution, the environmental state changes, and the immediate reward r resulting from the action is calculated based on this expression. t The reward value quantifies the quality of an action; if the action improves the weighted performance of communication capacity and positioning accuracy, the reward r is increased. t If the value is positive and relatively large, the reward may be negative if it leads to a decrease in performance, thus guiding the agent to learn a strategy of seeking benefits and avoiding harm.
[0124] c, observe the next state s t+1 Execute action a t Subsequently, the environmental state will be updated due to factors such as changes in resource allocation and the relative motion between the satellite and the user. The intelligent agent observes and records the new state s. t+1 It contains information such as new channel states, location relationships, and the effects of resource allocation, which is used to learn the rules of environmental state transitions.
[0125] d, will transfer tuple (s) t a t r t s t+1 Store the data in the experience replay buffer. Save the "state-action-reward-next state" data of the current time step to the experience replay buffer to accumulate interaction experience, provide diverse data samples for subsequent batch training, and avoid the agent overfitting to the environmental characteristics at the current moment.
[0126] e. Randomly sample a mini-batch of size B from the experience replay buffer. To make neural network training more stable and efficient, B transition tuples from different time steps and scenarios are randomly selected from the experience replay buffer to form a mini-batch of data. This can cover diverse environmental states and action feedback, making the trained policy more generalizable and adaptable to the complex and ever-changing environment in low-Earth orbit satellite scenarios.
[0127] f, according to the above expression Calculate the advantage function value. The advantage function measures the degree of advantage of the current action compared to other possible actions. Using this expression and sampled mini-batch data, calculate the advantage function value for each transition tuple. This helps the agent more accurately determine the long-term value of actions and avoid short-sighted decisions caused by relying solely on immediate rewards. For example, some actions may not have high immediate rewards, but they can bring higher returns in the long run; the advantage function can capture this characteristic.
[0128] g, using the above expression The neural network parameters θ are updated using stochastic gradient descent (SGD). Based on the calculated advantage function value, the parameters θ of the deep neural network are adjusted according to this expression using the stochastic gradient descent algorithm. Each update optimizes the neural network in the direction that yields higher cumulative rewards, gradually improving the quality of the strategy and making its output resource allocation actions more suitable for the integrated needs of low-Earth orbit satellite communication and navigation.
[0129] h, will s t Updated to s t+1 Then proceed to the next iteration. After completing the training at the current time step, update the current state to the next state s. t+1 Then, we move on to the next time step of interaction and training, continuously optimizing the strategy under different environmental conditions.
[0130] Step 3: Repeat training until termination. Continuously repeat the training process in Step 2 above until the pre-set maximum number of training rounds is reached. As the number of training rounds increases, the agent's interaction experience with the environment becomes richer, and the deep neural network parameters θ are gradually optimized. Ultimately, an optimal strategy is obtained in low-Earth orbit satellite communication and navigation scenarios, which can adaptively adjust the structured pilot distribution and efficiently allocate communication and navigation resources according to dynamic channel states and relative position changes.
[0131] The PPO algorithm in this application revolves around a closed loop of "agent-environment interaction - experience accumulation - training and optimization strategy". By leveraging the advantages of the PPO algorithm, it solves the problem of dynamic resource allocation in the integrated low-orbit satellite communication and navigation scenario, enabling the system to adapt to complex and ever-changing environments and balance communication and navigation performance.
[0132] In summary, this application addresses the low-Earth orbit (LEO) satellite communication and navigation integrated system by constructing an optimization problem model aimed at maximizing communication capacity and positioning accuracy, taking into account its rapidly changing geometric topology and environmental characteristics. This model aims to achieve efficient and coordinated allocation of communication and navigation resources. Starting from the time-frequency structure design of the signal, a method for optimizing pilot insertion is proposed to address the problem of low resource utilization efficiency in LEO satellite communication and navigation integrated systems. This method achieves fine-grained control of pilot resources by flexibly configuring the insertion positions of pilots in the time and frequency domains, thereby improving communication performance while ensuring navigation and positioning accuracy. In solving the problem of resource allocation in integrated communication and navigation, a deep reinforcement learning method is employed for non-convex problems. This method can continuously learn and optimize resource scheduling strategies through interaction with the environment without relying on precise mathematical models, achieving an efficient trade-off and dynamic coordination between communication capacity and positioning accuracy.
[0133] (1) Traditional communication and navigation systems are typically designed separately, making it difficult to balance communication performance and positioning accuracy with limited resources. Even when integration is achieved, it is often a static configuration, lacking dynamic adaptability. This application addresses the characteristics of low-Earth orbit satellite constellations, such as rapidly changing geometry and environment, and limited resources. It constructs a joint optimization model, explicitly defining "maximizing communication capacity and positioning accuracy" as the optimization objective function, and coordinating the allocation strategy of communication and navigation resources. This achieves dynamic equilibrium optimization of communication performance and navigation accuracy under resource constraints.
[0134] (2) Traditional pilot designs often use fixed structures or preset intervals, which are difficult to adapt flexibly to different tasks, resulting in wasted spectrum or degraded positioning performance. This application starts from the time-frequency structure of the signal and designs a flexible insertion mechanism for pilots in the time and frequency domains, which can dynamically adjust the pilot intervals and realize structured control of resources.
[0135] (3) Traditional static or rule-driven resource allocation strategies lack responsiveness when faced with high-speed movement of low-Earth orbit satellites and rapid channel changes, resulting in low resource utilization efficiency. This application introduces a deep reinforcement learning method to model the resource allocation problem as a Markov decision process (MDP). It utilizes an agent to adaptively learn the optimal resource allocation strategy in complex environments, thereby avoiding the modeling difficulties and rigid rule limitations faced by traditional methods. It possesses strong environmental adaptability and robustness, and can still achieve efficient resource allocation in complex and ever-changing satellite communication and navigation scenarios.
[0136] The various methods described above are based on embodiments of this application. Apparatus for implementing the above methods will now be provided.
[0137] Please refer to Figure 4 This application also provides a resource allocation device for an integrated communication and navigation system, comprising:
[0138] The first building module 41 is used to build a system for low-orbit satellite constellation to serve ground-based integrated communication and navigation ICaN users;
[0139] The second construction module 42 is used to construct an optimization objective function based on the channel state of the system and the relative positional relationship between the satellite and the user.
[0140] The first processing module 43 is used to reconstruct the resource allocation problem corresponding to the optimization objective function into a Markov decision process (MDP).
[0141] The second processing module 44 is used to optimize the PPO algorithm using a near-end strategy, solve the MDP, and dynamically allocate low-orbit satellite communication and navigation resources.
[0142] Optionally, in the system, low-Earth orbit satellites transmit ICaN signals modulated by orthogonal frequency division multiplexing (OFDM) to users. The ICaN signals include information data carrying communication data and satellite ephemeris information, as well as measurement data consisting of structured sequences used for synchronization and navigation ranging.
[0143] Optionally, the objective function is to maximize the weighted sum of communication capacity and positioning accuracy; wherein, the communication capacity is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the user's received channel, and the positioning accuracy is quantified by the Cramér-Rao lower bound of the standard deviation of the ranging error and the geometrical precision factor (GDOP).
[0144] Optionally, the MDP includes a state space, an action space, state transition probabilities, and a reward function;
[0145] The state space includes the geometric precision factor GDOP at the current time, the radial relative velocity between the user and the satellite, the signal-to-interference-plus-noise ratio (SINR), and the pilot insertion subcarrier spacing, pilot insertion symbol spacing, and communication resource allocation factor selected at the previous time step.
[0146] The action space in time slot t includes the pilot configuration scheme and the weight ratio of communication and navigation functions. The pilot configuration scheme includes the pilot insertion subcarrier interval and the pilot insertion symbol interval.
[0147] The reward function is constructed based on communication capacity and positioning accuracy.
[0148] Optionally, dynamic allocation of low-Earth orbit satellite communication and navigation resources can be performed, including:
[0149] Initialize the environment, action space, and deep neural network parameters; set up the experience replay buffer.
[0150] For each round of training, obtain the initial state, and at each time step execute: select an action based on the current state and policy, calculate the immediate reward, observe the next state, and store the state, action, reward and the transition tuple of the next state into the experience replay buffer.
[0151] Randomly sample small batches of transition tuples from the experience replay buffer, calculate the dominance function value, and update the parameters of the deep neural network through stochastic gradient descent.
[0152] Repeated training continues until the maximum number of training rounds is reached to obtain the optimal resource allocation strategy. This strategy adaptively adjusts the distribution of structured pilots based on the dynamic channel state and relative position changes of low-Earth orbit satellites and users, thereby achieving efficient allocation of communication and navigation resources.
[0153] It should be noted that the device in this embodiment corresponds to the method described above. The implementation methods in each of the above embodiments are also applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this application embodiment can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Therefore, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail here.
[0154] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the resource allocation method embodiment of the above-described integrated communication and navigation system, achieving the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0155] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the resource allocation method embodiment of the above-described integrated communication and navigation system and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0156] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0158] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A resource allocation method for an integrated communication and navigation system, characterized in that, include: To construct a system that integrates communication and navigation for ground-based ICaN users using a low-Earth orbit satellite constellation; Based on the channel state of the system and the relative positions of the satellite and the user, an optimization objective function is constructed. The resource allocation problem corresponding to the optimization objective function is reconstructed into a Markov Decision Process (MDP). The PPO algorithm is optimized using a near-end strategy to solve the MDP and dynamically allocate low-Earth orbit satellite communication and navigation resources.
2. The method according to claim 1, characterized in that, The system transmits ICaN signals modulated by orthogonal frequency division multiplexing (OFDM) to users from low-Earth orbit satellites. The ICaN signals include information data carrying communication data and satellite ephemeris information, as well as measurement data consisting of structured sequences used for synchronization and navigation ranging.
3. The method according to claim 1, characterized in that, The objective function is to maximize the weighted sum of communication capacity and positioning accuracy; where communication capacity is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the user's receiving channel, and positioning accuracy is quantified by the Cramér-Rao lower bound of the standard deviation of ranging error and the geometric precision factor (GDOP).
4. The method according to claim 1, characterized in that, The MDP includes a state space, an action space, state transition probabilities, and a reward function; The state space includes the geometric precision factor GDOP at the current time, the radial relative velocity between the user and the satellite, the signal-to-interference-plus-noise ratio (SINR), and the pilot insertion subcarrier spacing, pilot insertion symbol spacing, and communication resource allocation factor selected at the previous time step. The action space in time slot t includes the pilot configuration scheme and the weight ratio of communication and navigation functions. The pilot configuration scheme includes the pilot insertion subcarrier interval and the pilot insertion symbol interval. The reward function is constructed based on communication capacity and positioning accuracy.
5. The method according to claim 3, characterized in that, Dynamic allocation of low-Earth orbit satellite communication and navigation resources, including: Initialize the environment, action space, and deep neural network parameters; set up the experience replay buffer. For each round of training, obtain the initial state, and at each time step execute: select an action based on the current state and policy, calculate the immediate reward, observe the next state, and store the state, action, reward and the transition tuple of the next state into the experience replay buffer. Randomly sample small batches of transition tuples from the experience replay buffer, calculate the dominance function value, and update the parameters of the deep neural network through stochastic gradient descent. Repeated training continues until the maximum number of training rounds is reached to obtain the optimal resource allocation strategy. This strategy adaptively adjusts the distribution of structured pilots based on the dynamic channel state and relative position changes of low-Earth orbit satellites and users, thereby achieving efficient allocation of communication and navigation resources.
6. A resource allocation device for an integrated communication and navigation system, characterized in that, include: The first building block is used to build a system that integrates communication and navigation for ground-based ICaN users using a low-Earth orbit satellite constellation. The second construction module is used to construct an optimization objective function based on the channel state of the system and the relative positional relationship between the satellite and the user; The first processing module is used to reconstruct the resource allocation problem corresponding to the optimization objective function into a Markov Decision Process (MDP). The second processing module is used to optimize the PPO algorithm using a near-end strategy, solve the MDP, and dynamically allocate low-orbit satellite communication and navigation resources.
7. The apparatus according to claim 6, characterized in that, The system transmits ICaN signals modulated by orthogonal frequency division multiplexing (OFDM) to users from low-Earth orbit satellites. The ICaN signals include information data carrying communication data and satellite ephemeris information, as well as measurement data consisting of structured sequences used for synchronization and navigation ranging.
8. The apparatus according to claim 6, characterized in that, The objective function is to maximize the weighted sum of communication capacity and positioning accuracy; where communication capacity is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the user's receiving channel, and positioning accuracy is quantified by the Cramér-Rao lower bound of the standard deviation of ranging error and the geometric precision factor (GDOP).
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 5.
Citation Information
Cited By
Satellite communication and navigation integrated system and satellite communication and navigation integrated method
CN121276555A
Energy-saving routing method and system in low-orbit satellite constellation
CN121585241A