A multi-satellite multi-beam reliable precoding method and system
By constructing a multi-satellite, multi-beam communication error model and training a deep reinforcement learning model, the problem of communication precoding under incomplete channel state information is solved, and the technical problem of communication error model under incomplete channel state information is realized.
Patent Information
- Application Number
- CN202510044032.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In existing low-Earth orbit satellite communication systems, linear precoding methods struggle to obtain accurate channel state information under high dynamic conditions, resulting in low throughput in some beam coverage areas. Furthermore, existing precoding methods fail to balance overall system throughput and fairness.
A multi-satellite, multi-beam reliable precoding method is adopted. By constructing a communication error model, a precoding matrix is trained using a deep reinforcement learning model. By combining throughput and fairness optimization indicators, a precoding scheme under incomplete channel state information is generated. The deep reinforcement learning model is trained by optimizing the combination of optimization indicators to perform coding operations and generate the precoding matrix.
This study optimizes the throughput and fairness of precoding schemes for satellite communication under conditions of incomplete channel state information, reduces the dependence on the accuracy of real-time channel state information, and improves the throughput of the system and the throughput and fairness of ground users.
Smart Images

Figure CN120034218B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of low-orbit satellite communication, in particular to a multi-satellite multi-beam reliable precoding method and system. BACKGROUND
[0002] Low-orbit satellite communication system is an important part of the fifth generation non-terrestrial network, which has great advantages when the ground network cannot provide stable and efficient communication services due to natural disasters, economic conditions and other factors. In recent years, the rapid development of low-orbit satellite communication has led to a shortage of spectrum resources. As a key technology that can effectively improve system capacity and resource utilization, precoding technology has been widely researched and applied in academia and industry.
[0003] However, the performance of linear precoding directly applied to low-orbit satellite communication system is poor. Precoding is generally divided into two categories: linear precoding and nonlinear precoding. Nonlinear precoding has not been widely used due to its high complexity; although linear precoding is relatively simple, it requires high accuracy of channel state information (CSI), and the high dynamics of low-orbit satellites make it difficult to obtain perfect CSI. Moreover, most existing precoding methods for low-orbit satellite systems only focus on the overall performance of the system, and these scheme designs may result in low throughput in some beam coverage areas. Therefore, how to design a reliable precoding method that can balance the total throughput and fairness of the system and is robust to imperfect CSI according to the characteristics of low-orbit satellite communication system has become a research hotspot in the field of low-orbit satellite precoding technology. SUMMARY
[0004] The present application aims to at least partially solve one of the problems in the related art.
[0005] To this end, the present application proposes a multi-satellite multi-beam reliable precoding method to maximize the throughput of the system while taking into account the coverage of each beam.
[0006] Another object of the present application is to propose a multi-satellite multi-beam reliable precoding system.
[0007] A third object of the present application is to propose a computer device.
[0008] A fourth object of the present application is to propose a non-transitory computer readable storage medium.
[0009] To achieve the above objects, the present application proposes, in one aspect, a multi-satellite multi-beam reliable precoding method, comprising:
[0010] establishing a multi-satellite multi-beam communication model and constructing a corresponding communication error model based on the established multi-satellite multi-beam communication model;
[0011] An optimization index is constructed from the perspective of maximizing and minimizing the signal-to-interference-plus-noise ratio of each antenna beam, based on taking into account the throughput and fairness of the optimized coordinated multi-beam satellite system;
[0012] A deep reinforcement learning model is constructed, and the deep reinforcement learning model is trained based on the optimization index, to perform an encoding operation on the error communication model to be precoded according to the trained deep reinforcement learning model, to obtain a final precoding matrix.
[0013] The multi-satellite multi-beam reliable precoding method of the embodiment of the application can further have the following additional technical features:
[0014] In an embodiment of the application, a corresponding communication error model is constructed based on the established multi-satellite multi-beam communication model, including:
[0015] An array model of the multi-beam antenna is constructed based on the characteristics of the linear uniform antenna array;
[0016] A multi-satellite multi-beam communication model is constructed based on the joint transmission mode of multiple low-orbit satellites using multi-beam antenna technology;
[0017] Based on the multi-satellite multi-beam communication model, the overall phase shift caused by the inter-satellite synchronization time delay error, and the angle error in the antenna direction caused by the rotation of the satellite and the positioning error of the ground user are obtained, and the angle error is quantized to construct a corresponding communication error model.
[0018] In an embodiment of the application, the deep reinforcement learning model is trained based on the optimization index, including:
[0019] The action network and the evaluation network in the deep reinforcement learning model are initialized;
[0020] The communication error model is preprocessed to reduce the dimension to obtain a channel vector;
[0021] The reduced dimension channel vector is used to generate a corresponding action vector through the action network, and the action vector is mapped and power controlled to obtain a precoding matrix for training;
[0022] The reward value reflecting the performance of the precoding matrix is calculated according to the optimization index and using a multi-objective learning method;
[0023] The channel vector, the action vector, and the reward value are stored in a preset experience area, to train the action network and the evaluation network to obtain a trained deep reinforcement learning model.
[0024] In an embodiment of the application, the action network and the evaluation network are trained to obtain a trained deep reinforcement learning model, including:
[0025] Randomly extract part of the data in the experience area to train the evaluation network;
[0026] Generate evaluation value data of each group of experience data based on the evaluation network;
[0027] Train the action network using the evaluation value data, and update the network weights of the action network to obtain a trained deep reinforcement learning model.
[0028] To achieve the above object, another aspect of the present application provides a multi-satellite multi-beam reliable precoding system, comprising:
[0029] A communication error model construction module is configured to construct a corresponding communication error model based on the established multi-satellite multi-beam communication model;
[0030] An optimization index construction module is configured to construct an optimization index from the perspectives of maximizing the sum rate and minimizing the signal-to-interference-plus-noise ratio (SINR) of each antenna beam, while taking into account the throughput and fairness of the cooperative multi-beam satellite system;
[0031] A precoding matrix determination module is configured to construct a deep reinforcement learning model, train the deep reinforcement learning model based on the optimization index, and perform encoding operation on the error communication model to be precoded according to the trained deep reinforcement learning model, so as to obtain a final precoding matrix.
[0032] The multi-satellite multi-beam reliable precoding method and system can obtain a low-orbit satellite communication precoding scheme under incomplete channel state information conditions, thereby reducing the dependence on the accuracy of real-time channel state information, and can improve the overall throughput of the system while ensuring the fairness of the precoding method to ground users to a certain extent.
[0033] To achieve the above object, the third aspect embodiment of the present application provides a computer device, comprising a processor and a memory, the processor runs a program corresponding to executable program code stored in the memory by reading the executable program code, and the program is implemented when executed by the processor to realize the multi-satellite multi-beam reliable precoding method of the first aspect embodiment.
[0034] To achieve the above object, the fourth aspect embodiment of the present application provides a non-transitory computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the multi-satellite multi-beam reliable precoding method of the first aspect embodiment.
[0035] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0036] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0037] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention;
[0038] Figure 2 This is a flowchart of a multi-satellite multi-beam reliable precoding method according to an embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of a linear antenna array according to an embodiment of the present invention;
[0040] Figure 4 This is a flowchart illustrating the construction of the communication model and error model according to an embodiment of the present invention;
[0041] Figure 5 This is a schematic diagram of the precoding algorithm flow according to an embodiment of the present invention;
[0042] Figure 6 This is a schematic diagram of the deep reinforcement learning network training process according to an embodiment of the present invention;
[0043] Figure 7 This is a schematic diagram of the structure of a multi-satellite multi-beam reliable precoding system according to an embodiment of the present invention;
[0044] Figure 8 It is a computer device according to an embodiment of the present invention. Detailed Implementation
[0045] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0046] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0047] The following description, with reference to the accompanying drawings, outlines a multi-satellite, multi-beam reliable precoding method, system, device, and storage medium according to embodiments of the present invention.
[0048] Figure 1This is an application scenario diagram of a multi-satellite, multi-beam reliable precoding method for low-Earth orbit satellite communication that balances throughput and fairness, as proposed in this invention.
[0049] Figure 2 This is a flowchart of a multi-satellite multi-beam reliable precoding method according to an embodiment of the present invention, such as... Figure 2 As shown, the method includes, but is not limited to, the following steps:
[0050] S1, construct the corresponding communication error model based on the established multi-satellite multi-beam communication model.
[0051] For example, the construction steps of the communication error model of the present invention are as follows: Figure 3 As shown, it includes:
[0052] S101, an array model for a multi-beam antenna is constructed based on the characteristics of a linear uniform antenna array.
[0053] Specifically, consider a scenario where M multi-beam satellites collaboratively provide communication services to K users, with each satellite equipped with a linear uniform antenna array of N antennas. The antenna array model is as follows: Figure 4 As shown. Where, d n Let d be the distance between the nth antenna and its adjacent antennas, where n = 1, 2, 3, ..., N. The antennas are distributed outwards from the center, and the distance d from the nth antenna to the center antenna is... n (N+1-2n) / 2, therefore its transmitted signal beam has a difference of d compared to the central antenna. n (N+1-2n)cos(v k,m The time delay is ) / 2c, where c is the speed of light, and cos(v) k,m Let be the departure angle between the satellite antenna and the user. Further converting the time delay into a phase shift, we can obtain the steering vector between the nth antenna of the mth satellite and the kth user, which is:
[0054]
[0055] S102 is a multi-satellite multi-beam communication model constructed based on the joint transmission method of multiple low-orbit satellites using multi-beam antenna technology.
[0056] Specifically, the downlink transmission model between the satellite system and the k-th user (k = 1, 2, 3, ... K) is as follows:
[0057]
[0058] Among them, s k w is the data symbol required by the k-th user. k For the precoding matrix, h k To define the channel vector between the cooperative multi-beam satellite and the k-th user, The signal is transmitted to the user terminal through the channel after being pre-coded, y k The signal received by the kth user includes the interference signal of other beams in addition to the required signal and noise n k k n is an additive white Gaussian noise, which obeys the distribution The channel between each satellite and the user is Since the distance between the satellite and the user is quite large, in order to simplify the model, only the large-scale fading of the line-of-sight channel is considered, and the free space attenuation model is used. The channel vector between the mth satellite and the kth user is:
[0059]
[0060] S103, based on the multi-satellite multi-beam communication model, the overall phase shift caused by the inter-satellite synchronization delay error, and the angle error in the antenna direction caused by the rotation of the satellite and the positioning error of the ground user are obtained, and the angle error is quantified to construct a corresponding communication error model.
[0061] Specifically, considering that the error has strong randomness, and from the perspective of reducing the complexity of the model, a statistical model method is used for analysis, considering the channel error from two angles. First, the low-orbit satellite has high dynamicity, and the positional relationship between the satellite and the ground user changes rapidly, and the satellite itself also has a certain angle offset during operation, so there is a certain angle error in the beam departure angle between the satellite and the user. The angle error is measured by using a uniformly distributed error variable ε k,m ~ Y(-Δε, Δε), and the error model of the system is:
[0062]
[0063] wherein, is the Hadamard product. v k,m (cos(ε k,m )) is the error vector generated by the angle error through the steering vector, which is:
[0064]
[0065] On the other hand, although the cooperative multi-beam satellites can maintain good time synchronization and delay synchronization in the ideal state, in the actual multi-satellite joint transmission system, since the satellites run on different orbits, the relative speed and relative position between the satellites will change, which will cause the signals between the satellites to be unsynchronized, causing additional phase errors between the actual beams of each satellite in the cooperative multi-beam satellite system. A Gaussian distributed error variable The inter-satellite synchronization error is measured, and the complete error model of the system is:
[0066]
[0067] S2, based on considering the throughput and fairness of the cooperative multi-beam satellite system, respectively from the perspective of maximizing sum rate and minimizing the signal-to-interference-plus-noise ratio of each antenna beam, an optimization index is constructed.
[0068] Specifically, the present application considers the throughput performance optimization problem under the consideration of fairness. For optimizing the throughput of the cooperative multi-beam satellite system, an optimization index is constructed from the perspective of maximizing sum rate. The sum rate is the sum of the data transmission rates of all users in the multi-concurrent transmission system, reflecting the total capacity of the system, and its calculation formula is:
[0069]
[0070] For the measurement of system fairness, since a single antenna beam cannot measure the transmission rate from the sum rate, an optimization index is constructed from the perspective of minimizing the variance of the signal-to-interference-plus-noise ratio of each antenna beam. The signal-to-interference-plus-noise ratio is the ratio of the strength of the useful signal of the communication channel to the strength of the received interference signal plus noise signal, and according to Shannon's law, its size is proportional to the channel capacity. The calculation formula is:
[0071]
[0072] S3, a deep reinforcement learning model is constructed, and the deep reinforcement learning model is trained based on the optimization index, so as to perform encoding operation on the error communication model to be precoded according to the trained deep reinforcement learning model, and obtain the final precoding matrix.
[0073] Specifically, as shown in the pre-coding process, the following steps are included: Figure 5
[0074] S301, initializing the action network and evaluation network in the deep reinforcement learning model.
[0075] S302, preprocessing the communication error model to obtain a channel vector.
[0076] Specifically, according to the distance, position and other information between the satellite and the user, the error channel model under the beam is calculated from S1 step Where t is the current training step. At this time is a complex vector, in order to meet the input requirements of the action network and facilitate subsequent network training, it needs to be preprocessed. First, the is reduced in dimension and expanded into a one-dimensional vector of 1xMNK, and then the vector is expanded into a real vector z of 1x2MNK.t , z t The former half elements are the real parts of the elements of the one-dimensional vector, and the latter half elements are the imaginary parts of the one-dimensional vector, that is:
[0077]
[0078] The Re function and the Im function take the real part and the imaginary part of the complex number respectively.
[0079] S303, the dimension-reduced channel vector is used to generate a corresponding action vector through the action network, and the action vector is subjected to vector mapping and power control to obtain a precoding matrix for training.
[0080] Specifically, the channel vector passes through the action network, and the mean and variance of each element of the action vector are returned. After receiving these parameters, a random number is generated according to the normal distribution as the element value of the action vector a t . At this time it needs to be converted into a precoding matrix vector The latter half elements of a t are taken as the imaginary parts of complex numbers, added in order with the former half elements, and then dimension-adjusted, that is:
[0081]
[0082] The precoding matrix W t ′ obtained by the above formula still needs to be subjected to power control processing to make the power of the finally transmitted signal meet the rated power requirement of the satellite. Assuming that the signal symbol satisfies ||s|| 2 =1, the precoding matrix is subjected to the following power control processing:
[0083]
[0084] wherein, represents the sum of squares of all elements of the original precoding matrix, and P is the rated power of the satellite. This processing can ensure that the total power of the signal symbol after precoding processing is P.
[0085] S304, according to the optimization index, and using a multi-objective learning method, a reward value reflecting the performance of the precoding matrix is calculated.
[0086] Specifically, it is determined whether t is greater than the set maximum training step number T. If t is less than T, according to the optimization index set in S2, a reward value reflecting the performance of the precoding matrix is designed and calculated using a multi-objective learning method.
[0087] The sum rate in S2 is taken as the main task, and the variance of the signal-to-interference-plus-noise ratio is taken as the auxiliary task, and the weighted difference value is the reward value of the multi-objective reinforcement learning, as follows:
[0088]
[0089] wherein R t is the reward value of the current training step, sumrate t is the sum rate at the current training step, T is the total number of training steps, std t and std0 are the variance value at the current training step and the initial variance value respectively, and alpha is the control amount of the auxiliary task weight.
[0090] S305, the channel vector, the action vector and the reward value are stored in the preset experience area to train the action network and the evaluation network to obtain the trained deep reinforcement learning model.
[0091] The embodiment of the application stores the channel vector, the action vector and the reward value in the preset experience area, periodically trains the action network and the evaluation network, and optimizes the performance of the action network. Figure 6 As shown in the figure, the specific process is as follows:
[0092] S3051, part of the data in the experience area is randomly extracted to train the evaluation network.
[0093] Specifically, the evaluation index of the value network on the action strategy is based on the time difference, which includes the reward amount of the current state and the value estimation amount of the next state, and the latter is obtained by evaluating the value of the next state of the current experience item. The value network needs to generate a value estimation for each input action vector, but the network does not understand how to calculate the reward value and the estimated value of the next state, so the design goal of the loss function of the value network is to make its output approach the evaluation index. The value network adopts the gradient descent algorithm for weight updating, and needs to design a corresponding loss function. The loss function is designed as follows:
[0094]
[0095] wherein is the current value network and its weight, and gamma is the future reward control amount, and too large or too small will cause the training of the value network to delay convergence or local optimum.
[0096] S3052, the evaluation value data of each group of experience data is generated based on the evaluation network.
[0097] S3053, training the action network using the evaluation value data, and updating the network weight of the action network to obtain a trained deep reinforcement learning model.
[0098] Specifically, the weight update is also performed by using a gradient descent algorithm, and a loss function needs to be designed. The loss function of the action network is designed to enable the action vector generated by the action network to obtain a greater value evaluation. Specifically, the following is performed:
[0099]
[0100] In order to increase the exploration degree of the algorithm and avoid the phenomenon of overfitting, the maximum entropy is added to the loss function of the action network. The action network directly generates a vector of variables each of which is normally distributed, and therefore the entropy value of the action vector can be obtained. The maximum entropy value can be achieved by minimizing the negative entropy value, and the following is performed:
[0101]
[0102] Wherein, a is an entropy control factor, which is iteratively updated during the training process to keep the entropy approximately constant during the training. Whenever the entropy is higher or lower than the pre-set heuristic target value, the entropy is adjusted. The total loss function of the action network is:
[0103] Λ μ =Λ μ,1 +Λ μ,2
[0104] Thereafter, the steps of (S302)-(S305) are repeated until the maximum number of training steps is reached.
[0105] S306, after the maximum number of training steps is reached, the generated matrix is directly output, which is the precoding matrix under the channel.
[0106] The multi-satellite multi-beam reliable precoding method according to the embodiments of the present application can obtain a precoding scheme for low-orbit satellite communication under incomplete CSI conditions, thereby reducing the dependence on the accuracy of real-time channel state information. On the other hand, the multi-objective learning method also enables the precoding scheme to improve the overall throughput of the system while ensuring the fairness of the ground station to some extent. Compared with other precoding schemes, the technical scheme proposed in the present application has lower computational complexity and can achieve the performance under perfect CSI conditions in the case of large distances between satellites, which is a significant advantage for actual system deployment.
[0107] In order to realize the above-mentioned embodiments, as Figure 7As shown, the embodiment also provides a multi-satellite multi-beam reliable precoding system 10, which comprises a communication error model construction module 100, an optimization index construction module 200, and a precoding matrix determination module 300:
[0108] The communication error model construction module 100 is configured to construct a corresponding communication error model based on the established multi-satellite multi-beam communication model.
[0109] The optimization index construction module 200 is configured to construct an optimization index from the perspectives of maximizing the sum rate and minimizing the signal-to-interference-plus-noise ratio (SINR) of each antenna beam, while taking into account the throughput and fairness of the cooperative multi-beam satellite system.
[0110] The precoding matrix determination module 300 is configured to construct a deep reinforcement learning model and train the deep reinforcement learning model based on the optimization index, so as to perform encoding operation on the error communication model to be precoded according to the trained deep reinforcement learning model, and obtain a final precoding matrix.
[0111] Further, the communication error model construction module 100 is further configured to:
[0112] construct an array model of the multi-beam antenna based on the characteristics of the linear uniform antenna array;
[0113] construct a multi-satellite multi-beam communication model based on the joint transmission mode of multiple low-orbit satellites using multi-beam antenna technology;
[0114] obtain the overall phase shift caused by the inter-satellite synchronization delay error, and the angle error in the antenna direction caused by the rotation of the satellite and the positioning error of the ground user, and quantize the angle error to construct a corresponding communication error model.
[0115] Further, the precoding matrix determination module 300 comprises:
[0116] a model initialization unit configured to initialize the action network and the evaluation network in the deep reinforcement learning model;
[0117] a model preprocessing unit configured to preprocess the communication error model to obtain a channel vector by dimension reduction;
[0118] a mapping control unit configured to generate a corresponding action vector through the action network from the dimension-reduced channel vector, and perform vector mapping and power control on the action vector to obtain a precoding matrix for training;
[0119] a reward value calculation unit configured to calculate a reward value reflecting the performance of the precoding matrix according to the optimization index and by using a multi-objective learning method;
[0120] The network training unit is configured to store the channel vector, the action vector and the reward value into a preset experience area, and train the action network and the evaluation network to obtain a trained deep reinforcement learning model.
[0121] Further, the network training unit is further configured to:
[0122] randomly extract part of the data in the experience area to train the evaluation network;
[0123] generate evaluation value data of each group of experience data based on the evaluation network;
[0124] train the action network by using the evaluation value data, and update the network weight of the action network to obtain the trained deep reinforcement learning model.
[0125] The multi-satellite multi-beam reliable precoding system according to the embodiments of the present application can obtain a precoding scheme for low-orbit satellite communication under incomplete CSI conditions, thereby reducing the dependence on the accuracy of real-time channel state information. On the other hand, the multi-target learning method also enables the precoding scheme to improve the overall throughput of the system while ensuring the fairness of the ground station to a certain extent. Compared with other precoding schemes, the technical scheme proposed in the present application has lower computational complexity and can achieve the performance under perfect CSI conditions in the case of large distances between satellites, which is a significant advantage for actual system deployment.
[0126] In order to implement the method of the above-mentioned embodiments, the present application further provides a computer device, as shown in the accompanying drawings, the computer device 600 comprises a memory 601, a processor 602; wherein the processor 602 runs the program corresponding to the executable program code by reading the executable program code stored in the memory 601, so as to realize each step of the method described above. Figure 8
[0127] In order to implement the above-mentioned embodiments, the present application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method according to the above-mentioned embodiments.
[0128] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, without contradiction.
[0129] In addition, the terms "first", "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly specified.
Claims
1. A multi-satellite multi-beam reliable precoding method, characterized in that, The method comprises the steps of: constructing a corresponding communication error model based on the established multi-satellite multi-beam communication model; constructing an optimization index from the perspectives of maximizing sum rate and minimizing signal-to-interference-plus-noise ratio (SINR) of each antenna beam, in consideration of optimizing throughput and fairness of the cooperative multi-beam satellite system; constructing a deep reinforcement learning model and training the deep reinforcement learning model based on the optimization index, so as to perform coding operation on the communication error model to be precoded according to the trained deep reinforcement learning model, and obtain a final precoding matrix; constructing a corresponding communication error model based on the established multi-satellite multi-beam communication model, comprising: constructing an array model of the multi-beam antenna based on the characteristics of the linear uniform antenna array; constructing a multi-satellite multi-beam communication model based on the joint transmission mode of multiple low-orbit satellites adopting the multi-beam antenna technology; obtaining overall phase shift caused by inter-satellite synchronization time delay error and angle error in the antenna direction caused by rotation of the satellite and positioning error of the ground user, and quantifying the angle error to construct a corresponding communication error model based on the multi-satellite multi-beam communication model; training the deep reinforcement learning model based on the optimization index, comprising: initializing an action network and an evaluation network in the deep reinforcement learning model; performing preprocessing on the communication error model to obtain a channel vector by dimension reduction; generating a corresponding action vector through the action network from the dimension-reduced channel vector, and performing vector mapping and power control on the action vector to obtain a precoding matrix for training; calculating a reward value reflecting performance of the precoding matrix according to the optimization index and by using a multi-objective learning mode; storing the channel vector, the action vector and the reward value into a preset experience area to train the action network and the evaluation network to obtain the trained deep reinforcement learning model; training the action network and the evaluation network to obtain the trained deep reinforcement learning model, comprising: randomly extracting part of data in the experience area to train the evaluation network; generating evaluation value data of each group of experience data based on the evaluation network; training the action network by using the evaluation value data, and updating network weights of the action network to obtain the trained deep reinforcement learning model.
2. A multi-satellite multi-beam reliable precoding system, characterized in that, The method comprises the steps of: a communication error model construction module, configured to construct a corresponding communication error model based on an established multi-satellite multi-beam communication model; an optimization index construction module, configured to construct an optimization index from the perspectives of maximizing sum rate and minimizing signal-to-interference-plus-noise ratio (SINR) of each antenna beam, in consideration of optimizing throughput and fairness of the cooperative multi-beam satellite system; a precoding matrix determination module, configured to construct a deep reinforcement learning model and train the deep reinforcement learning model based on the optimization index, so as to perform coding operation on the communication error model to be precoded according to the trained deep reinforcement learning model, and obtain a final precoding matrix; the communication error model construction module is further configured to: construct an array model of the multi-beam antenna based on the characteristics of the linear uniform antenna array; construct a multi-satellite multi-beam communication model based on the joint transmission mode of multiple low-orbit satellites adopting the multi-beam antenna technology; Obtain an overall phase shift caused by an inter-satellite synchronization time delay error based on the multi-satellite multi-beam communication model, and an angle error in an antenna direction caused by rotation of a satellite and positioning inaccuracy of a ground user, and quantize the angle error to construct a corresponding communication error model; The pre-coding matrix determination module comprises: A model initialization unit configured to initialize an action network and an evaluation network in the deep reinforcement learning model; A model preprocessing unit configured to preprocess the communication error model to obtain a channel vector by dimension reduction; A mapping control unit configured to generate a corresponding action vector by the action network from the channel vector by dimension reduction, and obtain a pre-coding matrix for training by vector mapping and power control on the action vector; A reward value calculation unit configured to calculate a reward value reflecting performance of the pre-coding matrix according to the optimization index and by using a multi-objective learning manner; A network training unit configured to store the channel vector, the action vector, and the reward value in a preset experience area, to train the action network and the evaluation network, and to obtain a trained deep reinforcement learning model; The network training unit is further configured to: randomly extract part of data in the experience area to train the evaluation network; generate evaluation value data of each group of experience data based on the evaluation network; train the action network by using the evaluation value data, and update network weights of the action network to obtain the trained deep reinforcement learning model.
3. A computer device, comprising: comprise a processor and a memory; The processor runs a program corresponding to executable program code stored in the memory by reading the executable program code, to implement the multi-satellite multi-beam reliable pre-coding method of claim 1.
4. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the multi-satellite multi-beam reliable pre-coding method of claim 1.
Citation Information
Patent Citations
Robust precoding method for multi-beam satellite communication system based on machine learning
CN113765553A
RIS-assisted MU-MISO communication system intelligent beam forming method based on deep reinforcement learning
CN118764055A