Multi-satellite multi-beam reliable pre-coding method and system

By building a communication error model and optimization indicators in a low-orbit satellite communication system, and training the precoding matrix with a deep reinforcement learning model, the problem of difficult to take into account both system throughput and fairness is solved, and efficient communication under incomplete channel state information conditions is achieved.

CN120034218AActive Publication Date: 2025-05-23BEIJING UNIV OF POSTS & TELECOMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510044032.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-23
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In low-orbit satellite communication systems, existing linear precoding methods are difficult to obtain perfect channel state information in highly dynamic environments, making it difficult to take into account both system throughput and fairness.

Method used

A reliable precoding method for multi-star and multi-beam is proposed. By building a communication error model and optimization indicators, and combining a deep reinforcement learning model to train the precoding matrix, it optimizes system throughput and fairness.

Benefits of technology

Under the condition of incomplete channel state information, the overall throughput of the low-orbit satellite communication system is improved, and the fairness of ground users is guaranteed to a certain extent, reducing the dependence on the accuracy of real-time channel state information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034218A_ABST
    Figure CN120034218A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-satellite multi-beam reliable pre-coding method and a multi-satellite multi-beam reliable pre-coding system. The multi-satellite multi-beam reliable pre-coding method comprises the following steps: constructing a corresponding communication error model based on an established multi-satellite multi-beam communication model; based on consideration of throughput and fairness of optimization of the cooperative multi-beam satellite system, optimization indexes are constructed from the angles of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam; and constructing a deep reinforcement learning model, and training the deep reinforcement learning model based on the optimization index, so as to carry out coding operation on the error communication model to be pre-coded according to the trained deep reinforcement learning model, and obtain a final pre-coding matrix. According to the invention, a pre-coding scheme of low-orbit satellite communication can be obtained under the condition of incomplete channel state information, so that the dependence on the accuracy of real-time channel state information is reduced; and the fairness of the pre-coding method to ground users can be ensured to a certain extent while the overall throughput of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of low-orbit satellite communications, and in particular to a multi-satellite multi-beam reliable precoding method and system. Background Art

[0002] The low-orbit satellite communication system is an important part of the fifth-generation non-terrestrial network. It has great advantages when the terrestrial network cannot provide stable and efficient communication services due to natural disasters, economic conditions and other factors. In recent years, the rapid development of low-orbit satellite communications has led to an increasing shortage of spectrum resources. Precoding technology, as a key technology that can effectively improve system capacity and resource utilization, has been widely studied and applied in academia and industry.

[0003] However, the performance of directly applying linear precoding to low-orbit satellite communication systems is poor. Precoding is generally divided into two categories: linear precoding and nonlinear precoding. Nonlinear precoding has not been widely used due to its high complexity; although linear precoding is relatively simple, it has high requirements for the accuracy of channel state information (CSI), and the high dynamics of low-orbit satellites make it very difficult to obtain perfect CSI. In addition, most of the existing research on precoding methods for low-orbit satellite systems only focuses on the overall performance of the system, and these schemes will result in low throughput in some beam coverage areas. Therefore, how to design a reliable precoding method that can take into account the total throughput and fairness of the system and is robust to imperfect CSI based on the characteristics of low-orbit satellite communication systems has become a research hotspot in low-orbit satellite precoding technology. Summary of the invention

[0004] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0005] To this end, the present invention proposes a multi-satellite multi-beam reliable precoding method to maximize the system throughput while taking into account the coverage area of ​​each beam.

[0006] Another object of the present invention is to provide a multi-satellite multi-beam reliable precoding system.

[0007] A third object of the present invention is to provide a computer device.

[0008] A fourth object of the present invention is to provide a non-transitory computer-readable storage medium.

[0009] To achieve the above object, the present invention provides a multi-satellite multi-beam reliable precoding method, comprising:

[0010] Based on the established multi-satellite multi-beam communication model, a corresponding communication error model is constructed;

[0011] Based on the optimization of the throughput and fairness of the cooperative multi-beam satellite system, the optimization index is constructed from the perspectives of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam.

[0012] A deep reinforcement learning model is constructed, and the deep reinforcement learning model is trained based on the optimization index, so as to perform encoding operation on the error communication model to be precoded according to the trained deep reinforcement learning model to obtain a final precoding matrix.

[0013] The multi-satellite multi-beam reliable precoding method of the embodiment of the present invention may also have the following additional technical features:

[0014] In one embodiment of the present invention, a corresponding communication error model is constructed based on the established multi-satellite multi-beam communication model, including:

[0015] Construct an array model of multi-beam antenna based on the characteristics of linear uniform antenna array;

[0016] A multi-satellite multi-beam communication model is constructed based on the joint transmission of multiple low-orbit satellites using multi-beam antenna technology;

[0017] Based on the multi-satellite multi-beam communication model, the overall phase shift caused by the inter-satellite synchronization delay error and the angular error in the antenna direction caused by the rotation of the satellite and the positioning misalignment of the ground user are obtained, and the angular error is quantified to construct a corresponding communication error model.

[0018] In one embodiment of the present invention, training a deep reinforcement learning model based on an optimization index includes:

[0019] Initialize the action network and evaluation network in the deep reinforcement learning model;

[0020] Preprocessing the communication error model to reduce the dimension to obtain a channel vector;

[0021] The reduced-dimensional channel vector is passed through an action network to generate a corresponding action vector, and the action vector is vector-mapped and power-controlled to obtain a precoding matrix for training;

[0022] According to the optimization index, a reward value reflecting the performance of the precoding matrix is ​​calculated by using a multi-objective learning method;

[0023] The channel vector, action vector, and reward value are stored in a preset experience area to train the action network and the evaluation network to obtain a trained deep reinforcement learning model.

[0024] In one embodiment of the present invention, the action network and the evaluation network are trained to obtain a trained deep reinforcement learning model, including:

[0025] Randomly extract some data from the experience area to train the evaluation network;

[0026] Generate evaluation value data for each group of experience data based on the evaluation network;

[0027] The action network is trained using the evaluation value data, and the network weights of the action network are updated to obtain a trained deep reinforcement learning model.

[0028] To achieve the above object, the present invention provides a multi-satellite multi-beam reliable precoding system, comprising:

[0029] A communication error model building module is used to build a corresponding communication error model based on the established multi-satellite multi-beam communication model;

[0030] An optimization index building module is used to build optimization indicators from the perspectives of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam based on optimizing the throughput and fairness of the cooperative multi-beam satellite system;

[0031] The precoding matrix determination module is used to construct a deep reinforcement learning model and train the deep reinforcement learning model based on the optimization index, so as to perform encoding operations on the error communication model to be precoded according to the trained deep reinforcement learning model to obtain the final precoding matrix.

[0032] The multi-satellite multi-beam reliable precoding method and system of the embodiments of the present invention can allow a precoding scheme for low-orbit satellite communications to be obtained under conditions of incomplete channel state information, thereby reducing dependence on the accuracy of real-time channel state information; and can improve the overall system throughput while ensuring the fairness of the precoding method to ground users to a certain extent.

[0033] To achieve the above-mentioned purpose, the third aspect embodiment of the present application proposes a computer device, including a processor and a memory, wherein the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, and when the program is executed by the processor, the multi-star multi-beam reliable precoding method as described in the first aspect embodiment is implemented.

[0034] To achieve the above-mentioned objectives, the fourth aspect embodiment of the present application proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the multi-star multi-beam reliable precoding method as described in the first aspect embodiment is implemented.

[0035] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0037] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present invention;

[0038] Figure 2 is a flowchart of a multi-satellite multi-beam reliable precoding method according to an embodiment of the present invention;

[0039] Figure 3 is a schematic diagram of a linear antenna array according to an embodiment of the present invention;

[0040] Figure 4 is a flow chart of constructing a communication model and an error model according to an embodiment of the present invention;

[0041] Figure 5 is a schematic diagram of a precoding algorithm flow chart according to an embodiment of the present invention;

[0042] Figure 6 is a schematic diagram of a deep reinforcement learning network training process according to an embodiment of the present invention;

[0043] Figure 7 is a schematic structural diagram of a multi-satellite multi-beam reliable precoding system according to an embodiment of the present invention;

[0044] Figure 8 It is a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0046] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0047] The following describes a multi-satellite multi-beam reliable precoding method, system, device, and storage medium according to embodiments of the present invention with reference to the accompanying drawings.

[0048] Figure 1This is an application scenario diagram of a multi-satellite multi-beam reliable precoding method in low-orbit satellite communications that takes into account both throughput and fairness, as proposed by the present invention.

[0049] Figure 2 is a flow chart of a multi-satellite multi-beam reliable precoding method according to an embodiment of the present invention. Figure 2 As shown, the method includes but is not limited to the following steps:

[0050] S1, build the corresponding communication error model based on the established multi-satellite multi-beam communication model.

[0051] Exemplarily, the steps of constructing the communication error model of the present invention are as follows: Figure 3 As shown, including:

[0052] S101, constructing an array model of a multi-beam antenna based on the characteristics of a linear uniform antenna array.

[0053] Specifically, consider a total of M multi-beam satellites cooperating to provide communication services for K users, and each satellite is equipped with a linear uniform antenna array with N antennas. The antenna array model is as follows: Figure 4 As shown. Among them, d n is the distance between the nth antenna and the adjacent antenna, n = 1, 2, 3, ..., N. The antennas are distributed from the center to both sides, and the distance between the nth antenna and the center antenna is d n (N+1-2n) / 2, so the signal beam it transmits has d compared to the center antenna. n (N+1-2n)cos(v k,m ) / 2c, where c is the speed of light, cos(v k,m ) is the departure angle between the satellite antenna and the user. By further converting the time delay into a phase shift, we can get the steering vector between the nth antenna of the mth satellite and the kth user, which is:

[0054]

[0055] S102, constructing a multi-satellite multi-beam communication model based on the joint transmission of multiple low-orbit satellites using multi-beam antenna technology.

[0056] Specifically, the downlink transmission model between the satellite system and the kth user (k=1, 2, 3, ... K) is:

[0057]

[0058] Among them, s k is the data symbol required by the kth user, w k is the precoding matrix, h k is the channel vector between the collaborative multi-beam satellite and the kth user, After precoding, the signal is transmitted to the user end through the channel. k is the signal received by the kth user, which includes the interference signals from other beams in addition to the desired signal and noise n k .n k is additive Gaussian white noise, which follows the distribution The channel between each satellite and the user is Since the distance between the satellite and the user is quite large, in order to simplify the model, only the large-scale fading of the line-of-sight channel is considered here, and the free space attenuation model is adopted. The channel vector between the mth satellite and the kth user is:

[0059]

[0060] S103, based on the multi-satellite multi-beam communication model, obtain the overall phase shift caused by the inter-satellite synchronization delay error, as well as the angle error in the antenna direction caused by the rotation of the satellite and the positioning inaccuracy of the ground user, and quantify the angle error to construct a corresponding communication error model.

[0061] Specifically, considering that the error has a strong random nature, and from the perspective of reducing the complexity of the model, the statistical model method is used for analysis, and the channel error is considered from two perspectives. First, the low-orbit satellite is highly dynamic, and its position relationship with the ground user changes rapidly. In addition, the satellite itself has a certain angle offset during operation, so there is a certain angle error in the beam departure angle between the satellite and the user. Using a uniformly distributed error variable ε k,m ~Υ(-Δε,Δε) measures the angle error, then the error model of the system is:

[0062]

[0063] in, is the Hadamard product. k,m (cos(ε k,m )) is the error vector generated by the angle error through the steering vector, which is:

[0064]

[0065] On the other hand, although cooperative multi-beam satellites can maintain good time synchronization and delay synchronization in an ideal state, in an actual multi-satellite joint transmission system, since the satellites are running in different orbits, the relative speed and relative position between the satellites will change, which will cause the signals between the satellites to be out of sync, causing additional phase errors between the actual beams of each satellite in the cooperative multi-beam satellite system. Error variables using Gaussian distribution To measure the inter-satellite synchronization error, the complete error model of the system is:

[0066]

[0067] S2, based on optimizing both the throughput and fairness of the cooperative multi-beam satellite system, constructs optimization indicators from the perspectives of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam.

[0068] Specifically, the present invention considers the problem of optimizing throughput performance under fairness. For optimizing the throughput of the cooperative multi-beam satellite system, an optimization index is constructed from the perspective of maximizing the sum rate. The sum rate is the sum of the data transmission rates of all users in multiple concurrent transmission systems, reflecting the total capacity of the system, and its calculation formula is:

[0069]

[0070] As for the measurement of system fairness, since a single antenna beam cannot measure the transmission rate from the sum rate, an optimization index is constructed from the perspective of minimizing the variance of the signal to interference plus noise ratio of each antenna beam. The signal to interference plus noise ratio is the ratio of the strength of the useful signal in the communication channel to the strength of the received interference signal plus noise signal. According to Shannon's law, its size is proportional to the channel capacity. The calculation formula is:

[0071]

[0072] S3, constructing a deep reinforcement learning model, and training the deep reinforcement learning model based on the optimization index, so as to perform encoding operation on the error communication model to be precoded according to the trained deep reinforcement learning model to obtain a final precoding matrix.

[0073] Specifically, Figure 5 The precoding process shown includes the following steps:

[0074] S301, initialize the action network and evaluation network in the deep reinforcement learning model.

[0075] S302, preprocessing the communication error model to reduce the dimension to obtain a channel vector.

[0076] Specifically, according to the distance and position between the satellite and the user, the error channel model under the beam is calculated by step S1: Where t is the current training step. is a complex vector. In order to meet the input requirements of the action network and facilitate subsequent network training, it needs to be preprocessed. Perform dimensionality reduction processing, expand it into a one-dimensional vector of 1×MNK, and then expand the vector into a real vector z of 1×2MNKt , z t The first half of the elements are the real part of each element of the one-dimensional vector, and the second half of the elements are the imaginary part of the one-dimensional vector, that is:

[0077]

[0078] The Re function and the Im function take the real part and the imaginary part of the complex number respectively.

[0079] S303, generating a corresponding action vector from the dimension-reduced channel vector through an action network, and performing vector mapping and power control on the action vector to obtain a precoding matrix for training.

[0080] Specifically, the channel vector passes through the action network and returns the mean and variance of each element of the action vector. After receiving these parameters, a random number is generated according to the normal distribution as the action vector a t The element value of . It needs to be converted into a precoding matrix vector Using the reverse steps of the input channel state vector preprocessing, a t The second half of the elements are added to the first half of the elements in order as the imaginary part of the complex number, and then the dimension is adjusted. The specific steps are:

[0081]

[0082] The precoding matrix W obtained from the above formula is t ′Power control processing is also required to make the final transmitted signal power meet the rated power requirements of the satellite. Assume that the signal symbol satisfies ||s|| 2 =1, the following power control processing is performed on the precoding matrix:

[0083]

[0084] in, represents the sum of the squares of all elements of the original precoding matrix, and P is the rated power of the satellite. This process ensures that the total power of the signal symbol after precoding is P.

[0085] S304: Calculate a reward value reflecting the performance of the precoding matrix according to the optimization index and by using a multi-objective learning method.

[0086] Specifically, it is determined whether t is greater than the set maximum number of training steps T. If t is less than T, a reward value reflecting the performance of the precoding matrix is ​​designed and calculated using a multi-objective learning method according to the optimization index set in S2.

[0087] The sum rate in S2 is used as the main task, and the variance of the signal to interference plus noise ratio is used as the auxiliary task. The weighted difference is the reward value of multi-objective reinforcement learning, as follows:

[0088]

[0089] Where R t That is, the reward value of the current training step, sumrate t is the sum rate size under the current training step, T is the total number of training steps, std t and std 0 are the variance value and initial variance value in the current training step, respectively, and α is the control value of the auxiliary task weight. It can be seen that the weighted value of the sum rate approaches 1 as the number of training steps increases; the weighted value of the variance gradually decreases with its size.

[0090] S305, storing the channel vector, the action vector, and the reward value in a preset experience area to train the action network and the evaluation network to obtain a trained deep reinforcement learning model.

[0091] The embodiment of the present invention stores the channel vector, action vector, and reward value in a preset experience area, periodically trains the action network and the evaluation network, and optimizes the performance of the action network. Figure 6 As shown, the specific process is as follows:

[0092] S3051, randomly extracting part of the data in the experience area to train the evaluation network.

[0093] Specifically, the evaluation index of the action strategy of the value network is based on temporal difference, which includes the reward amount of the current state and the estimated value of the next state. The latter is obtained by evaluating the value of the next state of the current experience item. The value network needs to generate such a value estimate for each input action vector. However, the network does not know how to calculate the reward value and the estimated value of the next state. Therefore, the loss function design goal of the value network is to make its output close to this evaluation index. The value network uses the gradient descent algorithm to update the weights, and the corresponding loss function needs to be designed. Its loss function is designed as follows:

[0094]

[0095] in is the current value network and its weights, γ is the future reward control amount, and its excessive size or smallness will lead to delayed convergence or local optimality of the value network training.

[0096] S3052, generating evaluation value data for each group of experience data based on the evaluation network.

[0097] S3053, using the evaluation value data to train the action network, and updating the network weights of the action network to obtain a trained deep reinforcement learning model.

[0098] Specifically, the gradient descent algorithm is also used to update the weights, and a loss function needs to be designed. The loss function design of the action network is to make the action vectors it generates get a greater value evaluation. The details are as follows:

[0099]

[0100] In order to increase the exploration degree of the algorithm and avoid overfitting, the maximum entropy is added to the loss function of the action network. The action network directly generates a vector of variables whose elements are normally distributed, so the entropy value of the action vector can be obtained. Minimizing the negative entropy value can maximize the entropy value, as follows:

[0101]

[0102] Where α is an entropy control factor that is iteratively updated during training to keep the entropy roughly constant during training, and is adjusted whenever the entropy is above or below a pre-set heuristic target value. The total loss function of the action network is:

[0103] Λ μ =Λ μ,1 +Λ μ,2

[0104] Thereafter, steps (S302)-(S305) are repeated until the maximum number of training steps is reached.

[0105] S306, after reaching the maximum number of training steps, directly output the generated matrix, which is the precoding matrix under the channel.

[0106] The multi-satellite multi-beam reliable precoding method according to the embodiment of the present invention can allow the precoding scheme of low-orbit satellite communication to be obtained under incomplete CSI conditions, thereby reducing the dependence on the accuracy of real-time channel state information; on the other hand, the use of a multi-objective learning method also enables the precoding scheme to improve the overall system throughput while ensuring the fairness of ground stations to a certain extent. Compared with other precoding schemes, the technical solution proposed in this application has lower computational complexity and can achieve performance under perfect CSI conditions when there is a large distance between satellites, which is a significant advantage for actual system deployment.

[0107] In order to implement the above embodiment, Figure 7As shown, this embodiment also provides a multi-satellite multi-beam reliable precoding system 10, which includes a communication error model construction module 100, an optimization index construction module 200, and a precoding matrix determination module 300:

[0108] A communication error model building module 100 is used to build a corresponding communication error model based on the established multi-satellite multi-beam communication model;

[0109] The optimization index building module 200 is used to build the optimization index from the perspectives of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam based on optimizing the throughput and fairness of the cooperative multi-beam satellite system;

[0110] The precoding matrix determination module 300 is used to construct a deep reinforcement learning model and train the deep reinforcement learning model based on the optimization index, so as to perform encoding operations on the error communication model to be precoded according to the trained deep reinforcement learning model to obtain the final precoding matrix.

[0111] Furthermore, the communication error model building module 100 is also used for:

[0112] Construct an array model of multi-beam antenna based on the characteristics of linear uniform antenna array;

[0113] A multi-satellite multi-beam communication model is constructed based on the joint transmission of multiple low-orbit satellites using multi-beam antenna technology;

[0114] Based on the multi-satellite multi-beam communication model, the overall phase shift caused by the inter-satellite synchronization delay error and the angular error in the antenna direction caused by the rotation of the satellite and the positioning misalignment of the ground user are obtained, and the angular error is quantified to construct a corresponding communication error model.

[0115] Further, the precoding matrix determination module 300 includes:

[0116] Model initialization unit, used to initialize the action network and evaluation network in the deep reinforcement learning model;

[0117] A model preprocessing unit, used for preprocessing the communication error model to reduce the dimension to obtain a channel vector;

[0118] A mapping control unit, used to generate a corresponding action vector from the dimension-reduced channel vector through an action network, and perform vector mapping and power control on the action vector to obtain a precoding matrix for training;

[0119] A reward value calculation unit, used to calculate a reward value reflecting the performance of the precoding matrix according to the optimization index and by using a multi-objective learning method;

[0120] The network training unit is used to store the channel vector, action vector, and reward value in a preset experience area to train the action network and the evaluation network to obtain a trained deep reinforcement learning model.

[0121] Furthermore, the network training unit is also used for:

[0122] Randomly extract some data from the experience area to train the evaluation network;

[0123] Generate evaluation value data for each group of experience data based on the evaluation network;

[0124] The action network is trained using the evaluation value data, and the network weights of the action network are updated to obtain a trained deep reinforcement learning model.

[0125] The multi-satellite multi-beam reliable precoding system according to the embodiment of the present invention can allow the precoding scheme of low-orbit satellite communication to be obtained under incomplete CSI conditions, thereby reducing the dependence on the accuracy of real-time channel state information; on the other hand, the use of a multi-objective learning method also enables the precoding scheme to improve the overall system throughput while ensuring the fairness of ground stations to a certain extent. Compared with other precoding schemes, the technical solution proposed in this application has lower computational complexity and can achieve performance under perfect CSI conditions when there is a large distance between satellites, which is a significant advantage for actual system deployment.

[0126] In order to implement the method of the above embodiment, the present invention also provides a computer device, such as Figure 8 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein the processor 602 runs a program corresponding to the executable program code by reading the executable program code stored in the memory 601, so as to implement each step of the method described above.

[0127] In order to implement the above embodiments, the present application also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the above embodiments is implemented.

[0128] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0129] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

Claims

1. A multi-satellite multi-beam reliable precoding method, characterized in that: include: Based on the established multi-satellite multi-beam communication model, a corresponding communication error model is constructed; Based on the optimization of the throughput and fairness of the cooperative multi-beam satellite system, the optimization index is constructed from the perspectives of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam. A deep reinforcement learning model is constructed, and the deep reinforcement learning model is trained based on the optimization index, so as to perform encoding operations on the error communication model to be precoded according to the trained deep reinforcement learning model to obtain a final precoding matrix.

2. The method according to claim 1, characterized in that Based on the established multi-satellite multi-beam communication model, the corresponding communication error model is constructed, including: Construct an array model of multi-beam antenna based on the characteristics of linear uniform antenna array; A multi-satellite multi-beam communication model is constructed based on the joint transmission of multiple low-orbit satellites using multi-beam antenna technology; Based on the multi-satellite multi-beam communication model, the overall phase shift caused by the inter-satellite synchronization delay error and the angular error in the antenna direction caused by the rotation of the satellite and the positioning misalignment of the ground user are obtained, and the angular error is quantified to construct a corresponding communication error model.

3. The method according to claim 1, characterized in that Training deep reinforcement learning models based on optimization metrics, including: Initialize the action network and evaluation network in the deep reinforcement learning model; Preprocessing the communication error model to reduce the dimension to obtain a channel vector; The reduced-dimensional channel vector is passed through an action network to generate a corresponding action vector, and the action vector is vector-mapped and power-controlled to obtain a precoding matrix for training; According to the optimization index, a reward value reflecting the performance of the precoding matrix is ​​calculated by using a multi-objective learning method; The channel vector, action vector, and reward value are stored in a preset experience area to train the action network and the evaluation network to obtain a trained deep reinforcement learning model.

4. The method according to claim 3, characterized in that The action network and evaluation network are trained to obtain a trained deep reinforcement learning model, including: Randomly extract some data from the experience area to train the evaluation network; Generate evaluation value data for each group of experience data based on the evaluation network; The action network is trained using the evaluation value data, and the network weights of the action network are updated to obtain a trained deep reinforcement learning model.

5. A multi-satellite multi-beam reliable precoding system, characterized in that: include: A communication error model building module is used to build a corresponding communication error model based on the established multi-satellite multi-beam communication model; An optimization index building module is used to build optimization indicators from the perspectives of maximizing the sum rate and minimizing the signal to interference plus noise ratio of each antenna beam based on optimizing the throughput and fairness of the cooperative multi-beam satellite system; The precoding matrix determination module is used to construct a deep reinforcement learning model and train the deep reinforcement learning model based on the optimization index, so as to perform encoding operations on the error communication model to be precoded according to the trained deep reinforcement learning model to obtain the final precoding matrix.

6. The system according to claim 5, characterized in that The communication error model building block is also used to: Construct an array model of multi-beam antenna based on the characteristics of linear uniform antenna array; A multi-satellite multi-beam communication model is constructed based on the joint transmission of multiple low-orbit satellites using multi-beam antenna technology; Based on the multi-satellite multi-beam communication model, the overall phase shift caused by the inter-satellite synchronization delay error and the angular error in the antenna direction caused by the rotation of the satellite and the positioning misalignment of the ground user are obtained, and the angular error is quantified to construct a corresponding communication error model.

7. The system according to claim 5, characterized in that The precoding matrix determination module includes: Model initialization unit, used to initialize the action network and evaluation network in the deep reinforcement learning model; A model preprocessing unit, used for preprocessing the communication error model to reduce the dimension to obtain a channel vector; A mapping control unit, used to generate a corresponding action vector from the dimension-reduced channel vector through an action network, and perform vector mapping and power control on the action vector to obtain a precoding matrix for training; A reward value calculation unit, used to calculate a reward value reflecting the performance of the precoding matrix according to the optimization index and by using a multi-objective learning method; The network training unit is used to store the channel vector, action vector, and reward value in a preset experience area to train the action network and the evaluation network to obtain a trained deep reinforcement learning model.

8. The system according to claim 7, characterized in that The network training unit is also used to: Randomly extract some data from the experience area to train the evaluation network; Generate evaluation value data for each group of experience data based on the evaluation network; The action network is trained using the evaluation value data, and the network weights of the action network are updated to obtain a trained deep reinforcement learning model.

9. A computer device, characterized in that: including a processor and a memory; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the multi-satellite multi-beam reliable precoding method as described in any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multi-satellite multi-beam reliable precoding method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Robust precoding method for multi-beam satellite communication system based on machine learning

    CN113765553A

  • NOMA multi-beam satellite communication method based on deep reinforcement learning

    CN118075877A

  • Distributed RIS auxiliary millimeter wave communication system beam forming method and system based on deep reinforcement learning

    CN118539959A

  • RIS-assisted MU-MISO communication system intelligent beam forming method based on deep reinforcement learning

    CN118764055A

  • KR20240043924A