Satellite-to-ground data transmission method, device, equipment and medium
Through the LSTM network, predicting the future spatio-temporal data of satellites and combining the resource allocation model of deep reinforcement learning algorithms, the problem of inaccurate resource allocation in the existing technology in complex spatio-temporal environments is solved, and efficient and low-energy-consuming satellite-earth data transmission is achieved.
Patent Information
- Application Number
- CN202510197305.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing dynamic resource allocation technology based on channel state information (CSI) is difficult to achieve high-precision resource allocation in complex space-time environments, especially when satellites and ground stations move at high speed or change in locations significantly, resource allocation is inaccurate, resulting in poor communication efficiency and energy management effects.
The LSTM network is used to predict the future spatiotemporal data of satellites and based on the resource allocation model trained by the deep reinforcement learning algorithm, a dynamically adjusted resource allocation scheme is generated to meet the control constraints of communication delay, bandwidth resources and transmission power.
By improving the prediction accuracy of future spatio-temporal data, optimal resource allocation in complex environments can be achieved, energy consumption is reduced, transmission efficiency is improved, and system adaptability and robustness are enhanced.
Smart Images

Figure CN120017140A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data transmission, and in particular to a satellite-to-earth data transmission method, device, equipment and medium. Background Art
[0002] The basic principle of satellite-ground collaboration technology is based on the collaborative work between satellites and ground stations to achieve efficient data transmission, resource sharing and task execution. The key to satellite-ground collaboration is to ensure that satellites and ground stations can exchange data in real time and optimize the overall performance of the system through reasonably designed communication protocols and efficient collaboration mechanisms. Satellites need to transmit the collected data to the ground station or receive instructions from the ground station to perform specific tasks while in orbit. This process requires high reliability and low latency in data transmission. Satellite-ground collaboration technology ensures the real-time and integrity of data through a variety of transmission links (such as X-band, Ka-band) and multi-band communication technologies. Satellite-ground collaboration allows for the reasonable allocation and optimization of resources between ground stations and satellites. For example, ground stations can provide real-time computing power support for satellites, reduce the computing burden of satellites, and improve the efficiency of task execution. In this way, satellites can complete more complex tasks without increasing the burden on hardware.
[0003] In related technologies, there is a dynamic resource allocation technology based on channel state information (CSI), which can dynamically adjust resource allocation strategies by acquiring key parameters in the communication channel (such as channel gain, noise level, etc.) in real time. In this technology, the system uses CSI to monitor the channel conditions between the satellite and the ground station in real time, analyzes the instantaneous state of the channel, and dynamically allocates bandwidth resources and transmission power according to environmental changes to optimize transmission performance.
[0004] For example, when the channel conditions are good, the system will reduce the transmission power to save energy, or increase the transmission rate when there is sufficient bandwidth resources; when the channel conditions are poor, the system will increase the transmission power or adjust the bandwidth resources to ensure the stability of the communication link. CSI dynamic resource allocation technology is suitable for satellite communication scenarios with large fluctuations in the communication environment. The core of its design is to adjust resources in real time according to channel feedback information to achieve dynamic optimization of communication efficiency.
[0005] This dynamic resource allocation technology based on channel state information (CSI) can optimize bandwidth resources and transmission power usage in satellite communications to a certain extent, but the technology has limitations, especially in complex space-time environments.
[0006] First, CSI technology mainly relies on single channel feedback information such as channel gain and noise level, which makes it difficult to fully capture the dynamic relative position between satellites and ground stations, distance changes, and the impact of orbital factors. Therefore, when satellites and ground stations move at high speed or their positions change significantly, this technology may not be able to quickly adapt to fluctuations in channel characteristics, resulting in inaccurate resource allocation.
[0007] Secondly, CSI technology lacks high-precision perception and prediction methods for multi-dimensional spatiotemporal characteristics, and fails to integrate spatiotemporal information from multi-source data such as GPS, inertial measurement (IMU) and other sensors, which limits the system's adaptability in complex dynamic environments.
[0008] Furthermore, since no deep reinforcement learning or adaptive optimization algorithms are used, CSI technology lacks the intelligence to deal with sudden environmental changes and cannot flexibly optimize resource allocation strategies. Therefore, although CSI technology performs well in static or less variable channel environments, its effectiveness is limited in satellite communication scenarios that require high precision and dynamic response, and it is difficult to achieve optimal transmission efficiency and energy management effects. Summary of the invention
[0009] In view of this, embodiments of the present application provide a satellite-to-earth data transmission method, apparatus, device, and medium to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.
[0010] A first aspect of an embodiment of the present application provides a satellite-to-ground data transmission method, the method comprising: Collecting historical spatiotemporal data of the satellite, the historical spatiotemporal data including: historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite; According to the historical spatiotemporal data, the future spatiotemporal data of the satellite is predicted using an LSTM network; Generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmission power allocated at the current moment, and the noise transmission power spectrum density of the current position of the satellite; The future state vector is input into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward and next state as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite under the state; When the resource allocation scheme satisfies the constraint conditions, the resource allocation scheme is applied to the future time of the satellite, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
[0011] Optionally, the method further comprises: When the resource allocation scheme does not satisfy the constraint condition, readjusting the resource allocation scheme; The communication delay is constrained as follows: the total amount of data transmitted by the satellite Must be within the specified time window The communication delay constraint expression is: ; The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources The constraint expression of the bandwidth resource is: ; The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range The constraint expression of the transmit power control is: ; in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.
[0012] Optionally, before training the resource allocation model, the method further includes: Acquire the real time and space data of each position of the satellite in orbit; the real time and space data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data; According to the real antenna direction angle data, the phase change of the channel is obtained ; Calculate the dynamic distance using the real GPS location information ; ; in, , , are the three-dimensional coordinates of the satellites respectively; , , are the three-dimensional coordinates of the ground stations respectively; According to the dynamic distance , fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ; According to the total amount of data transmitted and the transmission rate R(t), based on the formula , determine the remaining amount of data to be transmitted ; According to the time window Total time and the current time t, based on the formula , determine the remaining time of the time window ; According to the real GPS location information, determine the current location noise emission power spectrum density ; According to the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) is obtained through the multidimensional vector It means that the multiple state vectors s(t) constitute a state space.
[0013] Optionally, before training the resource allocation model, the method further includes: The transmit power is discretely enumerated at a first interval to obtain a plurality of discrete transmit powers ; Discretely enumerate the allocatable bandwidth resource range at a second interval to obtain multiple discrete bandwidth resources ; The discrete transmission power and the discrete bandwidth resources Combine them to get the action space; The action space includes: multiple discrete transmission powers and multiple discrete bandwidth resources Composed of multiple actions .
[0014] Optionally, before training the resource allocation model, the method further includes: Set a reward function, the reward function is: ; ; Wherein, W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite; The derivation process of the influencing factors of energy consumption W(t,s) is as follows: According to the transmission power Determine W(t,s), the expression is: ; in, is the time window, is the transmission power, t is the current time; According to the communication time, the formula From calculation, we know that D is the amount of data transmitted, C is the channel capacity, and ; Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained; in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel; N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource; Determine .
[0015] Optionally, before training the resource allocation model, the method further includes: Define a deep reinforcement learning network structure of the resource allocation model, wherein the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ; Construct an experience replay pool, which is used to store multiple training samples, each training sample is a tuple , the tuple includes the satellite's state s, action a, reward r, and next state The training sample is determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.
[0016] Optionally, the resource allocation model training process is as follows: Extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained by using the small batch training samples; at each time step, use the epsilon-greedy strategy to select actions in the action space: An action is randomly selected from the action space with probability The probability of choosing the current action from the action space The action with the largest value ; Setting the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a; in, ; ; Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite; For each training sample, the target value y is calculated, which is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is: ; in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when The target value y is taken as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, the iterative update process is: ; β is the learning rate, which means Q(s , a) Update step size; During the training process, the mean square error is used as the loss function, and the calculation formula of the loss function is: ; Where N represents the number of small batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Next, take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Next, take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.
[0017] A second aspect of an embodiment of the present application provides a satellite-to-ground data transmission device, the device comprising: A collection module is used to collect the historical spatiotemporal data of the satellite, wherein the historical spatiotemporal data includes: the historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite; A prediction module, used to predict the future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; A future state vector generation module, used to generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmission power allocated at the current moment and the noise transmission power spectrum density of the current position of the satellite; An execution module is used to input the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the historical state, action, reward and next state of the satellite as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite under the state; The application module is used to apply the resource allocation scheme to the future time of the satellite when the resource allocation scheme meets the constraint conditions, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
[0018] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method described in the first aspect.
[0019] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0020] Beneficial effects of this application: The present application provides a satellite-to-ground data transmission method, device, equipment and medium, the method comprising: collecting historical spatiotemporal data of a satellite, the historical spatiotemporal data comprising: historical GPS position information, IMU acceleration data and antenna azimuth data of the satellite; predicting future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; generating a future state vector based on the future spatiotemporal data; the future state vector comprising: the remaining time of a specified time window, the remaining data volume of a data packet, a channel gain, a bandwidth resource allocated at the current moment, a transmission power allocated at the current moment and a noise transmission power spectrum density of the current position of the satellite; inputting the future state vector into a pre-trained resource allocation model to obtain The resource allocation model outputs a resource allocation plan; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward and next state as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite under the state; when the resource allocation plan meets the constraint conditions, the resource allocation plan is applied to the future moment of the satellite, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
[0021] Through the technical solution of the present application, the LSTM network can be used to analyze and predict the historical spatiotemporal data of the satellite, which can significantly improve the prediction accuracy of future spatiotemporal data. The resource allocation model obtained based on deep reinforcement learning training can dynamically adjust the resource allocation plan according to the future state vector, and achieve optimal resource allocation under the constraints of communication delay, bandwidth resources and transmission power control. The resource allocation model trained by the deep reinforcement learning algorithm can minimize energy consumption while ensuring data transmission tasks, improve transmission efficiency, adapt to complex spatiotemporal environment changes, and improve the adaptability and robustness of the system in the face of dynamic changes, providing strong technical support for the development of communications between satellites and the ground. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application.
[0023] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0024] Figure 1 It is a flowchart of a satellite-to-ground data transmission method shown in an embodiment of the present application; Figure 2 This is a schematic diagram of a satellite-to-earth data transmission device according to an embodiment of the present application; Figure 3 It is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0025] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other.
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] Based on the content proposed in the background technology, the technical solution of this application needs to achieve: 1. Use spatiotemporal data to model and define the power consumption problem of satellite data transmission.
[0028] 2. Design a model based on deep reinforcement learning, namely the resource allocation model, which can produce the optimal resource allocation strategy, namely the appropriate actions (transmit power and bandwidth resources) in real time.
[0029] First of all, it is necessary to model and analyze the relevant factors affecting the energy consumption of satellite data transmission.
[0030] Let the current time be t, which represents the time feature. The time feature includes the relative time change between the satellite and the ground station, such as the orbital position of the satellite, the influence of the earth's rotation, etc. s is the state, which represents the spatial feature. The spatial feature includes the relative position, distance, angle and other spatial dynamic information between the satellite and the ground station.
[0031] It is further necessary to determine the energy consumption that affects satellite data transmission, that is, power consumption. Let the energy consumption be , in joules, and the energy consumption is It represents the energy consumption of the satellite when it performs action a in state s.
[0032] By formula It can be seen that energy consumption By transmission power Decide, is the time window, is the transmission power, t is the current time, and the calculation formula of the communication time combined with the time window is , D is the amount of transmitted data, and C is the channel capacity.
[0033] According to Shannon's formula, , S is the signal power, N is the noise power, , is the noise power spectral density, and B is the bandwidth resource.
[0034] The signal power S can be expressed by the formula To express, is the channel gain, channel gain According to the formula get, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel.
[0035] Therefore, combined with the above formula, it is finally determined that the energy consumption of transmitting data by the satellite at the current time t and in state s can be expressed as; .
[0036] Regarding the above The derivation process shows that the energy consumption of satellite data transmission is affected by the transmission power At the same time, bandwidth resource B will also affect the data transmission rate and channel gain. It will change with time and space position, and the noise power spectrum density Affects channel capacity.
[0037] The optimization goal of the technical solution of this application is to meet the various constraints of the system within the time window. When the amount of data transmitted is D, minimize the energy consumption of satellite transmission D .
[0038] Based on this, this application abstracts the above optimization objectives into a decision-making process based on deep reinforcement learning, and uses the deep reinforcement learning (DRL) algorithm to train a resource allocation model for resource allocation in the process of satellite data transmission. The resource allocation model is a deep Q network (DQN) based on Q learning and deep neural networks, which is used to deal with reinforcement learning problems with high-dimensional state space. It is very powerful in dealing with continuous state and discrete action optimization tasks in complex environments.
[0039] The pre-trained resource allocation model can determine the optimal action, namely the resource allocation plan, based on the input satellite state vector. The resource allocation plan includes the transmission power and bandwidth resources allocated to the satellite.
[0040] In one embodiment, the present application proposes a satellite-to-ground data transmission method, which can predict the future spatiotemporal data of the satellite through the LSTM network, and predict the resource allocation plan for the future time in advance in combination with the above resource allocation model. Specifically, Figure 1 FIG. 1 is a flow chart of a satellite-to-ground data transmission method according to an embodiment of the present application. Figure 1 As shown, the method includes: Step S101, collecting historical spatiotemporal data of the satellite, the historical spatiotemporal data including: historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite.
[0041] Among them, the historical GPS location information is , the historical IMU acceleration data is , the historical antenna direction angle data is .
[0042] Step S102: predicting the future space-time data of the satellite using an LSTM network based on the historical space-time data.
[0043] First, you need to define the input layer, multiple LSTM layers, and output layer of the LSTM network (Long Short-Term Memory).
[0044] The historical spatiotemporal data of the satellite is expressed as: , and input the satellite's historical spatiotemporal data of the past period of time into the LSTM network to obtain the satellite's future spatiotemporal data , the future spatiotemporal data includes: the future GPS location information is , the future IMU acceleration data is , the future antenna direction angle data is .
[0045] in Indicates the historical time corresponding to the historical spatiotemporal data, Indicates past time The time after k minutes have passed.
[0046] Step S103, generating a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmission power allocated at the current moment and the noise transmission power spectrum density of the current position of the satellite.
[0047] Among them, the future GPS location information Contains the three-dimensional coordinate information of the satellite, that is , , and the three-dimensional coordinate information of the ground station ( , , ) is known.
[0048] Therefore, according to the formula , determine the dynamic distance between the satellite and the ground station at the future time .
[0049] According to the calculation formula of channel gain , determine the satellite channel gain at the future time , according to the noise power spectral density of the current position at the future time , the currently allocated transmit power at the future time and the currently allocated bandwidth resources at future times Combine the remaining data volume of the data packet at the future time and the remaining time in the specified time window , get the future state vector at the future time .
[0050] Step S104: input the future state vector into a pre-trained resource allocation model to obtain a resource allocation solution output by the resource allocation model, such as the transmit power and bandwidth resources that need to be allocated at a future moment.
[0051] The resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward and next state as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after the satellite executes the resource allocation plan for the satellite under the state.
[0052] Step S105, and further, it is necessary to determine whether the obtained resource allocation scheme satisfies the constraint conditions. When the constraint conditions are met, the resource allocation scheme is applied to the future time of the satellite. The constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
[0053] Optionally, in one embodiment, the constraints include the following three: The communication delay is constrained by: the total amount of data transmitted by the satellite Must be within the specified time window If the time window is exceeded, it means that the satellite data transmission fails. The constraint expression of communication delay is: ; The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources Therefore, it is necessary to determine whether the bandwidth resources in the resource allocation scheme are within the range of available bandwidth resources. If the bandwidth resources in the resource allocation scheme are larger than the range of available bandwidth resources, the size of the bandwidth resources in the resource allocation scheme is adjusted to the size of the range of available bandwidth resources.
[0054] The constraint expression of bandwidth resources is: ; The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range Therefore, it is necessary to determine whether the transmit power in the resource allocation scheme is within the transmit power range. If the transmit power in the resource allocation scheme is greater than the transmit power range, the transmit power in the resource allocation scheme is adjusted to the size of the transmit power range.
[0055] The constraint expression of the transmit power control is: ; in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.
[0056] Optionally, in an embodiment, before training the resource allocation model, a state space needs to be constructed, and the state space is used to generate training samples of the resource allocation model.
[0057] Specifically, first obtain the real time and space data of each position of the satellite in orbit; the real time and space data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data; Further, according to the real antenna direction angle data, the phase change of the channel is obtained ; Further, the dynamic distance is calculated through the real GPS location information ; ; in, , , are the three-dimensional coordinates of each position of the satellite in orbit; , , are the three-dimensional coordinates of the ground stations respectively; Further, according to the dynamic distance , Fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ; Further, according to the total amount of data transmitted and the transmission rate R(t), based on the formula , determine the remaining amount of data to be transmitted ; Furthermore, according to the time window Total time and the current time t, based on the formula , determine the remaining time of the time window ; Further, according to the real GPS location information, the noise emission power spectrum density at the current location is determined ; Finally, according to the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) is obtained through the multidimensional vector It means that the multiple state vectors s(t) constitute a state space.
[0058] Optionally, before training the resource allocation model, an action space needs to be constructed, and the action space is used to generate training samples of the resource allocation model.
[0059] First, the transmit power is discretely enumerated at a first interval, such as discretely enumerating at an interval of 1W, to obtain multiple discrete transmit powers. ; The allocatable bandwidth resource range is discretely enumerated at a second interval, such as discretely enumerating at an interval of 5 MHz, to obtain a plurality of discrete bandwidth resources. ; The discrete transmission power and the discrete bandwidth resources These two discrete variable sets are combined in many-to-many ways, that is, one bandwidth resource can be combined with multiple transmit powers to obtain multiple actions, and one transmit power can also be combined with multiple bandwidth resources to obtain multiple actions. These actions constitute an action space. Therefore, the action space contains: multiple discrete transmit powers and multiple discrete bandwidth resources Composed of multiple actions .
[0060] Optionally, before training the resource allocation model, a reward function needs to be set to achieve Minimize the energy consumption when the amount of data transmitted is D .
[0061] Among them, the reward function is: ; ; W(t,s) is the energy consumption of the satellite when performing action a in state s, which can be determined by the formula of related influencing factors derived in the previous article. Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite.
[0062] Optionally, before training the resource allocation model, the method further includes: First, define the deep reinforcement learning network structure of the resource allocation model, namely the deep Q network, where the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ; Furthermore, a pool for storing multiple training sample experience replays is constructed, where each training sample is a tuple , the tuple includes the satellite's state s, action a, reward r, and next state The training samples are determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.
[0063] Optionally, the resource allocation model training process is as follows: Extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained by using the small batch training samples; at each time step, use the epsilon-greedy strategy to select actions in the action space: An action is randomly selected from the action space with probability The probability of choosing the current action from the action space The action with the largest value ; Setting the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a; in, ; ; Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite; For each training sample, the target value y is calculated, which is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is: ; in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when The target value y is taken as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, the iterative update process is: ; β is the learning rate, which means Q(s , a) Update step size; During the training process, the mean square error is used as the loss function, and the calculation formula of the loss function is: ; Where N represents the number of small batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Next, take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Next, take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.
[0064] Specifically, in one embodiment, the process of using a deep Q network (DQN) to train a resource allocation model is as follows: 1. Initialization phase: First, two deep Q networks are initialized, one is the online network and the other is the target network.
[0065] The online network is the network that actually learns and makes decisions. It can input the Q value of each possible action according to the current state, that is, the expected cumulative reward. The expected cumulative reward is the expected cumulative reward that can be obtained starting from a certain state. It can measure the goodness of taking an action. The expected cumulative reward is expressed in this article through the objective function To express.
[0066] The target network is designed to stabilize the learning process. Its parameter update is slower than that of the online network. The target network provides a relatively stable Q-value estimate, which is used to calculate a stable target value y.
[0067] After initialization, the initial parameters of the two networks are the same. During the training process, the online network continuously updates its parameters through interaction with the environment, and uses the training samples in the experience replay pool to learn how to choose the best action based on the current state. The parameters of the online network are adjusted according to the target value y and the objective function of the online network. This process can be implemented by the gradient descent algorithm to minimize the target value y and the objective function of the online network The difference between.
[0068] As the training progresses, the online network gradually learns the strategy of selecting the best action under different states. Therefore, the fully trained online network is a resource allocation model, and after the training is completed, the obtained resource allocation model is deployed to the actual system (referring to the system that executes the data transmission method between the satellite and the ground station). The resource allocation model can determine the optimal resource allocation strategy, that is, the action, according to the current state of the input satellite.
[0069] 2. Experience replay pool initialization: Furthermore, an empty experience replay pool is created to store the experience of the agent's interaction with the environment, and these experiences are used to generate training samples for model training in reinforcement learning.
[0070] The generation of experience can be understood as: each interaction of the agent in the current environment is to choose an action and then observe the result. The result is the reward corresponding to the action and the next state generated. These results constitute the experience of the agent. Each interactive experience can form a training sample, and each training sample is a tuple , It consists of four parts: satellite state s, action a, reward r and next state , the satellite state s is the current state.
[0071] The generated training samples are then stored in an experience replay pool, which allows the agent to store and review past interactions, rather than learning based solely on the most recent experience.
[0072] 3. Initial exploration of the agent: In the early stages of training, the agent performs a series of random actions in the action space in the current environment to explore different states and possible actions. After each interaction, the agent records the satellite's state s, action a, reward r, and next state , and then store these experiences as training samples in the experience replay pool.
[0073] 4. Training process: During the training process, the agent randomly extracts a small batch of samples (minibatch) from the experience replay pool for training. This method can break the time correlation between training samples and make the training process more stable.
[0074] In addition, during the action selection process of the agent, at each time step, for each training sample, the agent uses the online network and epsilon-greedy strategy to randomly select an action from the action space. The probability of randomly selecting actions to explore is The probability of selecting the action with the largest current Q value .
[0075] 5. Objective Function Calculation: ; After the agent performs the actions in the training sample, the environment will feedback the immediate reward, that is, , the immediate reward is the sum of the energy consumption and penalty term generated by executing the action.
[0076] ; The penalty term is Penalty, which is used to limit the optimization goal of the training process to minimize power consumption and complete data transmission. When the data Dtotal is transmitted within the time window, there is no penalty, Penalty=0; otherwise, the penalty needs to be increased. .
[0077] And, in order to consider the weight of future rewards, in the objective function In the calculation of ,therefore, The calculation formula represents the current state =s, and the expected cumulative reward after a series of actions are executed, where the immediate reward for each action is multiplied by the corresponding discount factor. This process is performed by the online network.
[0078] 6. Target value calculation: For each action in the training sample performed by the agent, the corresponding target value y needs to be calculated, that is, ; The target value y is used to represent the immediate reward after taking action a in state s. and discount factor Multiply the maximum expected reward of taking the best action in the next state The sum of, among which, It represents the target network's next state The maximum Q value of all possible actions is obtained by the target network.
[0079] 7. Using the target value y and the objective function of the online network Update the parameters of the online network using the difference between: The target value y is taken as The goal of iterative update is to calculate the target value y and the objective function of the online network The difference between them is used to update the parameters of the online network using the learning rate β and this difference; thereby training the online network (i.e., the resource allocation model to be trained) so that the online network Close to the actual action value function, that is, the target value y, where the iterative update process is: ; β is the learning rate, which means Q(s , a) Update step size.
[0080] 8. Loss Function: Using the target value y and the online network The difference between the two is used to calculate the loss function. In this process, the mean square error (MSE) is used as the loss function, and this difference is used to update the parameters of the online network. This update process is completed through back propagation and gradient descent.
[0081] The loss function is calculated as: ; Here N is the number of mini-batch samples, is the target value y of the jth training sample, is the online network's response to the jth sample , that is, based on the state of the current parameter in the jth training sample Next, take the action in the jth training sample The expected cumulative reward after are the current parameters of the online network (i.e., the resource allocation model to be trained).
[0082] As the training progresses, the online network can gradually learn the strategy of selecting the best action under different states. Therefore, the fully trained online network is a resource allocation model, and after the training is completed, the obtained resource allocation model is deployed to the actual system (referring to the system that executes the data transmission method between the satellite and the ground station). The resource allocation model can determine the optimal resource allocation strategy, that is, the action, according to the current state of the input satellite.
[0083] Through the above embodiments, during the training process, the agent learns the optimal strategy for action selection through continuous exploration, sample extraction, action selection, target value calculation, network training and parameter update. This process involves the collaboration of two networks: the online network is responsible for generating the current Q value estimate, i.e. ; The target network is responsible for generating a stable target Q value, that is, the target value y, to stabilize the training process and ultimately obtain a trained resource allocation model.
[0084] It should be noted that, in addition to using a deep Q-learning network, the technical solution of the present application can also be implemented using a value function method, a policy gradient method, and a hybrid method.
[0085] The key points and intended protection points of the technical solution of this application are: 1. Mathematical model of satellite data transmission power consumption Key point: Multiple constraints such as communication delay, spectrum resources and power range are taken into consideration during the design to ensure the stability and efficiency of the system under different conditions.
[0086] Points to be protected: Constraint mechanisms and their applications in satellite communication resource allocation, including the specific design of maximum delay constraints, spectrum resource limitations, and power control methods.
[0087] 2. Optimized system architecture for multi-layer model integration Key point: Integrate spatiotemporal perception and deep reinforcement learning to form an integrated dynamic optimization transmission system.
[0088] Points to be protected: A systematic approach to protecting the overall architecture, including the structure, interaction mode, and information processing flow of sub-modules such as perception, prediction, and control, as well as the system's real-time prediction and tuning capabilities for future states.
[0089] Based on the same inventive concept, another embodiment of the present application further provides a satellite-to-ground data transmission device, Figure 2 FIG. 1 is a schematic diagram of a satellite-to-ground data transmission device according to an embodiment of the present application. Figure 2 As shown, the device comprises: The acquisition module 11 is used to acquire the historical spatiotemporal data of the satellite, wherein the historical spatiotemporal data includes: the historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite; A prediction module 12 is used to predict the future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; A future state vector generating module 13 is used to generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmission power allocated at the current moment, and the noise transmission power spectrum density of the current position of the satellite; The execution module 14 is used to input the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the historical state, action, reward and next state of the satellite as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite under the state; The application module 15 is used to apply the resource allocation plan to the future time of the satellite when the resource allocation plan meets the constraint conditions, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
[0090] Optionally, the device further comprises: A readjustment module, used for readjusting the resource allocation scheme when the resource allocation scheme does not meet the constraint condition; The communication delay is constrained as follows: the total amount of data transmitted by the satellite Must be within the specified time window The communication delay constraint expression is: ; The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources The constraint expression of the bandwidth resource is: ; The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range The constraint expression of the transmit power control is: ; in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.
[0091] Optionally, the device further comprises: A real spatiotemporal data acquisition module, used to acquire the real spatiotemporal data of each position of the satellite in orbit before training the resource allocation model; the real spatiotemporal data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data; A phase change determination module is used to obtain the phase change of the channel according to the real antenna direction angle data. ; A dynamic distance calculation module is used to calculate the dynamic distance through the real GPS location information. ; ; in, , , are the three-dimensional coordinates of the satellites respectively; , , are the three-dimensional coordinates of the ground stations respectively; A channel gain determination module is used to determine the dynamic distance , Fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ; The transmission data remaining amount determination module is used to determine the remaining amount of transmission data based on the total amount of transmission data and the transmission rate R(t) based on the formula , determine the remaining amount of data to be transmitted ; The remaining time determination module is used to determine the remaining time according to the time window. Total time and the current time t, based on the formula , determine the remaining time of the time window ; The current position noise emission power spectrum density determination module is used to determine the current position noise emission power spectrum density according to the real GPS position information. ; A state space determination module is used to determine the remaining time according to the , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) is obtained through the multidimensional vector It means that the multiple state vectors s(t) constitute a state space.
[0092] Optionally, the device further comprises: A transmission power enumeration module is used to discretely enumerate the transmission power at a first interval before training the resource allocation model to obtain multiple discrete transmission powers. ; The bandwidth resource enumeration module is used to discretely enumerate the allocatable bandwidth resource range at a second interval to obtain multiple discrete bandwidth resources. ; Combining module for converting the discrete transmission power and the discrete bandwidth resources Combine them to get the action space; The action space includes: multiple discrete transmission powers and multiple discrete bandwidth resources Composed of multiple actions .
[0093] Optionally, the device comprises: The reward function setting module is used to set the reward function before training the resource allocation model. The reward function is: ; ; Wherein, W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite; The derivation process of the influencing factors of energy consumption W(t,s) is as follows: According to the transmission power Determine W(t,s), the expression is: ; in, is the time window, is the transmission power, t is the current time; Determine module, used to determine the communication time according to the formula From calculation, we know that D is the amount of data transmitted, C is the channel capacity, and ; Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained; in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel; Influencing factor determination module, For N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource; Determine .
[0094] Optionally, the device further comprises: A network structure definition module is used to define a deep reinforcement learning network structure of the resource allocation model before training the resource allocation model, wherein the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ; An experience replay pool construction module is used to construct an experience replay pool, which is used to store multiple training samples, each of which is a tuple. , the tuple includes the satellite's state s, action a, reward r, and next state The training sample is determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.
[0095] Optionally, the device further comprises: The extraction module is used to extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained through the small batch training samples; at each time step, the epsilon-greedy strategy is used to select actions in the action space: An action is randomly selected from the action space with probability The probability of choosing the current action from the action space The action with the largest value ; Objective function setting module, used to set the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a; in, ; ; Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite; The target value calculation module is used to calculate the target value y for each training sample. The target value y is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is: ; in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when Iterative update module, used to take the target value y as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, the iterative update process is: ; β is the learning rate, which means Q(s , a) Update step size; The loss calculation module is used to use the mean square error as the loss function during the training process. The calculation formula of the loss function is: ; Where N represents the number of small batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Next, take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Next, take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.
[0096] Based on the same inventive concept, another embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the satellite-to-earth data transmission method as described in any of the above embodiments.
[0097] Among them, electronic equipment refers to Figure 3 , Figure 3 is a schematic diagram of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 300 includes: a memory 310 and a processor 320. The memory 310 and the processor 320 are connected via a bus communication. A computer program is stored in the memory 310. The computer program can be run on the processor 320 to implement the steps in the satellite-to-ground data transmission method disclosed in the above embodiment of the present application.
[0098] Based on the same inventive concept, another embodiment of the present application further provides a computer program product, including a computer program, which is executed by a processor to execute the steps in the satellite-to-earth data transmission method as described in any of the above embodiments.
[0099] Based on the same inventive concept, another embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein when the program is executed by a processor, the steps in the satellite-to-earth data transmission method as described in any of the above embodiments are implemented.
[0100] As for the device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0101] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0102] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0104] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0106] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0107] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0108] The above is a detailed introduction to a satellite-to-earth data transmission method, device, equipment and medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technicians in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A satellite-to-ground data transmission method, characterized in that: The method comprises: Collecting historical spatiotemporal data of the satellite, the historical spatiotemporal data including: historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite; According to the historical spatiotemporal data, the future spatiotemporal data of the satellite is predicted using an LSTM network; Generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmission power allocated at the current moment, and the noise transmission power spectrum density of the current position of the satellite; The future state vector is input into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward and next state as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite under the state; When the resource allocation scheme satisfies the constraint conditions, the resource allocation scheme is applied to the future time of the satellite, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
2. The satellite-to-ground data transmission method according to claim 1, characterized in that: The method further comprises: When the resource allocation scheme does not satisfy the constraint condition, readjusting the resource allocation scheme; The communication delay is constrained as follows: the total amount of data transmitted by the satellite Must be within the specified time window The communication delay constraint expression is: ; The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources The constraint expression of the bandwidth resource is: ; The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range The constraint expression of the transmit power control is: ; in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.
3. The satellite-to-ground data transmission method according to claim 2, characterized in that: Before training the resource allocation model, the method further includes: Acquire the real time and space data of each position of the satellite in orbit; the real time and space data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data; According to the real antenna direction angle data, the phase change of the channel is obtained ; Calculate the dynamic distance using the real GPS location information ; ; in, , , are the three-dimensional coordinates of the satellites respectively; , , are the three-dimensional coordinates of the ground stations respectively; According to the dynamic distance , fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ; According to the total amount of data transmitted and the transmission rate R(t), based on the formula , determine the remaining amount of data to be transmitted ; According to the time window Total time and the current time t, based on the formula , determine the remaining time of the time window ; According to the real GPS location information, determine the current location noise emission power spectrum density ; According to the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) is obtained through the multidimensional vector It means that the multiple state vectors s(t) constitute a state space.
4. The satellite-to-ground data transmission method according to claim 3, characterized in that: Before training the resource allocation model, the method further includes: The transmit power is discretely enumerated at a first interval to obtain a plurality of discrete transmit powers ; Discretely enumerate the allocatable bandwidth resource range at a second interval to obtain multiple discrete bandwidth resources ; The discrete transmission power and the discrete bandwidth resources Combine them to get the action space; The action space includes: multiple discrete transmission powers and multiple discrete bandwidth resources Composed of multiple actions .
5. The satellite-to-ground data transmission method according to claim 4, characterized in that: Before training the resource allocation model, the method further includes: Set a reward function, the reward function is: ; ; Wherein, W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite; The derivation process of the influencing factors of energy consumption W(t,s) is as follows: According to the transmission power Determine W(t,s), the expression is: ; in, is the time window, is the transmission power, t is the current time; According to the communication time, the formula From calculation, we know that D is the amount of data transmitted, C is the channel capacity, and ; Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained; in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel; N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource; Determine .
6. The satellite-to-ground data transmission method according to claim 5, characterized in that: Before training the resource allocation model, the method further includes: Define a deep reinforcement learning network structure of the resource allocation model, wherein the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ; Construct an experience replay pool, which is used to store multiple training samples, each training sample is a tuple , the tuple includes the satellite's state s, action a, reward r, and next state The training sample is determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.
7. The satellite-to-ground data transmission method according to claim 6, characterized in that: The resource allocation model training process is as follows: Extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained by using the small batch training samples; at each time step, use the epsilon-greedy strategy to select actions in the action space: The probability of randomly selecting an action from the action space is The probability of choosing the current action from the action space The action with the largest value ; Setting the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a; in, ; ; Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite; For each training sample, the target value y is calculated, which is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is: ; in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when The target value y is taken as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, the iterative update process is: ; β is the learning rate, which means Q(s , a) Update step size; During the training process, the mean square error is used as the loss function, and the calculation formula of the loss function is: ; Where N represents the number of small batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Next, take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Next, take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.
8. A satellite-to-ground data transmission device, characterized in that: The device comprises: A collection module is used to collect the historical spatiotemporal data of the satellite, wherein the historical spatiotemporal data includes: the historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite; A prediction module, used to predict the future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; A future state vector generation module, used to generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmission power allocated at the current moment and the noise transmission power spectrum density of the current position of the satellite; An execution module is used to input the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the historical state, action, reward and next state of the satellite as training samples; the action refers to the resource allocation plan for the satellite under the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite under the state; The application module is used to apply the resource allocation scheme to the future time of the satellite when the resource allocation scheme meets the constraint conditions, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the satellite-to-ground data transmission method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, wherein when the computer program is executed by a processor, the satellite-to-ground data transmission method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Wave hopping satellite service prediction and multi-dimensional link dynamic resource allocation method and device
CN116546624A
Satellite-ground convergence network resource allocation method
CN116981091A
Satellite edge computing task unloading and resource allocation method based on deep reinforcement learning
CN118250750A
Space-air-ground integrated network task unloading method based on hierarchical deep reinforcement learning
CN118870435A
Method, apparatus, and recording medium for processing data of e-commerce services
KR1020250045076A
Cited By
Wireless communication power prediction method and system based on machine learning
CN121664269A