Satellite-to-ground data transmission method, device, equipment, and medium

The resource allocation model trained by LSTM network and deep reinforcement learning solves the problem of insufficient adaptability of CSI technology to dynamic position changes in satellite communications, realizes efficient resource allocation and energy management, and adapts to complex environmental changes.

CN120017140BActive Publication Date: 2025-09-30YIWEI AEROSPACE (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197305.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-09-30
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Existing dynamic resource allocation technologies based on channel state information (CSI) in satellite communications have difficulty adapting quickly to the dynamic position changes of satellites and ground stations. They lack the perception and prediction of multi-dimensional spatiotemporal characteristics and cannot achieve efficient resource allocation in complex environments.

Method used

An LSTM network is used to predict the future spatiotemporal data of satellites, and a deep reinforcement learning algorithm is used to train a resource allocation model to generate future state vectors and perform resource allocation to meet the constraints of communication delay, bandwidth resources, and transmission power.

Benefits of technology

It improves the prediction accuracy of future spatiotemporal data, dynamically adjusts resource allocation plans, optimizes transmission efficiency, adapts to changes in complex spatiotemporal environments, reduces energy consumption, and improves the adaptability and robustness of the system under dynamic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017140B_ABST
    Figure CN120017140B_ABST
Patent Text Reader

Abstract

The present application provides a satellite-to-ground data transmission method, apparatus, device, and medium, the method comprising: collecting historical spatiotemporal data of a satellite, and predicting future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; generating a future state vector based on the future spatiotemporal data; inputting the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples; the action is the resource allocation plan for the satellite under the state, the reward is the immediate reward after the satellite executes the resource allocation plan for the satellite under the state, and the next state is the state of the satellite after executing the resource allocation plan for the satellite under the next state; and when the resource allocation plan satisfies the constraint conditions, the resource allocation plan is applied to the satellite at a future moment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data transmission technology, and in particular to a satellite-to-ground data transmission method, apparatus, device, and medium. Background Art

[0002] The fundamental principle of satellite-ground collaboration technology is the collaborative work between satellites and ground stations to achieve efficient data transmission, resource sharing, and mission execution. The key to satellite-ground collaboration lies in ensuring real-time data exchange between satellites and ground stations through well-designed communication protocols and efficient collaboration mechanisms, optimizing overall system performance. While in orbit, satellites need to transmit collected data to ground stations or receive instructions from them to perform specific tasks. This process requires high data reliability and low latency. Satellite-ground collaboration technology ensures real-time data integrity and integrity through multiple transmission links (such as X-band and Ka-band) and multi-band communication technologies. Satellite-ground collaboration allows for the rational allocation and optimization of resources between ground stations and satellites. For example, ground stations can provide satellites with real-time computing power, reducing the satellite's computational burden and improving mission efficiency. This allows satellites to complete more complex missions without increasing hardware overhead.

[0003] Among related technologies, there's a dynamic resource allocation technique based on channel state information (CSI). This technique dynamically adjusts resource allocation strategies by acquiring key communication channel parameters (such as channel gain and noise level) in real time. This technique uses CSI to monitor channel conditions between satellites and ground stations in real time. By analyzing the instantaneous channel state, it dynamically allocates bandwidth resources and transmit power based on environmental changes to optimize transmission performance.

[0004] For example, when channel conditions are favorable, the system will reduce transmit power to save energy or increase transmission rates when bandwidth resources are sufficient. When channel conditions are poor, the system will increase transmit power or adjust bandwidth resources to ensure communication link stability. CSI dynamic resource allocation technology is suitable for satellite communication scenarios with highly volatile communication environments. Its core design relies on real-time resource adjustment based on channel feedback to achieve dynamic optimization of communication efficiency.

[0005] This dynamic resource allocation technology based on channel state information (CSI) can optimize bandwidth resources and transmission power usage in satellite communications to a certain extent, but the technology has limitations, especially in complex spatiotemporal environments.

[0006] First, CSI technology relies primarily on single-channel feedback information, such as channel gain and noise level. This makes it difficult to fully capture the dynamic relative position between satellites and ground stations, distance changes, and the impact of orbital factors. Consequently, when satellites and ground stations are moving at high speeds or experiencing significant position changes, this technology may not be able to quickly adapt to fluctuations in channel characteristics, leading to inaccurate resource allocation.

[0007] Secondly, CSI technology lacks high-precision perception and prediction methods for multi-dimensional spatiotemporal features and fails to integrate spatiotemporal information from multiple sources such as GPS and inertial measurement units (IMUs). This limits the system's adaptability in complex dynamic environments.

[0008] Furthermore, because it lacks deep reinforcement learning or adaptive optimization algorithms, CSI technology lacks the intelligence to respond to sudden environmental changes and cannot flexibly optimize resource allocation strategies. Therefore, although CSI technology performs well in static or slightly changing channel environments, its effectiveness is limited in satellite communication scenarios that require high precision and dynamic response, making it difficult to achieve optimal transmission efficiency and energy management. Summary of the Invention

[0009] In view of this, embodiments of the present application provide a satellite-to-ground data transmission method, apparatus, device, and medium to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.

[0010] A first aspect of an embodiment of the present application provides a satellite-to-ground data transmission method, the method comprising:

[0011] Collecting historical spatiotemporal data of the satellite, the historical spatiotemporal data including: historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite;

[0012] Based on the historical spatiotemporal data, an LSTM network is used to predict the future spatiotemporal data of the satellite;

[0013] Generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of a specified time window, the remaining data volume of a data packet, a channel gain, a bandwidth resource allocated at a current moment, a transmit power allocated at a current moment, and a noise transmit power spectral density at a current location of the satellite;

[0014] Inputting the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples; the action refers to the resource allocation plan for the satellite in the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite in the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite in the state;

[0015] When the resource allocation scheme satisfies a constraint condition, the resource allocation scheme is applied to the future time of the satellite, wherein the constraint condition is used to constrain the communication delay, bandwidth resources and transmit power control of the satellite during data transmission.

[0016] Optionally, the method further includes:

[0017] When the resource allocation plan does not meet the constraint conditions, readjusting the resource allocation plan;

[0018] The communication delay is constrained as follows: the total amount of data transmitted by the satellite Must be within the specified time window The communication delay constraint expression is:

[0019] ;

[0020] The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources The constraint expression of the bandwidth resource is:

[0021] ;

[0022] The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range The constraint expression of the transmit power control is:

[0023] ;

[0024] in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.

[0025] Optionally, before training the resource allocation model, the method further includes:

[0026] Acquire real time and space data of each position of the satellite in orbit; the real time and space data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data;

[0027] According to the real antenna direction angle data, the phase change of the channel is obtained ;

[0028] Calculate dynamic distance using the real GPS location information ; ;

[0029] in, 、 、 are the three-dimensional coordinates of the satellites respectively; 、 、 are the three-dimensional coordinates of the ground stations respectively;

[0030] According to the dynamic distance , fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ;

[0031] According to the total amount of data transmitted and the transmission rate R(t), based on the formula , determine the remaining amount of transmission data ;

[0032] According to the time window Total time and the current time t, based on the formula , determine the remaining time of the time window ;

[0033] According to the real GPS location information, determine the current location noise emission power spectrum density ;

[0034] According to the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) through the multidimensional vector It indicates that the multiple state vectors s(t) constitute a state space.

[0035] Optionally, before training the resource allocation model, the method further includes:

[0036] The transmit power is discretely enumerated at the first interval to obtain multiple discrete transmit powers ;

[0037] Discretely enumerate the allocatable bandwidth resource range at the second interval to obtain multiple discrete bandwidth resources ;

[0038] The discrete transmit power and the discrete bandwidth resources Combine to get the action space;

[0039] The action space includes: multiple discrete transmission powers and multiple discrete bandwidth resources Composed of multiple actions .

[0040] Optionally, before training the resource allocation model, the method further includes:

[0041] Set the reward function, which is:

[0042] ;

[0043] ;

[0044] Where W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite;

[0045] The derivation process of the factors affecting energy consumption W(t,s) is as follows:

[0046] According to the transmission power Determine W(t,s), the expression is:

[0047] ;

[0048] in, is the time window, is the transmission power, t is the current time;

[0049] According to the communication time, the formula Calculation shows that D is the amount of data transmitted, C is the channel capacity, and ;

[0050] Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained;

[0051] in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel;

[0052] N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource;

[0053] Determine .

[0054] Optionally, before training the resource allocation model, the method further includes:

[0055] Define the deep reinforcement learning network structure of the resource allocation model, wherein the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ;

[0056] Construct an experience replay pool, which is used to store multiple training samples, each training sample is a tuple , the tuple includes the satellite's state s, action a, reward r and next state The training sample is determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.

[0057] Optionally, the resource allocation model training process is as follows:

[0058] Extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained through the small batch training samples; at each time step, use the epsilon-greedy strategy to select the action in the action space: An action is randomly selected from the action space with a probability of The probability of choosing the current action from the action space The action with the largest value ;

[0059] Setting the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a;

[0060] in, ;

[0061] ;

[0062] Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite;

[0063] For each training sample, calculate the target value y, which is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is:

[0064] ;

[0065] in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when

[0066] The target value y is taken as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, where the iterative update process is:

[0067] ;

[0068] β is the learning rate, which means Q(s , a) Update step size;

[0069] During the training process, the mean square error is used as the loss function, and the calculation formula of the loss function is:

[0070] ;

[0071] Where N represents the number of mini-batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.

[0072] A second aspect of an embodiment of the present application provides a satellite-to-ground data transmission device, the device comprising:

[0073] An acquisition module is used to acquire historical spatiotemporal data of the satellite, wherein the historical spatiotemporal data includes: historical GPS position information, IMU acceleration data, and antenna direction angle data of the satellite;

[0074] A prediction module is used to predict the future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data;

[0075] a future state vector generation module, configured to generate a future state vector based on the future spatiotemporal data; the future state vector including: the remaining time of a specified time window, the remaining amount of data in a data packet, a channel gain, bandwidth resources allocated at the current moment, transmit power allocated at the current moment, and a noise transmit power spectral density at the current location of the satellite;

[0076] An execution module is configured to input the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples; the action refers to the resource allocation plan for the satellite in the state; the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite in the state; and the next state refers to the state of the satellite after the resource allocation plan for the satellite in the state is executed;

[0077] An application module is configured to apply the resource allocation scheme to a future moment of the satellite when the resource allocation scheme satisfies a constraint condition, wherein the constraint condition is used to constrain communication delay, bandwidth resources, and transmit power control of the satellite during data transmission.

[0078] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method described in the first aspect.

[0079] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0080] Beneficial effects of this application:

[0081] The present application provides a satellite-to-ground data transmission method, apparatus, device, and medium. The method comprises: collecting historical spatiotemporal data of a satellite, wherein the historical spatiotemporal data comprises: historical GPS position information, IMU acceleration data, and antenna azimuth data of the satellite; predicting future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; generating a future state vector based on the future spatiotemporal data; the future state vector comprises: the remaining time of a specified time window, the remaining data volume of a data packet, a channel gain, bandwidth resources allocated at the current moment, transmit power allocated at the current moment, and noise transmit power spectrum density at the current position of the satellite; inputting the future state vector into a pre-trained resource allocation model to obtain The resource allocation model outputs a resource allocation plan; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples; the action refers to the resource allocation plan for the satellite in the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite in the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite in the state; when the resource allocation plan meets the constraint conditions, the resource allocation plan is applied to the future time of the satellite, and the constraint conditions are used to constrain the communication delay, bandwidth resources, and transmission power control of the satellite during data transmission.

[0082] Through the technical solution of the present application, the LSTM network can be used to analyze and predict the historical spatiotemporal data of the satellite, which can significantly improve the prediction accuracy of future spatiotemporal data. The resource allocation model obtained based on deep reinforcement learning training can dynamically adjust the resource allocation plan according to the future state vector, and achieve optimal resource allocation while meeting constraints such as communication delay, bandwidth resources and transmission power control. The resource allocation model trained by the deep reinforcement learning algorithm can minimize energy consumption while ensuring data transmission tasks, improve transmission efficiency, adapt to complex spatiotemporal environment changes, and improve the adaptability and robustness of the system in the face of dynamic changes, providing strong technical support for the development of communications between satellites and the ground. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] The drawings that constitute a part of this application are used to provide further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute improper limitations on this application.

[0084] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0085] Figure 1 This is a flowchart of a satellite-to-ground data transmission method according to an embodiment of the present application;

[0086] Figure 2 This is a schematic diagram of a satellite-to-ground data transmission device according to an embodiment of the present application;

[0087] Figure 3 Schematic diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0088] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.

[0089] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0090] Based on the content proposed in the background technology, the technical solution of this application needs to achieve:

[0091] 1. Use spatiotemporal data to model and define the power consumption of satellite data transmission.

[0092] 2. Design a deep reinforcement learning-based model, namely a resource allocation model, that can generate the optimal resource allocation strategy, namely the appropriate actions (transmit power and bandwidth resources), in real time.

[0093] First, it is necessary to model and analyze the relevant factors affecting the energy consumption of satellite data transmission.

[0094] Let the current time be t, which represents the temporal characteristics. Temporal characteristics include the relative time changes between the satellite and the ground station, such as the satellite's orbital position and the influence of the Earth's rotation. s is the state, which represents the spatial characteristics. Spatial characteristics include spatial dynamic information such as the relative position, distance, and angle between the satellite and the ground station.

[0095] It is further necessary to determine the energy consumption that affects satellite data transmission, that is, power consumption. Let the energy consumption be , in joules, and the energy consumption is It represents the energy consumption of the satellite when it performs action a in state s.

[0096] By the formula It can be seen that energy consumption By transmission power Decide, is the time window, is the transmission power, t is the current moment, and the calculation formula for the communication time combined with the time window is , D is the amount of transmitted data, and C is the channel capacity.

[0097] According to Shannon's formula, , S is the signal power, N is the noise power, , is the noise power spectral density, and B is the bandwidth resource.

[0098] The signal power S can be expressed by the formula To express, is the channel gain, channel gain According to the formula get, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel.

[0099] Therefore, combined with the above formula, it is finally determined that the energy consumption of transmitting data by the satellite at the current time t and in state s can be expressed as; .

[0100] Regarding the above The derivation process shows that the energy consumption of satellite data transmission is affected by the transmission power At the same time, bandwidth resource B will also affect the data transmission rate and channel gain. It will change with time and spatial position, and the noise power spectrum density Affects channel capacity.

[0101] The optimization goal of the technical solution of this application is to meet the various constraints of the system within the time window. When the amount of data transmitted is D, minimize the energy consumption of satellite transmission D .

[0102] Based on this, this application abstracts the above optimization objectives into a decision-making process based on deep reinforcement learning. Using the deep reinforcement learning (DRL) algorithm, a resource allocation model is trained to allocate resources during satellite data transmission. The resource allocation model is a deep Q-network (DQN) based on Q-learning and deep neural networks. It is used to handle reinforcement learning problems with high-dimensional state spaces and is very powerful in handling continuous state and discrete action optimization tasks in complex environments.

[0103] The pre-trained resource allocation model can determine the optimal action, namely the resource allocation plan, based on the input satellite state vector. The resource allocation plan includes the transmission power and bandwidth resources allocated to the satellite.

[0104] In one embodiment, the present application proposes a satellite-to-ground data transmission method that can predict the future spatiotemporal data of a satellite through an LSTM network, and combine the above-mentioned resource allocation model to predict the resource allocation plan for the future in advance. Specifically, Figure 1 FIG. 1 is a flow chart of a satellite-to-ground data transmission method according to an embodiment of the present application. Figure 1 As shown, the method includes:

[0105] Step S101 : collecting historical spatiotemporal data of a satellite, wherein the historical spatiotemporal data includes: historical GPS position information, IMU acceleration data, and antenna direction angle data of the satellite.

[0106] Among them, the historical GPS location information is , the historical IMU acceleration data is , the historical antenna direction angle data is .

[0107] Step S102: Based on the historical spatiotemporal data, the future spatiotemporal data of the satellite is predicted using an LSTM network.

[0108] First, we need to define the input layer, multiple LSTM layers, and output layer of the LSTM network (Long Short-Term Memory).

[0109] The historical spatiotemporal data of the satellite is expressed as: And input the satellite's historical spatiotemporal data of the past period into the LSTM network to obtain the satellite's future spatiotemporal data , the future spatiotemporal data includes: the future GPS location information is , the future IMU acceleration data is , the future antenna direction angle data is .

[0110] in Indicates the historical time corresponding to the historical spatiotemporal data, Indicates past time The time after k minutes have passed.

[0111] Step S103: Generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of the specified time window, the remaining data volume of the data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmit power allocated at the current moment, and the noise transmit power spectral density at the current location of the satellite.

[0112] Among them, future GPS location information Contains the three-dimensional coordinate information of the satellite, that is 、 、 and the three-dimensional coordinate information of the ground station ( 、 、 ) is known.

[0113] Therefore, according to the formula , determine the dynamic distance between the satellite and the ground station at the future time .

[0114] Then according to the calculation formula of channel gain , determine the satellite channel gain at the future time , according to the noise power spectrum density of the current position at the future time , the currently allocated transmit power at the future time and the currently allocated bandwidth resources at future times Combined with the remaining data volume of the data packet at the future time and the remaining time in the specified time window , get the future state vector at the future time .

[0115] Step S104: input the future state vector into a pre-trained resource allocation model to obtain a resource allocation solution output by the resource allocation model, such as transmit power and bandwidth resources to be allocated at a future moment.

[0116] The resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples. The action refers to the resource allocation plan for the satellite in the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite in the state, and the next state refers to the state of the satellite after the satellite executes the resource allocation plan for the satellite in the state.

[0117] In step S105, it is also necessary to determine whether the obtained resource allocation plan meets the constraint conditions. When the constraint conditions are met, the resource allocation plan is applied to the future time of the satellite. The constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.

[0118] Optionally, in one embodiment, the constraints include the following three:

[0119] The communication delay is constrained to express: the total amount of data transmitted by the satellite Must be within the specified time window If the time window is exceeded, it means that the satellite data transmission fails. The constraint expression of communication delay is:

[0120] ;

[0121] The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources Therefore, it is necessary to determine whether the bandwidth resources in the resource allocation plan are within the range of available bandwidth resources. If the bandwidth resources in the resource allocation plan are larger than the range of available bandwidth resources, the size of the bandwidth resources in the resource allocation plan is adjusted to the size of the range of available bandwidth resources.

[0122] The constraint expression of bandwidth resources is:

[0123] ;

[0124] The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range Therefore, it is necessary to determine whether the transmit power in the resource allocation scheme is within the transmit power range. If the transmit power in the resource allocation scheme is greater than the transmit power range, the transmit power in the resource allocation scheme is adjusted to the size of the transmit power range.

[0125] The constraint expression of the transmit power control is:

[0126] ;

[0127] in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.

[0128] Optionally, in one embodiment, before training the resource allocation model, a state space needs to be constructed, and the state space is used to generate training samples of the resource allocation model.

[0129] Specifically, first obtain the real time and space data of each position of the satellite in orbit; the real time and space data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data;

[0130] Further, according to the real antenna direction angle data, the phase change of the channel is obtained ;

[0131] Furthermore, the dynamic distance is calculated using the real GPS location information. ; ;

[0132] in, 、 、 are the three-dimensional coordinates of each position of the satellite in orbit; 、 、 are the three-dimensional coordinates of the ground stations respectively;

[0133] Further, according to the dynamic distance , fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ;

[0134] Further, according to the total amount of data transmitted and the transmission rate R(t), based on the formula , determine the remaining amount of transmission data ;

[0135] Furthermore, according to the time window Total time and the current time t, based on the formula , determine the remaining time of the time window ;

[0136] Further, according to the real GPS location information, the noise emission power spectrum density of the current location is determined ;

[0137] Finally, according to the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) through the multidimensional vector It indicates that the multiple state vectors s(t) constitute a state space.

[0138] Optionally, before training the resource allocation model, an action space needs to be constructed, and the action space is used to generate training samples of the resource allocation model.

[0139] First, the transmit power is discretely enumerated at a first interval, such as discretely enumerating at intervals of 1W, to obtain multiple discrete transmit powers. ;

[0140] The allocatable bandwidth resource range is discretely enumerated at a second interval, such as discrete enumeration at an interval of 5 MHz, to obtain multiple discrete bandwidth resources. ;

[0141] The discrete transmit power and the discrete bandwidth resources These two discrete variable sets are combined in a many-to-many manner, that is, one bandwidth resource can be combined with multiple transmit powers to obtain multiple actions, and one transmit power can also be combined with multiple bandwidth resources to obtain multiple actions. These actions constitute an action space. Therefore, the action space contains: multiple discrete transmit powers and multiple discrete bandwidth resources Composed of multiple actions .

[0142] Optionally, before training the resource allocation model, a reward function needs to be set to achieve Minimize energy consumption when the amount of data transmitted is D .

[0143] Among them, the reward function is:

[0144] ;

[0145] ;

[0146] W(t,s) is the energy consumption of the satellite when performing action a in state s, which can be determined by the formula of relevant influencing factors derived above. Penalty represents the penalty term. is the penalty factor; is the total amount of data transmitted by the satellite.

[0147] Optionally, before training the resource allocation model, the method further includes:

[0148] First, define the deep reinforcement learning network structure of the resource allocation model, namely the deep Q network, where the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ;

[0149] Furthermore, a replay pool is constructed to store multiple training sample experiences, where each training sample is a tuple. , the tuple includes the satellite's state s, action a, reward r, and next state , the training samples are determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.

[0150] Optionally, the resource allocation model training process is as follows:

[0151] Extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained through the small batch training samples; at each time step, use the epsilon-greedy strategy to select the action in the action space: An action is randomly selected from the action space with a probability of The probability of choosing the current action from the action space The action with the largest value ;

[0152] Setting the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a;

[0153] in, ;

[0154] ;

[0155] Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite;

[0156] For each training sample, calculate the target value y, which is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is:

[0157] ;

[0158] in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when

[0159] The target value y is taken as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, where the iterative update process is:

[0160] ;

[0161] β is the learning rate, which means Q(s , a) Update step size;

[0162] During the training process, the mean square error is used as the loss function, and the calculation formula of the loss function is:

[0163] ;

[0164] Where N represents the number of mini-batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.

[0165] Specifically, in one embodiment, the process of using a deep Q-network (DQN) to train a resource allocation model is as follows:

[0166] 1. Initialization phase:

[0167] First, two deep Q networks are initialized, one is the online network and the other is the target network.

[0168] The online network is the network that actually learns and makes decisions. It can input the Q value of each possible action according to the current state, that is, the expected cumulative reward. The expected cumulative reward is the expected cumulative reward that can be obtained from a certain state. It can measure the goodness of taking an action. The expected cumulative reward is expressed in this article through the objective function To express.

[0169] The target network is designed to stabilize the learning process. Its parameter updates are slower than those of the online network. The target network provides a relatively stable Q-value estimate, which is used to calculate a stable target value y.

[0170] After initialization, the initial parameters of the two networks are the same. During the training process, the online network continuously updates its parameters through interaction with the environment, and uses the training samples in the experience replay pool to learn how to choose the best action based on the current state. The parameters of the online network will be adjusted according to the target value y and the objective function of the online network. This process can be achieved by gradient descent algorithm to minimize the target value y and the objective function of the online network The difference between.

[0171] As training progresses, the online network gradually learns the strategy for selecting the optimal action under different states. Therefore, the fully trained online network becomes a resource allocation model. After training, the resulting resource allocation model is deployed to the actual system (the system that performs data transmission between satellites and ground stations). The resource allocation model can determine the optimal resource allocation strategy, i.e., the action, based on the current state of the satellite input.

[0172] 2. Experience replay pool initialization:

[0173] Furthermore, an empty experience replay pool is created to store the experience of the agent's interaction with the environment, and these experiences are used to generate training samples for model training in reinforcement learning.

[0174] The generation of experience can be understood as: each interaction of the agent in the current environment is to choose an action and then observe the result. The result is the reward corresponding to the action and the next state generated. These results constitute the experience of the agent. Each interaction experience can form a training sample, and each training sample is a tuple , It consists of four parts: satellite state s, action a, reward r and next state , the state s of the satellite is the current state.

[0175] The generated training samples are then stored in an experience replay pool, which allows the agent to store and review past interactions, rather than learning based solely on the most recent experience.

[0176] 3. Initial exploration of the agent:

[0177] At the beginning of training, the agent performs a series of random actions in the action space in the current environment to explore different states and possible actions. After each interaction, the agent records the satellite's state s, action a, reward r, and next state. , and then store these experiences as training samples in the experience replay pool.

[0178] 4. Training process:

[0179] During the training process, the agent randomly extracts a minibatch of samples from the experience replay pool for training. This method can break the temporal correlation between training samples and make the training process more stable.

[0180] In addition, during the action selection process of the agent, at each time step, for each training sample, the agent uses the online network and epsilon-greedy strategy to randomly select an action from the action space. The probability of randomly selecting actions to explore is The probability of selecting the action with the largest current Q value .

[0181] 5. Objective Function Calculation:

[0182] ;

[0183] After the agent performs the actions in the training sample, the environment will feedback immediate rewards, that is, , the immediate reward is the sum of the energy consumption and penalty generated by executing the action.

[0184] ;

[0185] The penalty term is Penalty, which is used to limit the optimization goal of the training process to minimize power consumption and complete data transmission. When the data Dtotal is transmitted within the time window, there is no penalty and Penalty = 0; otherwise, the penalty needs to be increased.

[0186] .

[0187] And, in order to consider the weight of future rewards, in the objective function The discount factor is also introduced in the calculation of ,therefore, The calculation formula represents the current state =s, and the expected cumulative reward obtained after executing a series of actions, where the immediate reward of executing each action is multiplied by the corresponding discount factor. This process is performed by the online network.

[0188] 6. Target value calculation:

[0189] For each action in the training sample performed by the agent, the corresponding target value y needs to be calculated, that is, ;

[0190] The target value y is used to represent the immediate reward after taking action a in state s and discount factor Multiply the maximum expected reward of taking the best action in the next state The sum of, It represents the target network's next state The maximum Q value of all possible actions is obtained by the target network.

[0191] 7. Using the target value y and the objective function of the online network Update the parameters of the online network using the difference between:

[0192] The target value y is taken as The goal of iterative update is to calculate the target value y and the objective function of the online network The difference between them is used to update the parameters of the online network using the learning rate β and this difference; thereby training the online network (i.e., the resource allocation model to be trained) so that the online network Close to the actual action value function, that is, the target value y, where the iterative update process is:

[0193] ;

[0194] β is the learning rate, which means Q(s , a) Update step size.

[0195] 8. Loss Function:

[0196] Using the target value y and the online network The loss function is calculated by taking the difference between the two. In this process, the mean square error (MSE) is used as the loss function, and this difference is used to update the parameters of the online network. This update process is completed through back propagation and gradient descent.

[0197] The calculation formula of the loss function is: ;

[0198] Here N is the number of mini-batch samples, is the target value y of the jth training sample, is the online network's response to the jth sample , that is, the state of the current parameters in the jth training sample Take the action in the jth training sample The expected cumulative reward after are the current parameters of the online network (i.e., the resource allocation model to be trained).

[0199] As training progresses, the online network gradually learns the strategy for selecting the optimal action under different states. Therefore, the fully trained online network becomes a resource allocation model. After training, the resulting resource allocation model is deployed to the actual system (the system that performs data transmission between satellites and ground stations). The resource allocation model determines the optimal resource allocation strategy, or action, based on the current state of the satellite input.

[0200] Through the above embodiment, during the training process, the agent learns the optimal strategy for action selection through continuous exploration, sample extraction, action selection, target value calculation, network training and parameter update. This process involves the collaboration of two networks: the online network is responsible for generating the current Q value estimate, i.e. The target network is responsible for generating a stable target Q value, that is, the target value y, to stabilize the training process and ultimately obtain a trained resource allocation model.

[0201] It should be noted that, in addition to using a deep Q-learning network, the technical solution of this application can also be implemented using a value function method, a policy gradient method, and a hybrid method.

[0202] The key points and intended protection points of the technical solution of this application are:

[0203] 1. Mathematical model of satellite data transmission power consumption

[0204] Key Point: The design takes into account multiple constraints such as communication latency, spectrum resources, and power range to ensure system stability and efficiency under different conditions.

[0205] Protection point: Constraint mechanism and its application in satellite communication resource allocation, including the specific design of maximum delay constraint, spectrum resource limitation and power control method.

[0206] 2. Optimized system architecture for multi-layer model integration

[0207] Key point: Integrate spatiotemporal perception and deep reinforcement learning to form an integrated dynamic optimization transmission system.

[0208] Points to be protected: A systematic approach to protecting the overall architecture, including the structure, interaction mode, and information processing flow of sub-modules such as perception, prediction, and control, as well as the system's real-time prediction and optimization capabilities for future states.

[0209] Based on the same inventive concept, another embodiment of the present application further provides a satellite-to-ground data transmission device. Figure 2 FIG. 1 is a schematic diagram of a satellite-to-ground data transmission device according to an embodiment of the present application. Figure 2 As shown, the device includes:

[0210] The acquisition module 11 is used to collect the historical spatiotemporal data of the satellite, wherein the historical spatiotemporal data includes: the historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite;

[0211] A prediction module 12 is configured to predict the future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data;

[0212] A future state vector generation module 13 is configured to generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of a specified time window, the remaining amount of data in a data packet, the channel gain, the bandwidth resources allocated at the current moment, the transmit power allocated at the current moment, and the noise transmit power spectral density at the current location of the satellite;

[0213] An execution module 14 is configured to input the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical states, actions, rewards, and next states as training samples; the action refers to the resource allocation plan for the satellite under the state; the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite under the state; and the next state refers to the state of the satellite after the resource allocation plan for the satellite under the state is executed.

[0214] The application module 15 is used to apply the resource allocation plan to the future time of the satellite when the resource allocation plan meets the constraint conditions, and the constraint conditions are used to constrain the communication delay, bandwidth resources and transmission power control of the satellite during data transmission.

[0215] Optionally, the device further comprises:

[0216] A readjustment module, configured to readjust the resource allocation scheme when the resource allocation scheme does not meet the constraint conditions;

[0217] The communication delay is constrained as follows: the total amount of data transmitted by the satellite Must be within the specified time window The communication delay constraint expression is:

[0218] ;

[0219] The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources The constraint expression of the bandwidth resource is:

[0220] ;

[0221] The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range The constraint expression of the transmit power control is:

[0222] ;

[0223] in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.

[0224] Optionally, the device further comprises:

[0225] A real spatiotemporal data acquisition module is used to acquire the real spatiotemporal data of each position of the satellite in orbit before training the resource allocation model; the real spatiotemporal data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data;

[0226] A phase change determination module is used to obtain the phase change of the channel based on the real antenna direction angle data. ;

[0227] Dynamic distance calculation module, used to calculate the dynamic distance through the real GPS location information ; ;

[0228] in, 、 、 are the three-dimensional coordinates of the satellites respectively; 、 、 are the three-dimensional coordinates of the ground stations respectively;

[0229] A channel gain determination module is configured to determine a channel gain according to the dynamic distance , fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ;

[0230] The module for determining the remaining amount of transmission data is used to determine the remaining amount of transmission data based on the formula according to the total amount of transmission data and the transmission rate R(t). , determine the remaining amount of transmission data ;

[0231] The remaining time determination module is used to determine the remaining time according to the time window. Total time and the current time t, based on the formula , determine the remaining time of the time window ;

[0232] The current position noise emission power spectrum density determination module is used to determine the current position noise emission power spectrum density based on the real GPS position information. ;

[0233] A state space determination module is used to determine the state space based on the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) through the multidimensional vector It indicates that the multiple state vectors s(t) constitute a state space.

[0234] Optionally, the device further comprises:

[0235] The transmit power enumeration module is used to discretely enumerate the transmit power at a first interval before training the resource allocation model to obtain multiple discrete transmit powers. ;

[0236] The bandwidth resource enumeration module is used to discretely enumerate the allocable bandwidth resource range at a second interval to obtain multiple discrete bandwidth resources. ;

[0237] Combining module for combining the discrete transmit power and the discrete bandwidth resources Combine to get the action space;

[0238] The action space includes: multiple discrete transmission powers and multiple discrete bandwidth resources Composed of multiple actions .

[0239] Optionally, the device comprises:

[0240] The reward function setting module is used to set the reward function before training the resource allocation model. The reward function is:

[0241] ;

[0242] ;

[0243] Where W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite;

[0244] The derivation process of the factors affecting energy consumption W(t,s) is as follows:

[0245] W(t,s) is determined according to the transmit power P(t,s), and the expression is:

[0246] ;

[0247] in, is the time window, is the transmission power, t is the current time;

[0248] Determine module, used to determine the communication time according to the formula Calculation shows that D is the amount of data transmitted, C is the channel capacity, and ;

[0249] Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained;

[0250] in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel;

[0251] Influencing factor determination module, For N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource;

[0252] Determine .

[0253] Optionally, the device further comprises:

[0254] A network structure definition module is used to define the deep reinforcement learning network structure of the resource allocation model before training the resource allocation model, wherein the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ;

[0255] The experience replay pool construction module is used to build an experience replay pool, which is used to store multiple training samples, each training sample is a tuple , the tuple includes the satellite's state s, action a, reward r and next state The training sample is determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.

[0256] Optionally, the device further comprises:

[0257] The extraction module is used to extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained through the small batch training samples; at each time step, the epsilon-greedy strategy is used to select the action in the action space: An action is randomly selected from the action space with a probability of The probability of choosing the current action from the action space The action with the largest value ;

[0258] Objective function setting module, used to set the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a;

[0259] in, ;

[0260] ;

[0261] Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite;

[0262] The target value calculation module is used to calculate the target value y for each training sample. The target value y is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is:

[0263] ;

[0264] in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when

[0265] Iterative update module, used to take the target value y as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, where the iterative update process is:

[0266] ;

[0267] β is the learning rate, which means Q(s , a) Update step size;

[0268] The loss calculation module is used to use the mean square error as the loss function during the training process. The calculation formula of the loss function is:

[0269] ;

[0270] Where N represents the number of mini-batch samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.

[0271] Based on the same inventive concept, another embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the satellite-to-ground data transmission method as described in any of the above embodiments.

[0272] Among them, electronic equipment refers to Figure 3 , Figure 3 Schematic diagram of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 300 includes: a memory 310 and a processor 320. The memory 310 and the processor 320 are connected via a bus communication. The memory 310 stores a computer program, which can be run on the processor 320 to implement the steps in the satellite-to-ground data transmission method disclosed in the above embodiment of the present application.

[0273] Based on the same inventive concept, another embodiment of the present application further provides a computer program product, including a computer program, which is used by a processor to execute the steps of the satellite-to-ground data transmission method described in any of the above embodiments.

[0274] Based on the same inventive concept, another embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein when the program is executed by a processor, the steps in the satellite-to-ground data transmission method as described in any of the above embodiments are implemented.

[0275] As for the device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0276] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0277] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0278] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0279] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.

[0280] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0281] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0282] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0283] The above is a detailed introduction to a satellite-to-ground data transmission method, device, equipment, and medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present application.

Claims

1. A satellite-to-ground data transmission method, characterized in that: The method comprises: Collecting historical spatiotemporal data of the satellite, the historical spatiotemporal data including: historical GPS position information, IMU acceleration data and antenna direction angle data of the satellite; Based on the historical spatiotemporal data, an LSTM network is used to predict the future spatiotemporal data of the satellite; Generate a future state vector based on the future spatiotemporal data; the future state vector includes: the remaining time of a specified time window, the remaining data volume of a data packet, a channel gain, a bandwidth resource allocated at a current moment, a transmit power allocated at a current moment, and a noise transmit power spectral density at a current location of the satellite; Inputting the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples; the action refers to the resource allocation plan for the satellite in the state, the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite in the state, and the next state refers to the state of the satellite after executing the resource allocation plan for the satellite in the state; When the resource allocation plan satisfies a constraint condition, applying the resource allocation plan to a future time instant of the satellite, wherein the constraint condition is used to constrain communication delay, bandwidth resources, and transmit power control of the satellite during data transmission; Before training the resource allocation model, the method further includes: Set the reward function, which is: ; ; Where W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite; The derivation process of the factors affecting energy consumption W(t,s) is as follows: According to the transmission power Determine W(t,s), the expression is: ; in, is the time window, is the transmission power, t is the current time; According to the communication time, the formula Calculation shows that D is the amount of data transmitted, C is the channel capacity, and ; Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained; in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel; N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource; Determine .

2. The satellite-to-ground data transmission method according to claim 1, wherein: The method further comprises: When the resource allocation plan does not meet the constraint conditions, readjusting the resource allocation plan; The communication delay is constrained as follows: the total amount of data transmitted by the satellite Must be within the specified time window The communication delay constraint expression is: ; The bandwidth resource is constrained to represent: the bandwidth resource used at the current time t Must be within the available bandwidth resources The constraint expression of the bandwidth resource is: ; The transmit power control is constrained to indicate that the transmit power used at the current time t is The transmission power of the satellite and ground station must be within the range The constraint expression of the transmit power control is: ; in, Indicates the total amount of data transmitted by the satellite time window, is the transmission rate, which is used to indicate the amount of data actually transmitted at the current time t.

3. The satellite-to-ground data transmission method according to claim 2, wherein: Before training the resource allocation model, the method further includes: Acquire real time and space data of each position of the satellite in orbit; the real time and space data includes: the real GPS position information of the satellite, the real IMU acceleration data and the real antenna direction angle data; According to the real antenna direction angle data, the phase change of the channel is obtained ; Calculate dynamic distance using the real GPS location information ; ; in, 、 、 are the three-dimensional coordinates of the satellites respectively; 、 、 are the three-dimensional coordinates of the ground stations respectively; According to the dynamic distance , fixed gain constant , path loss index n, attenuation factor empirical constant , the phase change of the channel ,wavelength , determine the channel gain ; The channel gain The calculation formula is: ; According to the total amount of data transmitted and the transmission rate R(t), based on the formula , determine the remaining amount of transmission data ; According to the time window Total time and the current time t, based on the formula , determine the remaining time of the time window ; According to the real GPS location information, determine the current location noise emission power spectrum density ; According to the remaining time , the remaining amount of the transmission data , the channel gain , the current position noise emission power spectrum density , and the bandwidth resources allocated at the current moment and transmit power , determine the state vector s(t), s(t) through the multidimensional vector It indicates that the multiple state vectors s(t) constitute a state space.

4. The satellite-to-ground data transmission method according to claim 3, wherein: Before training the resource allocation model, the method further includes: The transmit power is discretely enumerated at the first interval to obtain multiple discrete transmit powers ; Discretely enumerate the allocatable bandwidth resource range at the second interval to obtain multiple discrete bandwidth resources ; The discrete transmit power and the discrete bandwidth resources Combine to get the action space; The action space includes: multiple discrete transmission powers and multiple discrete bandwidth resources Composed of multiple actions .

5. The satellite-to-ground data transmission method according to claim 4, characterized in that: Before training the resource allocation model, the method further includes: Define the deep reinforcement learning network structure of the resource allocation model, wherein the input layer of the deep reinforcement learning network structure is the state vector s(t); the output layer is the action in the action space ; Construct an experience replay pool, which is used to store multiple training samples, each training sample is a tuple , the tuple includes the satellite's state s, action a, reward r and next state The training sample is determined based on the state vector s(t) in the state space, the action in the action space, the reward obtained by performing the action, and the next state after performing the action.

6. The satellite-to-ground data transmission method according to claim 5, characterized in that: The resource allocation model training process is as follows: Extract small batch training samples from multiple training samples in the experience replay pool, and train the resource allocation model to be trained through the small batch training samples; at each time step, use the epsilon-greedy strategy to select the action in the action space: An action is randomly selected from the action space with a probability of The probability of choosing the current action from the action space The action with the largest value ; Setting the objective function , It is used to represent the expected cumulative reward of the satellite after taking action a in state s after the resource allocation model selects action a; in, ; ; Wherein, E represents the computational expectation; s represents the current state of the satellite; a represents the action selected by the satellite in state s; is the discount factor; represents the initial state of the satellite in the deep reinforcement learning process; Indicates that the satellite is in the initial state The first action taken ; Indicates the total amount of data transmitted by the satellite; For each training sample, calculate the target value y, which is the sum of the immediate reward after taking the action a in state s and the maximum expected return of taking the optimal action in the next state. The calculation formula of the target value y is: ; in, is a discount factor used to measure the impact of future rewards; Indicates that from the next state Start by taking the best action The maximum expected return when The target value y is taken as The goal of iterative updating is to train the resource allocation model so that Close to the actual action value function, where the iterative update process is: ; β is the learning rate, which means Q(s , a) Update step size; During the training process, the mean square error is used as the loss function, and the calculation formula of the loss function is: ; Wherein, N represents the number of mini-batch training samples, are the current parameters of the resource allocation model, Indicates the state of the resource allocation model based on the current parameters in the jth training sample Take the action in the jth training sample The expected cumulative reward after Represents the state in the jth training sample Take the action in the jth training sample Immediate rewards and status after The sum of the maximum expected rewards of taking the best action in the next state.

7. A satellite-to-ground data transmission device, characterized in that: The device comprises: An acquisition module is used to acquire historical spatiotemporal data of the satellite, wherein the historical spatiotemporal data includes: historical GPS position information, IMU acceleration data, and antenna direction angle data of the satellite; A prediction module is used to predict the future spatiotemporal data of the satellite using an LSTM network based on the historical spatiotemporal data; a future state vector generation module, configured to generate a future state vector based on the future spatiotemporal data; the future state vector including: the remaining time of a specified time window, the remaining amount of data in a data packet, a channel gain, bandwidth resources allocated at the current moment, transmit power allocated at the current moment, and a noise transmit power spectral density at the current location of the satellite; An execution module is configured to input the future state vector into a pre-trained resource allocation model to obtain a resource allocation plan output by the resource allocation model; the resource allocation model is trained based on a deep reinforcement learning algorithm using the satellite's historical state, action, reward, and next state as training samples; the action refers to the resource allocation plan for the satellite in the state; the reward refers to the immediate reward after the satellite executes the resource allocation plan for the satellite in the state; and the next state refers to the state of the satellite after the resource allocation plan for the satellite in the state is executed; an application module, configured to apply the resource allocation plan to a future moment of the satellite when the resource allocation plan satisfies a constraint condition, wherein the constraint condition is used to constrain communication delay, bandwidth resources, and transmit power control of the satellite during data transmission; Optionally, the device comprises: The reward function setting module is used to set the reward function before training the resource allocation model. The reward function is: ; ; Where W(t,s) represents the energy consumption of the satellite when performing action a in state s; Penalty represents the penalty term; is the penalty factor; is the total amount of data transmitted by the satellite; The derivation process of the factors affecting energy consumption W(t,s) is as follows: According to the transmission power Determine W(t,s), the expression is: ; in, is the time window, is the transmission power, t is the current time; Determine module, used to determine the communication time according to the formula Calculation shows that D is the amount of data transmitted, C is the channel capacity, and ; Where S is the signal power, which is given by the formula express, is the channel gain, channel gain According to the formula obtained; in, is a fixed gain constant, is the empirical constant of the attenuation factor that varies with time and space, is the dynamic distance between the satellite and the ground station, n is the path loss exponent, is the phase change of the channel; Influencing factor determination module, For N is the noise power, given by express, is the noise power spectrum density, B is the bandwidth resource; Determine .

8. An electronic device, characterized in that: The system comprises a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the satellite-to-ground data transmission method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, wherein when the computer program is executed by a processor, the satellite-to-ground data transmission method according to any one of claims 1 to 6 is implemented.