Communication satellite handover model training method and satellite communication connection method
By training a communication satellite handover evaluation network using a dual-network model, and selecting the switching satellite based on channel load and energy loss, the problems of excessive satellite load and high energy consumption are solved, thereby improving communication quality and energy efficiency.
Patent Information
- Application Number
- CN202510214256.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional satellite handover strategies and algorithms can easily lead to excessive satellite load, degraded communication quality, and excessive energy consumption of terminal devices during frequent handovers.
By acquiring satellite status information from terminal devices and satellite status information after handover, the channel load and energy loss are calculated. A communication satellite handover evaluation network is trained using a dual-network model, and satellites with less load and lower energy loss are selected for handover.
It effectively reduces satellite channel load, lowers energy consumption, and improves communication quality and energy efficiency of terminal equipment.
Smart Images

Figure CN120150787B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication satellite handover technology, specifically to a communication satellite handover model training method and a satellite communication connection method. Background Technology
[0002] Low Earth Orbit (LEO) satellite networks, as a non-terrestrial communication platform, offer advantages over Medium Earth Orbit (MEO) and Geostationary Orbit (GEO) satellite networks, such as lower latency, stronger signal strength, and more economical deployment costs. However, due to the rapid and periodic movement of LEO satellites, their Earth coverage area is constantly changing, making it difficult to maintain long-term connections between satellites and user terminal devices. This necessitates frequent satellite handovers to ensure continuous communication over extended periods. Traditional satellite handover strategies, such as single-attribute decision-making methods, rely on a single attribute factor as the decision criterion. However, single-attribute handover algorithms are prone to overloading the satellite with the best attribute, leading to excessive satellite load. When the satellite channel is overloaded, communication quality degrades, requiring terminal devices to switch satellites again for reliable communication. Furthermore, the high energy consumption required during handover results in excessive power consumption for the terminal devices. Summary of the Invention
[0003] The purpose of this application is to overcome the shortcomings and deficiencies in the existing technology and to provide a communication satellite handover model training method and a satellite communication connection method, which can...
[0004] The first aspect of this application provides a method for training a communication satellite handover model, including:
[0005] Acquire several first satellite status information corresponding to the terminal device and several second satellite status information corresponding to the terminal device after it performs the action of switching satellite communication;
[0006] From the aforementioned second satellite status information, obtain the number of second channel loads and energy consumption of the second satellite communicating after the handover;
[0007] If the number of second channel loads is greater than the preset channel load threshold, the preset negative reward will be determined as the reward for the action of switching satellite communication.
[0008] If the number of loads on the second channel is less than or equal to the channel load threshold, an action reward for switching satellite communication is obtained based on a preset switching communication duration value, the data transmission rate between the terminal device and the second satellite, the number of loads on the second channel, and the switching time and energy consumption of the terminal device when switching satellite communication.
[0009] Through the first network model, a first action score corresponding to several first satellite status information and the action of switching satellite communication is obtained, as well as a second action score corresponding to several second satellite status information.
[0010] From the several second action scores output by the first network model, the action corresponding to the highest second action score, and the corresponding several second satellite state information are input into the second network model to obtain the third action score output by the second network model.
[0011] Based on the third action score and the action reward, the target action score is obtained;
[0012] The first network model is trained based on the first action score and the target action score;
[0013] The first network model after training was determined as the communication satellite handover evaluation network.
[0014] The second aspect of this application provides a satellite communication connection method for obtaining a communication satellite handover evaluation network trained according to the communication satellite handover model training method described above;
[0015] The real-time satellite status information corresponding to the terminal device is input into the communication satellite handover evaluation network to obtain the strategy action corresponding to the highest action score.
[0016] The terminal device is driven to execute the policy action to connect to the satellite targeted by the policy action.
[0017] Compared to existing technologies, this application calculates the corresponding action reward based on the number of second channel loads after the action of switching satellite communication and the energy loss of performing the satellite switching. It then trains the first network model by combining the third action score output by the first and second network models to obtain a communication satellite handover evaluation network. The third action score is a score obtained by the second network model from the action with the highest second action score predicted by the first network model. Therefore, the third action score is a comprehensive score that combines the model characteristics of the first and second network models. The action reward is obtained based on the number of second channel loads and energy loss. Therefore, the communication satellite handover evaluation network trained based on the third action score, action reward, and the first action score output by the first network model can output an action score by combining the number of second channel loads and energy loss. This is beneficial for selecting actions with fewer channel loads and lower energy loss for satellite communication based on the multiple satellite status information corresponding to the terminal device.
[0018] To provide a clearer understanding of this application, the specific embodiments of this application will be described below in conjunction with the accompanying drawings. Attached Figure Description
[0019] Figure 1 This is a step diagram of a communication satellite handover model training method according to an embodiment of this application.
[0020] Figure 2 This is a flowchart illustrating a communication satellite handover model training method according to an embodiment of this application.
[0021] Figure 3 This is a step diagram of a satellite communication connection method according to an embodiment of this application.
[0022] Figure 4 This is a schematic diagram of step S200 of a satellite communication connection method according to an embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0024] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.
[0025] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. The singular forms "a," "the," and "the" used in this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. The word "if" as used herein can be interpreted as "when," "when," or "in response to determination."
[0026] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0027] Please see Figure 1This is a flowchart of a communication satellite handover model training method according to an embodiment of this application, including:
[0028] S1: Obtain several first satellite status information corresponding to the terminal device and several second satellite status information corresponding to the terminal device after it performs the action of switching satellite communication.
[0029] In this context, "terminal device" refers to an electronic device with wireless communication capabilities, such as a smartphone. Before switching satellite communication, the terminal device connects to the first satellite; after switching, it connects to the second satellite. Specifically, when the terminal device receives a satellite switching signal from the base station, it obtains the satellite ephemeris information and satellite ID required for the switching. The terminal device converts the received ephemeris information into a geocentric coordinate system to obtain satellite coordinates. Then, based on the terminal device's coordinates and those of each satellite, it constructs a communication system model to obtain data such as the candidate satellite set, the reference signal received signal strength set, the remaining visible time set, and the load channel set, thus obtaining several satellite status information entries. The first satellite status information entries correspond to the satellite status information before the switching, while the second satellite status information entries correspond to the satellite status information after the switching. The satellite status information corresponding to the terminal device can be represented as follows:
[0030]
[0031] in, This refers to the satellite status information of u satellites corresponding to the terminal device at time t, where t is the time and u is the total number of satellites. Let u be the u-th satellite corresponding to the terminal device at time t. Let be the reference signal received signal strength value of the terminal device at time t corresponding to the u-th satellite. Let l be the remaining visible time between the terminal device and the u-th satellite at time t. t Let be the number of channel loads for the satellite at time t.
[0032] Among them, the first satellite status information corresponding to the terminal device refers to the satellite status of the several satellites before the terminal device switches satellite communication; the second satellite status information refers to the satellite status of the several satellites after the terminal device performs the action of switching satellite communication.
[0033] S2: Obtain the number of second channel loads and energy consumption of the second satellite after the handover from the aforementioned second satellite status information.
[0034] The number of second channel loads of the second satellite after the handover can be directly obtained from the status information of several second satellites. The number of second channel loads is the number of load channels of the satellite currently connected to the terminal device after the handover. The energy loss of the satellite handover can be obtained from the handover power and handover time of the terminal device when handover satellite communication.
[0035] The terminal device switches satellite communication within a switching frame. The switching frame has two phases: the first phase is the preparation phase, in which the signaling for whether to switch satellites and the signaling of the information required for switching satellites are transmitted; the second phase is the transmission phase, in which the service data is transmitted.
[0036] A handover frame must be located at the beginning of each time slot, and the connection link between the user and the satellite is considered unchanged within each time slot. Within the current handover frame, the terminal device can transmit the signaling information required for satellite handover multiple times until successful transmission, completing the handover action. If the terminal device has not completed the satellite handover action by the end of the current handover frame, the handover is considered a failure, and the signaling information required for satellite handover must be transmitted again in the next handover frame. After the terminal device completes the satellite handover action within the current handover frame, the remaining time of the handover frame is used for service data transmission. If the terminal device does not need to handover the satellite within the current handover frame, the entire current handover frame is used for service data transmission. Therefore, handover time refers to the time taken for the terminal device to complete the satellite handover action within the current handover frame.
[0037] S3: If the number of second channel loads is greater than the preset channel load threshold, the preset negative reward will be determined as the action reward for switching satellite communication.
[0038] The channel load threshold is the total number of channels used for communication by the corresponding satellite. If the number of channels loaded in the second channel is greater than the preset channel load threshold, it will cause excessive channel load on the satellite, increasing the communication pressure on the satellite and reducing the communication quality of the satellite. Therefore, the reward for the corresponding satellite switching action is a negative reward, which can be -50, -100, -120, etc.
[0039] S4: If the number of second channel loads is less than or equal to the channel load threshold, obtain the action reward for switching satellite communication based on the preset switching communication duration value, the data transmission rate between the terminal device and the second satellite, the number of second channel loads, and the switching time and energy consumption of the terminal device when switching satellite communication.
[0040] Specifically, when the number of second channel loads is less than or equal to the channel load threshold, it indicates that the satellite's communication pressure is within its capacity. In this case, the reward for switching satellite communication can be calculated based on the switching communication duration, data transmission rate, number of second channel loads, and the switching time and energy consumption of the terminal device. The switching communication duration is the duration of the switching frame.
[0041] The action reward is calculated based on the number of second channel loads after the action of switching satellite communication and the energy loss of performing the satellite switching. The calculated action reward combines the information of the number of second channel loads and energy loss, and can be used to guide the training direction of the network model.
[0042] S5: Using the first network model, obtain a first action score corresponding to several first satellite status information and the action of switching satellite communication, and a second action score corresponding to several second satellite status information.
[0043] like Figure 2 As shown, Figure 2 The estimated network is the first network model of this application. Figure 2 The target network is the second network model of this application, the environment is satellite state information, and the experience pool is a training data storage space used to store the first satellite state information, the action of switching satellite communication, the second satellite state information, and the action reward. Several pieces of the first satellite state information can be stored in the pool. The action of switching satellite communication Several second satellite status information and the action reward R t As a group The experience is saved to the experience pool. Then, several sets of experience are selected from the experience pool. The first satellite status information and the action of switching satellite communication in each set of experience are input into the first network model to obtain the first action score output by the first network model. Then, the second satellite status information is input into the first network model to obtain the second action score output by the first network model.
[0044] S6: From the several second action scores output by the first network model, input the action corresponding to the highest second action score and the corresponding several second satellite state information into the second network model to obtain the third action score output by the second network model.
[0045] The network parameters of the second network model are updated according to a preset time period. The update process uses the network parameters of the first network model as the new network parameters of the second network model. Therefore, the second network model can score the switching satellite communication action with the highest second action score output by the first network model by updating the network parameters with a longer update period, so as to obtain the third action score.
[0046] S7: Obtain the target action score based on the third action score and the action reward.
[0047] The sum of the third action score after discounting and the action reward is the target action score.
[0048] S8: Train the first network model based on the first action score and the target action score.
[0049] Training the first network model based on the first action score and the target action score can make the output of the first network model after training closer to the target action score.
[0050] S9: The first network model after training is determined as the communication satellite handover evaluation network.
[0051] Compared to existing technologies, this application calculates the corresponding action reward based on the number of second channel loads after the action of switching satellite communication and the energy loss of performing the satellite switching. It then trains the first network model using the third action score output by the first and second network models to obtain a communication satellite handover evaluation network. The third action score is a score predicted by the second network model based on the action with the highest second action score predicted by the first network model. Therefore, the third action score is a comprehensive score combining the model characteristics of both the first and second network models. The action reward is obtained based on the number of second channel loads and energy loss. Thus, the communication satellite handover evaluation network trained based on the third action score and action reward can output action scores based on the number of second channel loads and energy loss, which is beneficial for selecting actions with fewer channel loads and lower energy loss for satellite communication based on the multiple satellite status information corresponding to the terminal device. Furthermore, since the training process incorporates the action scoring of the second network model with a longer network parameter update cycle, it avoids the instability and singular parameter training direction caused by the first network model's excessively fast training speed.
[0052] In a feasible embodiment, step S4: obtaining the action reward for switching satellite communication based on a preset switching communication duration value, the data transmission rate between the terminal device and the second satellite, the load of the second channel, and the switching time and energy consumption of the terminal device switching satellite communication, includes:
[0053] S41: Based on the switching communication duration value, the switching time, and the data transmission rate between the terminal device and the second satellite, obtain the second maximum communication data value.
[0054] The maximum value of the second communication data can be obtained using the following formula:
[0055] D = C t ×(T t -T′ t )
[0056] In the above formula, D is the maximum value of the second communication data, and C t T represents the data transmission rate. t To switch the communication duration value, T′ t This refers to the switching time.
[0057] S42: Obtain the second channel load impact value based on the preset amplification factor and the second channel load quantity.
[0058] The second channel load impact value is obtained using the following formula:
[0059] M = η × l t
[0060] M is the second channel load impact value, η is the preset amplification factor, and l t The number of channel loads for the satellites connected for communication is specified in step S42, which specifies the number of second channel loads. The amplification factor can be 30, 40, 50, etc.
[0061] S43: Obtain the action reward based on the maximum value of the second communication data, the energy loss, and the impact value of the second channel load.
[0062] The reward for the action is obtained using the following formula:
[0063] R t =DEM
[0064] R t E represents the reward for the action, and E represents the energy loss.
[0065] In this embodiment, the maximum value of the second communication data is calculated by switching the communication duration value, the switching time, and the data transmission rate between the terminal device and the second satellite. After calculating the impact value of the second channel load based on the preset amplification factor and the number of second channel loads, the corresponding action reward is obtained more comprehensively and accurately by combining the energy loss of switching satellite communication.
[0066] In a feasible embodiment, the energy loss of the terminal device switching satellite communication is obtained through the following steps:
[0067] S401: Obtain the switching power of the terminal device when switching satellite communication.
[0068] The switching power of the terminal device when switching satellite communication can be obtained by testing the terminal device multiple times when switching satellite communication, or it can be obtained from the device parameters provided by the manufacturer.
[0069] S402: The energy loss is obtained based on the switching power and the switching time.
[0070] The energy loss is obtained through the following formula:
[0071] E = T' t ×P t
[0072] E represents energy loss, P t For the action power; T′ t This refers to the switching time.
[0073] In a feasible embodiment, before step S6: inputting the action corresponding to the highest second action score from among several second action scores output by the first network model, along with several corresponding second satellite state information, into the second network model to obtain the third action score output by the second network model, the method further includes:
[0074] S601: Obtain the action reward for not switching satellite communication based on the first channel load of the first satellite connected to the terminal device when the satellite communication is not switched, the data transmission rate, and the preset switching communication duration value.
[0075] Since there is no energy loss associated with switching satellite communications, the reward for not switching satellite communications can be calculated based on the first channel load of the first satellite connected to the terminal device, the data transmission rate, and the preset switching communication duration.
[0076] S602: Input the action of not switching satellite communication and the corresponding number of first satellite status information into the first network model to obtain the first action score and the second action score output by the first network model.
[0077] In this embodiment, by obtaining the action reward for not switching satellite communication and inputting the corresponding action and several first satellite state information into the first network model, the second action score for not switching satellite communication output by the first network model can be combined with the situation of not switching satellite communication, so as to obtain a more comprehensive second action score for switching satellite communication and not switching satellite communication.
[0078] In a feasible embodiment, step S601, which involves obtaining the action reward for not switching satellite communication based on the first channel load of the first satellite connected to the terminal device without switching satellite communication, the data transmission rate, and a preset switching communication duration value, includes:
[0079] S6011: Based on the switching communication duration value and the data transmission rate between the terminal device and the first satellite, obtain the first maximum value of communication data between the terminal device and the first satellite.
[0080] D′=C t ×T t
[0081] D′ is the maximum value of the first communication data, C t T represents the data transmission rate. t To switch the communication duration value.
[0082] S6012: Obtain the first channel load impact value based on the preset amplification factor and the first channel load quantity.
[0083] M′=η×l t
[0084] M′ is the first channel load impact value, η is the preset amplification factor, and l t The number of channel loads for the satellites connected for communication is specified in step S42, which represents the number of loads for the first channel. The amplification factor can be 30, 40, 50, etc.
[0085] S6013: Obtain the action reward based on the maximum value of the first communication data and the first channel load impact value.
[0086] R t =D′-M′
[0087] R t The action is rewarded.
[0088] Based on the above, the reward function for the actions of switching satellite communication and not switching satellite communication by the terminal device can be represented by the following reward function:
[0089]
[0090] Among them, R t E represents the action reward, and C represents the energy loss. t T represents the data transmission rate. t To switch the communication duration value, T′ t The switching time is η, where η is the preset amplification factor, and l t The number of channels for satellites connected for communication.
[0091] In a feasible embodiment, step S8: training the first network model based on the first action score and the target action score includes:
[0092] S81: Construct a target function based on the first action score and the target action score, and obtain the first function output of the target function.
[0093] Specifically, the absolute value operation can be performed on the difference function between the first action score and the target action score to obtain the target function, and the first function output of the target function can be obtained through the following formula:
[0094] δ=|Q(s t ,a t )-y t |=|Q(s t ,a t ,θ)-(R t +γQ′(s t+1 argmax a Q(s t+1 ,a t ;θ);θ′))|
[0095] Where δ is the output of the first function, Q(s) t ,a t ) is the score for the first action, y t Score the target action, s t Here, θ represents the state information of the first satellite, θ represents the network parameters of the first network model, and a represents the network parameters of the first network model. t For the action, R t The reward is γ, the discount rate is argmax. a Q(s t+1 ,a t ;θ) represents the action corresponding to the highest second action score output by the first network model, Q′(s t+1 argmax a Q(s t+1 ,a t ;θ);θ′) are the scores for the third action.
[0096] S82: Update the network parameters of the first network model according to the output of the first function to obtain the trained first network model.
[0097] Since the first function output combines the first action score and the target action score, and the target action score is obtained by combining the channel load and energy loss of switching satellite communication and the channel load of not switching satellite channels, the comprehensiveness of training the first network model can be improved by using the first function output based on the first action score and the target action score.
[0098] In a feasible embodiment, step S82: updating the network parameters of the first network model according to the output of the first function to obtain the trained first network model includes:
[0099] S821: Determine the loss function based on the output of the first function, and obtain the second function output of the loss function.
[0100] The output of the second function is obtained through the following formula:
[0101]
[0102] Where loss is the output of the second function and δ is the output of the first function.
[0103] S822: Based on the output of the second function, the network parameters of the first network model are updated using the gradient descent algorithm to obtain a trained first network model where the output of the second function is less than or equal to a preset function threshold.
[0104] The function threshold is set by the user.
[0105] In this embodiment, by updating the network parameters of the first network model using the gradient descent algorithm, the updated output of the first network model can be made closer to the target action score.
[0106] In a feasible embodiment, after step S8: training the first network model based on the first action score and the target action score, the method further includes:
[0107] S83: Based on a preset time period, update the network parameters of the second network model according to the network parameters of the first network model.
[0108] In this embodiment, since the network parameters of the first network model are updated in real time as training progresses, while the network parameters of the second network model are updated according to a preset time period, the situation of unstable training and single-direction parameter training caused by the training speed of the first network model being too fast can be avoided, making the training of the first network model more comprehensive and stable.
[0109] Please see Figure 3 and 4 The second embodiment of this application provides a satellite communication connection method, including:
[0110] S100: Obtain the communication satellite handover evaluation network trained according to the communication satellite handover model training method described above.
[0111] S200: Input the real-time satellite status information corresponding to the terminal device into the communication satellite handover evaluation network to obtain the strategy action corresponding to the highest action score.
[0112] Please see Figure 4 Step S200 can be broken down into: Figure 4 The handover process involves four steps: information collection, information processing, handover decision, and handover execution. The real-time satellite status information corresponding to the terminal equipment can be obtained through... Figure 4 The handover information collection and processing steps shown are obtained, and step S200, which involves obtaining the highest action score through the communication satellite handover evaluation network, corresponds to... Figure 4 The steps shown in the switching decision process yield the strategy action corresponding to the highest action score. Figure 4 The steps for switching execution are shown.
[0113] S300: Drive the terminal device to execute the policy action to connect to the satellite targeted by the policy action.
[0114] It should be noted that the satellite communication connection method provided in the second embodiment of this application and the communication satellite switching model training method in the first embodiment of this application belong to the same concept. The implementation process is detailed in the first embodiment and will not be repeated here.
[0115] The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0116] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0117] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function selected in one or more boxes.
[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function selected in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0120] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0121] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0122] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0123] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for training a communication satellite handover model, characterized in that, include: Acquire several first satellite status information corresponding to the terminal device and several second satellite status information corresponding to the terminal device after the terminal device performs the action of switching satellite communication; wherein, the several first satellite status information corresponding to the terminal device refers to the satellite status of several satellites before the terminal device switches satellite communication; the several second satellite status information refers to the satellite status of several satellites after the terminal device performs the action of switching satellite communication. From the aforementioned second satellite status information, obtain the number of second channel loads and energy consumption of the second satellite communicating after the handover; If the number of second channel loads is greater than the preset channel load threshold, the preset negative reward will be determined as the reward for the action of switching satellite communication. If the number of loads on the second channel is less than or equal to the channel load threshold, an action reward for switching satellite communication is obtained based on a preset switching communication duration value, the data transmission rate between the terminal device and the second satellite, the number of loads on the second channel, and the switching time and energy consumption of the terminal device when switching satellite communication. Through the first network model, a first action score corresponding to several first satellite status information and the action of switching satellite communication is obtained, as well as a second action score corresponding to several second satellite status information. From the several second action scores output by the first network model, the action corresponding to the highest second action score, and the corresponding several second satellite state information are input into the second network model to obtain the third action score output by the second network model. Based on the third action score and the action reward, the target action score is obtained; The first network model is trained based on the first action score and the target action score; The first network model after training is identified as the communication satellite handover evaluation network; The step of obtaining the action reward for switching satellite communication based on a preset switching communication duration value, the data transmission rate between the terminal device and the second satellite, the load of the second channel, and the switching time and energy consumption of the terminal device switching satellite communication includes: The second maximum communication data value is obtained based on the switching communication duration value, the switching time, and the data transmission rate between the terminal device and the second satellite; The second channel load impact value is obtained based on the preset amplification factor and the number of second channel loads; The action reward is obtained based on the maximum value of the second communication data, the energy loss, and the impact value of the second channel load.
2. The communication satellite handover model training method according to claim 1, characterized in that, The energy loss of the terminal device when switching satellite communication is obtained through the following steps: Obtain the switching power of the terminal device when switching satellite communication; The energy loss is obtained based on the switching power and the switching time.
3. The communication satellite handover model training method according to claim 1, characterized in that, Before the step of inputting the action corresponding to the highest second action score from the plurality of second action scores output by the first network model, along with the corresponding plurality of second satellite state information, into the second network model to obtain the third action score output by the second network model, the method further includes: The reward for not switching satellite communication is obtained based on the first channel load of the first satellite connected to the terminal device when the terminal device does not switch satellite communication, the data transmission rate, and the preset switching communication duration value. The action of not switching satellite communication and the corresponding number of first satellite status information are input into the first network model to obtain the first action score and the second action score output by the first network model.
4. The communication satellite handover model training method according to claim 3, characterized in that, The step of obtaining the action reward for not switching satellite communication based on the first channel load of the first satellite connected to the terminal device without switching satellite communication, the data transmission rate, and a preset switching communication duration value includes: Based on the switching communication duration value and the data transmission rate between the terminal device and the first satellite, the first maximum value of communication data between the terminal device and the first satellite is obtained; The first channel load impact value is obtained based on the preset amplification factor and the first channel load quantity; The action reward is obtained based on the maximum value of the first communication data and the impact value of the first channel load.
5. The communication satellite handover model training method according to claim 1, characterized in that, The step of training the first network model based on the first action score and the target action score includes: Construct a target function based on the first action score and the target action score, and obtain the first function output of the target function; The network parameters of the first network model are updated based on the output of the first function to obtain the trained first network model.
6. The communication satellite handover model training method according to claim 5, characterized in that, The step of constructing a target function based on the first action score and the target action score, and obtaining the first function output of the target function, includes: The difference function between the first action score and the target action score is processed by absolute value operation to obtain the target function and the first function output of the target function.
7. The communication satellite handover model training method according to claim 5, characterized in that, The step of updating the network parameters of the first network model according to the output of the first function to obtain the trained first network model includes: Based on the output of the first function, the loss function is determined, and the second function output of the loss function is obtained; Based on the output of the second function, the network parameters of the first network model are updated using the gradient descent algorithm to obtain a trained first network model where the output of the second function is less than or equal to a preset function threshold.
8. The communication satellite handover model training method according to any one of claims 1-7, characterized in that, After the step of training the first network model based on the first action score and the target action score, the method further includes: Based on a preset time period, the network parameters of the second network model are updated according to the network parameters of the first network model.
9. A satellite communication connection method, characterized in that, include: Obtain the communication satellite handover evaluation network trained by the communication satellite handover model training method according to any one of claims 1-8; The real-time satellite status information corresponding to the terminal device is input into the communication satellite handover evaluation network to obtain the strategy action corresponding to the highest action score. The terminal device is driven to execute the policy action to connect to the satellite targeted by the policy action.
Citation Information
Patent Citations
Serial Q learning distributed switching method and system under large-scale LEO satellite network
CN114698045A
Low-orbit giant constellation satellite switching method and device based on deep reinforcement learning
CN118611728A