An Adaptive Routing Intelligent Decision-making Method for Multi-mode Satellite Terminals
Through the intelligent decision-making method of adaptive routing of multi-mode satellite terminals, the gradient enhancement tree model and intelligent decision-making algorithm are used to optimize resource allocation and path selection, and the energy consumption and communication quality problems of satellite-ground dual-mode communication terminals in complex network environments are solved, and efficient resource utilization and fault-tolerant processing are achieved.
Patent Information
- Application Number
- CN202510487918.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Existing satellite-ground dual-mode communication terminals are difficult to achieve efficient resource allocation and fault-tolerant processing in complex network environments, resulting in communication quality and terminal energy consumption problems. Traditional routing mechanisms cannot adjust resource allocation in time when facing dynamic network changes, resulting in link congestion or idle resources, and cannot quickly switch to backup paths, affecting communication continuity.
The intelligent decision-making method of adaptive routing of multi-mode satellite terminals is adopted to construct a network state prediction model based on gradient enhancement tree by collecting multi-source data, determine the optimal dynamic resource allocation scheme and link failure prediction probability, determine the candidate path based on business needs, and optimize the routing decision through the intelligent decision-making model and select the optimal path.
It realizes reducing terminal energy consumption while ensuring communication quality, enhances network fault tolerance and resource utilization efficiency, and improves the performance of communication terminals.
Smart Images

Figure CN120034249B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and specifically to a multi-mode satellite terminal power adaptive management method. Background Art
[0002] In today's communication networks, satellite-ground dual-mode communication terminals are widely used in fields such as agricultural production and emergency rescue. However, the complex network environment and limited terminal energy restrict the performance of these terminals. Traditional routing mechanisms have many deficiencies in dealing with network dynamic changes. For example, in a large-scale network, the controller configuration of software-defined network (SDN) is complex and the processing delay is high; the training cost of energy-saving routing strategies based on deep reinforcement learning is high and the convergence is slow; the traffic prediction accuracy and real-time performance of traffic-aware routing algorithms are poor; energy-aware routing algorithms are prone to falling into local optima and it is difficult to predict energy; the cross-layer energy-saving routing strategy has a large inter-layer coordination overhead and poor protocol adaptability. These problems make it difficult for existing routing mechanisms to achieve terminal energy saving while ensuring communication quality, severely restricting the performance improvement of communication terminals.
[0003] In addition, traditional routing mechanisms have serious deficiencies in resource allocation and fault tolerance processing. It is difficult to achieve efficient dynamic resource allocation and cannot quickly respond to network failures and ensure the continuity of communication. For example, when network traffic suddenly increases, traditional mechanisms cannot adjust resource allocation in a timely manner, resulting in congestion on some links while some link resources are idle; when a link fails, it cannot quickly switch to a backup path, causing data transmission interruption. Summary of the Invention
[0004] In order to overcome the above technical defects, the present application provides a multi-mode satellite terminal adaptive routing intelligent decision-making method. To achieve the above object, the present application is implemented according to the following technical solutions:
[0005] The present application provides a multi-mode satellite terminal adaptive routing intelligent decision-making method, including:
[0006] Collect multi-source data;
[0007] Based on the multi-source data, construct a network state prediction model based on gradient boosting trees;
[0008] Solve the network state prediction model to obtain a real-time network state evaluation value;
[0009] Based on the real-time network state evaluation value, determine an optimal dynamic resource allocation scheme and a link failure prediction probability;
[0010] Obtain service requirements;
[0011] Determine a first number of candidate paths based on the real-time evaluation value of the network state, the service requirements, the optimal dynamic resource allocation scheme, and the link failure prediction probability;
[0012] Obtain a predefined state space and a predefined action space;
[0013] Calculate the reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space;
[0014] Construct a terminal routing intelligent decision-making model based on the reward value, the predefined state space, and the predefined action space;
[0015] Solve the terminal routing intelligent decision-making model to determine the optimal path for transmission.
[0016] Optionally, the constructing a network state prediction model based on gradient boosting trees based on the multi-source data includes:
[0017] Use the principal component analysis method to screen and process the multi-source data to determine the main influencing features;
[0018] Construct a network state prediction model based on gradient boosting trees based on the main influencing features.
[0019] Optionally, the determining an optimal dynamic resource allocation scheme based on the real-time evaluation value of the network state includes:
[0020] Construct an objective function for maximizing the overall benefit of the network based on the real-time evaluation value of the network state;
[0021] Solve the objective function to determine the optimal dynamic resource allocation scheme.
[0022] Optionally, the determining a link failure prediction probability based on the real-time evaluation value of the network state includes:
[0023] Obtain the real-time state parameters of the network and the historical state parameters of the network;
[0024] Determine the link failure prediction probability based on the real-time evaluation value of the network state, the real-time state parameters of the network, and the historical state parameters of the network.
[0025] Optionally, the determining a first number of candidate paths based on the real-time evaluation value of the network state, the service requirements, the optimal dynamic resource allocation scheme, and the link failure prediction probability includes:
[0026] Determine a second number of candidate paths based on the service requirements;
[0027] Determine the terminal resource allocation vector based on the optimal dynamic resource allocation scheme;
[0028] Based on the real-time evaluation value of the network state, the link failure prediction probability, and the terminal resource allocation vector, determine the comprehensive cost function of each link of the second number of candidate paths;
[0029] Based on the comprehensive cost function of each link, determine the first number of candidate paths.
[0030] Optionally, the determining the first number of candidate paths based on the comprehensive cost function of each link includes:
[0031] Based on the comprehensive cost function of each link, determine the cumulative cost of each candidate path among the second number of candidate paths;
[0032] Screen in ascending order of the cumulative cost to determine the first number of candidate paths.
[0033] Optionally, it further includes determining an alternative path for each candidate path among the first number of candidate paths.
[0034] Optionally, the predefined state space includes the following:
[0035] The real-time evaluation value of the network state, candidate path characteristics, terminal energy consumption, service requirements;
[0036] The terminal energy consumption includes battery content and energy consumption rate;
[0037] The service requirements include data volume and real-time requirements.
[0038] Optionally, the calculating the reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space includes:
[0039] Based on the first number of candidate paths, the predefined state space, and the predefined action space, construct the reward function of each candidate path;
[0040] Solve the reward function to calculate the reward value corresponding to each candidate path.
[0041] The present application has the following beneficial effects:
[0042] The method proposed in the present application perceives the network state through multi-source data fusion, constructs an accurate evaluation model, optimizes the routing decision using intelligent algorithms, optimizes resource utilization, enhances the fault tolerance rate, and realizes the reduction of terminal energy consumption while ensuring the communication quality.
[0043] In addition to the purposes, features, and advantages described above, this application has other purposes, features, and advantages. The following will refer to the accompanying drawings for a further detailed description of this application. Description of the Drawings
[0044] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0045] Figure 1 is a schematic flow diagram of a multi-mode satellite terminal adaptive routing intelligent decision-making method provided by an embodiment of this application. Detailed Embodiment
[0046] The following will describe the embodiments of this application in detail with reference to the accompanying drawings. However, this application can be implemented in many different ways defined and covered by the claims.
[0047] To solve the problems raised in the background art, as Figure 1 shown, this application provides a multi-mode satellite terminal adaptive routing intelligent decision-making method, including:
[0048] Step S101: Collect multi-source data;
[0049] Based on the multi-source data, construct a network state prediction model based on a gradient boosting tree;
[0050] The multi-source data collection is realized by integrating a satellite communication module, a ground network monitoring module, and an environmental sensor in the terminal.
[0051] The satellite communication module integrated in the terminal collects real-time monitoring satellite communication signal parameters such as signal strength, signal-to-noise ratio, and bit error rate, and then calculates the signal strength according to the above parameters. The calculation formula is as follows:
[0052] (1)
[0053] In the formula, is the received signal strength, is the transmitted signal power, and are the transmitting and receiving antenna gains respectively, is the feeder loss, is the path loss.
[0054] The ground network monitoring module obtains parameters such as ground network signal strength, bandwidth utilization rate, and delay by means of the Link Layer Discovery Protocol (LLDP) or the IPv6 Neighbor Discovery Protocol (NDP). The bandwidth utilization rate calculation formula is:
[0055] (2)
[0056] Environmental sensors collect temperature ,humidity , air pressure Environmental data such as temperature, temperature and signal propagation loss are related to each other. For example, according to the empirical formula, the temperature-induced propagation loss formula is:
[0057] (3)
[0058] in, is the temperature-dependent propagation loss, is the communication link distance, is the signal operating frequency, is the temperature correction function, which is implemented by table lookup or polynomial fitting.
[0059] After collecting the above multi-source data, the principal component analysis (PCA) method is used to screen and process the multi-source data to determine the main influencing features. The main influencing features can be understood as the key features that affect the network status. The judgment method is to regard the features that exceed the first threshold as the main influencing features. In order to accurately estimate the network status, the first threshold is generally set to a relatively large value. The above process of determining the main influencing features is as follows:
[0060] For satellite signal strength, signal-to-noise ratio, bit error rate, link quality, bandwidth utilization, delay, wind speed, wind direction, altitude and other multi-source data, the principal component analysis method is used to map high-dimensional data to low-dimensional space by linear transformation, while retaining the variability of the data as much as possible. The multi-source data X is standardized, and the first k principal components with cumulative variance contribution rate ≥ 85% are retained. The standardized data is projected onto the principal components to obtain the reduced-dimensional data Y:
[0061] (4)
[0062] In the formula, is the eigenvalue of the ith principal component, p is the total number of original features, Is included before k The matrix of eigenvectors, k is the number of principal components selected, X′ is the standardized original data, and W is the number of principal components including the k The matrix of eigenvectors.
[0063] After determining the main influencing features, a network state prediction model based on Gradient Boosting Decision Tree (GBDT) is constructed according to the main influencing features. GBDT is an iterative decision tree algorithm that forms a strong learner by constructing multiple weak learners (decision trees). Its core idea is to gradually optimize the model, and each step of learning aims to correct the errors of the previous step, thereby improving the prediction accuracy of the model. The specific construction process is as follows:
[0064] Since the network state prediction model of GBDT requires an initial value, after determining the main influencing features, the initial prediction value is defined and calculated according to the main influencing features. The calculation and definition process is as follows:
[0065] (5)
[0066] Among them, the weight of the signal strength S is , and the maximum and minimum values are respectively ; the weight of the delay D is , and the maximum and minimum values are respectively ; the weight of the packet loss rate L is , and the maximum and minimum values are respectively ; the weight of the bit error rate E is , and the maximum and minimum values are respectively .
[0067] The weights sum to 1. Through this calculation method, the values of each index are mapped to the interval [0, 1], and are integrated according to their weights. The finally obtained value is also within the range of [0, 1]. The closer the value is to 1, the better the network state; the closer it is to 0, the worse the network state.
[0068] After determining the initial prediction value, use it as the mean c, Select as the target variable y , and obtain the initial GBDT model , that is:
[0069] (6)
[0070] Take the main influencing features as the feature vector of the model, and the process of iteratively constructing the network state prediction model based on GBDT is as follows:
[0071] ① Residual calculation , is the residual of the th sample in the th round.
[0072] ② Fit a new decision tree Prediction residual .
[0073] ③ Update the model, , where is the learning rate, which controls the pace of model update.
[0074] After M rounds of iteration, a network state prediction model based on gradient boosting decision tree is obtained:
[0075] (7)
[0076] Step S102: Solve the network state prediction model to obtain a real-time network state evaluation value;
[0077] Apply the trained GBDT model to new network state data, input the real-time feature vector , and the obtained is the real-time network state evaluation value , and the finally obtained value is in the range of [0, 1]. The closer the value is to 1, the better the network state; the closer it is to 0, the worse the network state.
[0078] Step S103: Based on the real-time network state evaluation value, determine the optimal dynamic resource allocation scheme and link failure prediction probability;
[0079] After determining the calculated real-time network state evaluation value, according to this real-time network state evaluation value, construct an objective function for maximizing the overall network benefit, and use the linear programming algorithm to achieve dynamic resource allocation. The specific process is as follows:
[0080] Suppose there are terminals in the network, and the resource requirement of each terminal is , and the total network resource is . Let represent the amount of resources allocated to the th terminal, including bandwidth, power, time slot, etc., where . To maximize the overall network benefit, let represent the benefit function of the th terminal when the allocated resource amount is and the network state evaluation result is . Here, a simple assumption is made:
[0081] (8)
[0082] That is, the better the network state, the higher the benefit brought by the same amount of resources, and the benefit decreases with the marginal benefit. Then the objective function for maximizing the overall network benefit is:
[0083] (9)
[0084] The constraints include the total resource constraint, the lower limit constraint of the terminal resource demand, etc. The total amount of resources allocated to all terminals cannot exceed the total network resources , and the amount of resources allocated to this terminal cannot be lower than this minimum resource demand , that is:
[0085] (10)
[0086] Solve the above objective function to determine the optimal dynamic resource allocation scheme, and thus determine the resource vector allocated to each terminal.
[0087] In addition, to ensure communication reliability and enhance the fault tolerance rate, perform fault prediction in advance, capture sudden changes in link quality in real time, and be able to switch to the backup path in time after a link failure to avoid service terminals and ensure communication reliability. Use the Bayesian inference model to combine the network historical state parameters and the network real-time state parameters for early fault prediction, and dynamically adjust the link fault prediction probability model according to the real-time evaluation value N of the network state.
[0088] Let the normal state of the network link be , and the fault state be . According to the detected variable E that has changed (such as delay, packet loss rate, bit error rate, etc.), calculate the posterior probability, that is, the fault prediction probability as follows:
[0089] (11)
[0090] Among them, and are the likelihood probabilities, obtained using the network real-time state parameters, and are the prior probabilities, obtained using the network historical parameters.
[0091] Dynamically adjust the prior probability according to the network state evaluation value N :
[0092] (12)
[0093] Among them, is to control the influence intensity of the network state on the fault probability, optimized according to the grid search method to minimize the area under the ROC curve , usually taking 0.8; is the basic failure rate, calibrated by historical statistical values:
[0094] (13)
[0095] Step S104: Obtain service requirements;
[0096] Determine a first number of candidate paths based on the real-time evaluation value of the network state, the service requirements, the optimal dynamic resource allocation scheme, and the link failure prediction probability;
[0097] At this time, obtain specific service requirements, then confirm the source node and the destination node, and then correspondingly determine a second number of candidate paths between the source node and the destination node.
[0098] After determining the second number of candidate paths, then improve the Dijkstra algorithm by considering multi-dimensional parameters. According to the network state N, the terminal resource allocation vector , the link failure probability , define the link cost functions of each of the second number of candidate paths:
[0099] (14)
[0100] Wherein, is the source node, is the destination node, is the network state of the destination node, used to reflect the overall reliability. is the link bandwidth, provided by the terminal resource allocation vector , is the link delay, is the link historical failure probability. The weight coefficient is dynamically adjusted by the service, is the reliability weight, is the bandwidth weight, is the delay weight, is the failure probability weight.
[0101] After calculating the link cost functions of each, then calculate the cumulative cost of each path in the second number of candidate paths:
[0102] (15)
[0103] After calculating the cumulative cost of each path, screen them in ascending order of the cumulative cost to determine the first number of candidate paths. That is to say, the value of the second number is larger than the value of the first number.
[0104] To ensure that the entire network can operate normally after detecting a link failure, determine an alternative path for each path in the first number of candidate paths. The specific process is as follows:
[0105] Similarly, the improved Dijkstra algorithm is adopted, and multi-dimensional parameters are comprehensively planned to link relevant paths, and the optimal backup path is selected for each candidate path within the entire network. During the planning process, the selection of the backup path considers the network state N, the bit error rate E, and the main path repeatability penalty term , the link failure probability , and the comprehensive cost function is defined as:
[0106] (16)
[0107] where the penalty term , is the penalty coefficient.
[0108] To ensure the reliability of the backup path, the failure probability weight is increased and set to 0.8 to avoid the risk of secondary failures. Calculate the cumulative cost, and select the path with the minimum cumulative cost as the alternative path.
[0109] Step S105: Obtain the predefined state space and the predefined action space;
[0110] Based on the first number of candidate paths, the predefined state space, and the predefined action space, calculate the reward value corresponding to each candidate path;
[0111] At this time, the predefined state space and the predefined action space are obtained. The predefined state space covers multi-faceted information such as the network state, candidate path characteristics, terminal energy consumption, and service requirements. The network state is the real-time evaluation value N of the network state; the candidate path characteristic is the minimum cumulative cost value of the candidate path ; the terminal energy consumption includes the battery power , the energy consumption rate ; the service requirements include the data volume , the real-time requirement . Therefore, the state can be represented as a vector:
[0112] (17)
[0113] The predefined action space is the set of optional routing paths. Each action represents the routing selection from the current node to the next node.
[0114] According to the first number of candidate paths, the predefined state space, and the predefined action space, construct the reward function for each candidate path. The reward function comprehensively considers the transmission delay, energy consumption, and reliability. The specific construction process is as follows:
[0115] (18)
[0116] Among them, delay is the actual delay, max_delay is the maximum delay. energy_consumed is the actual energy consumption, and max_energy is the maximum energy consumption. success_count is the number of successfully transmitted data packets, and total_count is the total number of transmitted data packets. w1, w2, and w3 are weight coefficients, satisfying w1 + w2 + w3 = 1.
[0117] To obtain the optimal solution of the above weight coefficients, the following process is adopted:
[0118] Through the weight allocation method, comprehensive optimization is carried out on multiple objectives such as transmission delay D, energy consumption E, and reliability R. The multi-objective optimization problem is expressed as:
[0119] (19)
[0120] Among them, x is the decision variable, is the i-th objective function.
[0121] By assigning weights to each objective , after transforming the multi-objective optimization problem into a single-objective optimization problem, it is then solved:
[0122] (20)
[0123] By solving the objective function of Equation (20), the specific values of w1, w2, and w3 are obtained, so as to solve the reward function of Equation (18) and calculate the reward value corresponding to each candidate path.
[0124] Step S106: Based on the reward value, the predefined state space, and the predefined action space, construct a terminal routing intelligent decision-making model;
[0125] Solve the terminal routing intelligent decision-making model to determine the optimal transmission path.
[0126] After determining the reward value, according to the reward value, the predefined state space, and the predefined action space, establish a terminal routing intelligent decision-making model as follows:
[0127] (21)
[0128] Among them, s is the current state, a is the action taken. r is the reward obtained after taking action a. s′ is the new state. a′ is any action that can be taken from the new state s′. α is the learning rate (0 < α < 1), which determines the influence degree of new information on the update of the Q value. The closer α is to 1, the greater the influence of new information. is the discount factor (0 ≤ γ < 1), which reflects the importance the agent attaches to future rewards. The closer it is to 1, the more it values future rewards.
[0129] At each decision-making moment, the terminal selects the routing path with the largest Q value in the action space according to the current state as the optimal path for the current transmission, and continuously updates the Q value and optimizes the decision-making model as the network state changes and data is continuously transmitted.
[0130] In summary, the method proposed in this application perceives the network state through multi-source data fusion, constructs an accurate evaluation model, optimizes the routing decision using intelligent algorithms, optimizes resource utilization, enhances the fault tolerance rate, and realizes the reduction of terminal energy consumption while ensuring the communication quality.
[0131] The above are only the preferred embodiments of this application and are not intended to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. An intelligent decision-making method for adaptive routing of a multi-mode satellite terminal, characterized in that Including: Collecting multi-source data; Based on the multi-source data, constructing a network state prediction model based on gradient boosting trees; Solving the network state prediction model to obtain a real-time network state evaluation value; Based on the real-time network state evaluation value, determining an optimal dynamic resource allocation scheme and a link failure prediction probability; Obtaining service requirements; Based on the real-time network state evaluation value, the service requirements, the optimal dynamic resource allocation scheme, and the link failure prediction probability, determining a first number of candidate paths; Obtaining a predefined state space and a predefined action space; Based on the first number of candidate paths, the predefined state space, and the predefined action space, calculating a reward value corresponding to each candidate path; Based on the reward value, the predefined state space, and the predefined action space, constructing a terminal routing intelligent decision-making model; Solving the terminal routing intelligent decision-making model to determine an optimal transmission path; Wherein, the determining the optimal dynamic resource allocation scheme based on the real-time network state evaluation value includes: Based on the real-time network state evaluation value, constructing an objective function for maximizing the overall network benefit; Solving the objective function to determine the optimal dynamic resource allocation scheme; The determining the link failure prediction probability based on the real-time network state evaluation value includes: Obtaining real-time network state parameters and historical network state parameters; Based on the real-time network state evaluation value, the real-time network state parameters, and the historical network state parameters, determining the link failure prediction probability; The determining the first number of candidate paths based on the real-time network state evaluation value, the service requirements, the optimal dynamic resource allocation scheme, and the link failure prediction probability includes: Based on the service requirements, determining a second number of candidate paths; Based on the optimal dynamic resource allocation scheme, determining a terminal resource allocation vector; Based on the real-time network state evaluation value, the link failure prediction probability, and the terminal resource allocation vector, determining a comprehensive cost function for each link of the second number of candidate paths; Based on the comprehensive cost functions of each link, determining the first number of candidate paths.
2. The method according to claim 1, wherein The constructing the network state prediction model based on gradient boosting trees based on the multi-source data includes: Using the principal component analysis method to screen and process the multi-source data to determine the main influencing features, and the main influencing features are the key features that affect the network state; Based on the main influencing features, constructing a network state prediction model based on gradient boosting trees.
3. The method according to claim 1, wherein The determining the first number of candidate paths based on the comprehensive cost functions of each link includes: Based on the comprehensive cost functions of each link, determining the cumulative cost of each candidate path among the second number of candidate paths; Screening in ascending order of the cumulative cost to determine the first number of candidate paths.
4. The method according to claim 1, wherein It further includes determining an alternative path for each of the first number of candidate paths.
5. The method according to claim 1, wherein The predefined state space includes the following: The real-time network state evaluation value, candidate path features, terminal energy consumption, and the service requirements; The terminal energy consumption includes battery content and energy consumption rate; The business requirements include the data volume and real-time requirements.
6. The method according to claim 1, wherein Calculating the reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space includes: Constructing a reward function for each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space; Solving the reward function to calculate the reward value corresponding to each candidate path.
Citation Information
Patent Citations
Adaptive QoS intelligent routing method based on deep reinforcement learning
CN116614437A
Satellite network adaptive routing method and system for time delay optimization
CN119743423A