Multi-mode satellite terminal adaptive routing intelligent decision-making method

By building a network state prediction model based on gradient enhancement tree and intelligent terminal routing decision model, optimizing the routing decision of multi-mode satellite terminals, the problem of difficulty in ensuring communication quality and terminal energy saving in the existing technology is solved, and the terminal energy consumption reduction and fault tolerance increase are achieved.

CN120034249AActive Publication Date: 2025-05-23HUNAN UNIV

Patent Information

Application Number
CN202510487918.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

When facing complex network environments and limited terminal energy, the existing routing mechanism is difficult to achieve communication quality assurance and terminal energy saving, and there are serious shortcomings in resource allocation and fault tolerance processing.

Method used

By collecting multi-source data, a network state prediction model based on gradient enhancement tree is constructed, the optimal dynamic resource allocation scheme and link failure prediction probability are determined, and the reward value of candidate paths is calculated based on business needs and terminal energy consumption, and the terminal routing intelligent decision-making model is constructed to optimize routing decisions.

Benefits of technology

It realizes that the terminal energy consumption is reduced, resource utilization is optimized, fault tolerance is enhanced, and the dynamic resource allocation efficiency of the network and the continuity of communication are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034249A_ABST
    Figure CN120034249A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode satellite terminal adaptive routing intelligent decision-making method. The method comprises the following steps: collecting multi-source data; based on the multi-source data, constructing a network state prediction model based on a gradient boosting tree; solving the network state prediction model to obtain a network state real-time evaluation value; determining an optimal dynamic resource allocation scheme and a link fault prediction probability based on the network state real-time evaluation value; determining a first number of candidate paths based on the network state real-time evaluation value, the service demand, the optimal dynamic resource allocation scheme and the link fault prediction probability; calculating a reward value corresponding to each candidate path based on the first number of candidate paths, a predefined state space and a predefined action space; based on the reward value, the predefined state space and the predefined action space, constructing a terminal routing intelligent decision model; and solving the terminal routing intelligent decision model, and determining an optimal transmission path. The resource utilization is optimized, and the energy consumption of the terminal is reduced on the premise of ensuring the communication quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communications, and in particular to a method for adaptively managing power supply of a multi-mode satellite terminal. Background Art

[0002] In today's communication networks, satellite-ground dual-mode communication terminals are widely used in agricultural production, emergency rescue and other fields, but the complex network environment and limited terminal energy restrict their performance. Traditional routing mechanisms have many shortcomings in dealing with dynamic changes in the network, such as complex controller configuration and high processing delay in software-defined networks (SDN) in large-scale networks; high training cost and slow convergence of energy-saving routing strategies based on deep reinforcement learning; poor traffic prediction accuracy and real-time performance of traffic-aware routing algorithms; energy-aware routing algorithms are prone to fall into local optimality and energy prediction is difficult; cross-layer energy-saving routing strategies have large inter-layer coordination overhead and poor protocol adaptability. These problems make it difficult for existing routing mechanisms to achieve terminal energy saving while ensuring communication quality, which seriously restricts the performance improvement of communication terminals.

[0003] In addition, traditional routing mechanisms have serious deficiencies in resource allocation and fault tolerance, making it difficult to achieve efficient dynamic resource allocation, and unable to respond quickly and ensure communication continuity when facing network failures. For example, when network traffic increases suddenly, traditional mechanisms cannot adjust resource allocation in a timely manner, resulting in congestion on some links and idle resources on some links; when a link fails, it cannot quickly switch to an alternate path, causing data transmission interruption. Summary of the invention

[0004] In order to overcome the above technical defects, the present application provides a multi-mode satellite terminal adaptive routing intelligent decision method. To achieve the above purpose, the present application is implemented according to the following technical solutions: The present application provides a multi-mode satellite terminal adaptive routing intelligent decision method, including: Collect data from multiple sources; Based on the multi-source data, a network status prediction model based on a gradient boosting tree is constructed; Solving the network status prediction model to obtain a real-time evaluation value of the network status; Determining an optimal dynamic resource allocation scheme and a link failure prediction probability based on the real-time evaluation value of the network status; Obtain business requirements; Determine a first number of candidate paths based on the real-time evaluation value of the network status, the service demand, the optimal dynamic resource allocation scheme and the link failure prediction probability; Obtain a predefined state space and a predefined action space; Calculate a reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space; Based on the reward value, the predefined state space and the predefined action space, construct a terminal routing intelligent decision model; Solve the terminal routing intelligent decision model to determine the optimal transmission path.

[0005] Optionally, constructing a network status prediction model based on a gradient boosting tree based on the multi-source data includes: The principal component analysis method is used to screen and process the multi-source data to determine the main influencing features; Based on the main influencing features, a network status prediction model based on a gradient boosting tree is constructed.

[0006] Optionally, determining an optimal dynamic resource allocation scheme based on the real-time evaluation value of the network status includes: Based on the real-time evaluation value of the network status, construct an objective function for maximizing the overall network benefit; The objective function is solved to determine the optimal dynamic resource allocation solution.

[0007] Optionally, determining the link failure prediction probability based on the real-time evaluation value of the network status includes: Obtain network real-time status parameters and network historical status parameters; Based on the real-time evaluation value of the network status, the real-time network status parameter and the historical network status parameter, a link failure prediction probability is determined.

[0008] Optionally, the determining a first number of candidate paths based on the real-time evaluation value of the network status, the service demand, the optimal dynamic resource allocation scheme, and the link failure prediction probability includes: Based on the business requirement, determining a second number of candidate paths; Determining a terminal resource allocation vector based on the optimal dynamic resource allocation solution; Determine the comprehensive cost function of each link of the second number of candidate paths based on the real-time evaluation value of the network status, the link failure prediction probability and the terminal resource allocation vector; Based on the comprehensive cost functions of the links, a first number of candidate paths are determined.

[0009] Optionally, the determining a first number of candidate paths based on the comprehensive cost functions of the links includes: Determining the cumulative cost of each candidate path in the second number of candidate paths based on the comprehensive cost function of each link; The candidate paths of the first number are determined by screening in order of cumulative cost from small to large.

[0010] Optionally, the method further includes determining an alternative path for each candidate path among the first number of candidate paths.

[0011] Optionally, the predefined state space includes the following contents: The real-time evaluation value of the network status, candidate path characteristics, terminal energy consumption, and service requirements; The terminal energy consumption includes battery content and energy consumption rate; The business requirements include data volume and real-time requirements.

[0012] Optionally, the calculating a reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space includes: constructing a reward function for each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space; The reward function is solved to calculate the reward value corresponding to each candidate path.

[0013] This application has the following beneficial effects: The method proposed in this application perceives the network status through multi-source data fusion, builds an accurate evaluation model, uses intelligent algorithms to optimize routing decisions, optimizes resource utilization, and enhances fault tolerance, thereby achieving terminal energy consumption reduction while ensuring communication quality.

[0014] In addition to the above-described purposes, features and advantages, the present application has other purposes, features and advantages. The present application will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 It is a flow chart of a multi-mode satellite terminal adaptive routing intelligent decision-making method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] The embodiments of the present application are described in detail below with reference to the accompanying drawings; however, the present application can be implemented in many different ways as defined and covered by the claims.

[0017] In order to solve the problems raised by the background technology, such as Figure 1 As shown, the present application provides a multi-mode satellite terminal adaptive routing intelligent decision method, comprising: Step S101: Collect multi-source data; Based on the multi-source data, a network status prediction model based on a gradient boosting tree is constructed; Multi-source data collection is achieved by integrating satellite communication modules, ground network monitoring modules and environmental sensors in the terminal.

[0018] The terminal integrated satellite communication module collects real-time monitoring satellite communication signal parameters such as signal strength, signal-to-noise ratio, bit error rate, etc., and then calculates the signal strength based on the above parameters. The calculation formula is as follows: (1) In the formula, is the received signal strength, is the transmitted signal power, and are the transmitting and receiving antenna gains, respectively, is the feeder loss, is the path loss.

[0019] The ground network monitoring module uses the Link Layer Discovery Protocol (LLDP) or IPv6 Neighbor Discovery Protocol (NDP) to obtain ground network signal strength, bandwidth utilization, delay and other parameters. The bandwidth utilization calculation formula is: (2) Environmental sensors collect temperature ,humidity , air pressure Environmental data such as temperature, temperature and signal propagation loss are related to each other. For example, according to the empirical formula, the temperature-induced propagation loss formula is: (3) in, is the temperature-dependent propagation loss, is the communication link distance, is the signal operating frequency, is the temperature correction function, which is implemented by table lookup or polynomial fitting.

[0020] After collecting the above multi-source data, the principal component analysis (PCA) method is used to screen and process the multi-source data to determine the main influencing features. The main influencing features can be understood as the key features that affect the network status. The judgment method is to regard the features that exceed the first threshold as the main influencing features. In order to accurately estimate the network status, the first threshold is generally set to a relatively large value. The above process of determining the main influencing features is as follows: For satellite signal strength, signal-to-noise ratio, bit error rate, link quality, bandwidth utilization, delay, wind speed, wind direction, altitude and other multi-source data, the principal component analysis method is used to map high-dimensional data to low-dimensional space by linear transformation, while retaining the variability of the data as much as possible. The multi-source data X is standardized, and the first k principal components with cumulative variance contribution rate ≥ 85% are retained. The standardized data is projected onto the principal components to obtain the reduced-dimensional data Y: (4) In the formula, is the eigenvalue of the ith principal component, p is the total number of original features, Is included before k The matrix of eigenvectors, k is the number of principal components selected, X′ is the standardized original data, and W is the number of principal components including the k The matrix of eigenvectors.

[0021] After determining the main influencing features, a network status prediction model based on the gradient boosting tree is constructed according to the main influencing features. The gradient boosting tree (GBDT) is an iterative decision tree algorithm that forms a strong learner by constructing multiple weak learners (decision trees). Its core idea is to gradually optimize the model, and each step of learning aims to correct the errors of the previous step, thereby improving the prediction accuracy of the model. The specific process of its construction is as follows: Since the network status prediction model of the gradient boosting tree requires an initial value, after determining the main influencing features, the initial prediction value is defined and calculated according to the main influencing features. The calculation definition process is as follows: (5) Among them, the signal strength S weight is , the maximum and minimum values ​​are ; The delayed D weight is , the maximum and minimum values ​​are ; The weight of packet loss rate L is , the maximum and minimum values ​​are ; The bit error rate E weight is , the maximum and minimum values ​​are .

[0022] Weight The sum is 1. Through this calculation method, the values ​​of each indicator are mapped to the interval [0, 1] and integrated according to their weights. The final result is The value is also in the range of [0, 1]. The closer the value is to 1, the better the network status is, and the closer the value is to 0, the worse the network status is.

[0023] After the initial forecast value is determined, it is used as the meanc, choose The target variable y , get the GBDT initial model ,Right now: (6) The main influencing features are used as the feature vectors of the model, and the process of iteratively constructing the network status prediction model based on the gradient boosting tree is as follows: ① Residual calculation , It is The sample in The residual of the wheel.

[0024] ② Fitting a new decision tree Prediction residuals .

[0025] ③Update the model, , is the learning rate, which controls the pace of model updates.

[0026] After M rounds of iterations, the network status prediction model based on the gradient boosting tree is obtained: (7) Step S102: solving the network status prediction model to obtain a real-time evaluation value of the network status; Apply the trained GBDT model to new network status data and input the real-time feature vector , the income That is, the real-time evaluation value of the network status , and finally obtained The value is in the range of [0, 1]. The closer the value is to 1, the better the network status is, and the closer the value is to 0, the worse the network status is.

[0027] Step S103: determining an optimal dynamic resource allocation scheme and a link failure prediction probability based on the real-time evaluation value of the network status; After determining and calculating the real-time evaluation value of the network status, the objective function of maximizing the overall network benefit is constructed according to the real-time evaluation value of the network status, and the linear programming algorithm is used to realize dynamic resource allocation. The specific process is as follows: Assume that there is terminals, and the resource requirement of each terminal is , the total network resources are ,set up Indicates that the The amount of resources per terminal, including bandwidth, power, time slot, etc. In order to maximize the overall benefits of the network, Indicates The number of terminals allocated resources And the network status evaluation result is The benefit function when , here is a simple assumption: (8) That is, the better the network status, the higher the benefits brought by the same amount of resources, and the benefits decrease with the boundary. Then the objective function of maximizing the overall benefits of the network is: (9) Constraints include total resource constraints, lower limit constraints on terminal resource requirements, etc. The total amount of resources allocated to all terminals cannot exceed the total network resources. , and the amount of resources allocated to the terminal cannot be lower than this minimum resource requirement ,Right now: (10) The above objective function is solved to determine the optimal dynamic resource allocation solution, thereby determining the resource vector allocated to each terminal.

[0028] In addition, in order to ensure communication reliability and enhance fault tolerance, fault prediction is carried out in advance, link quality mutations are captured in real time, and backup paths can be switched in time after link failure to avoid service termination and ensure communication reliability. The Bayesian inference model is used to combine network historical state parameters with network real-time state parameters for early fault prediction, and the link fault prediction probability model is dynamically adjusted according to the real-time evaluation value N of the network state.

[0029] Assume that the normal state of the network link is , the fault status is , based on the detected changed variables E (such as delay, packet loss rate, bit error rate, etc.), the posterior probability, i.e. the fault prediction probability, is calculated as follows: (11) in, and is the likelihood probability, obtained using the real-time network status parameters, and is the prior probability, obtained using the historical parameters of the network.

[0030] Dynamically adjust the prior probability according to the network status evaluation value N : (12) in, In order to control the impact of network status on the probability of failure, the grid search method is used to minimize the area under the ROC curve. , usually 0.8; The basic failure rate is calibrated by historical statistical values: (13) Step S104: Obtaining business requirements; Determine a first number of candidate paths based on the real-time evaluation value of the network status, the service demand, the optimal dynamic resource allocation scheme and the link failure prediction probability; At this time, the specific business requirements are obtained, and then the source node and the target node are confirmed, and then the second number of candidate paths between the source node and the target node are correspondingly determined.

[0031] After determining the second number of candidate paths, the Dijkstra algorithm is improved to comprehensively consider multi-dimensional parameters. , link failure probability , define the cost function of each link of the second number of candidate paths: (14) in, is the source node, is the target node, It is the network status of the target node, which is used to reflect the overall reliability. For Link Bandwidth, determined by the terminal resource allocation vector supply, is the link delay, is the historical failure probability of the link. Weight coefficient Adjusted by business dynamics, is the reliability weight, is the bandwidth weight, is the delay weight, is the failure probability weight.

[0032] After calculating the cost function of each link, the cumulative cost of each path in the second number of candidate paths is then calculated: (15) After calculating the cumulative cost of each path, the first number of candidate paths are screened in order of the cumulative cost from small to large, that is, the value of the second number is greater than the value of the first number.

[0033] To ensure that the entire network can operate normally after a link failure is detected, an alternative path is determined for each path in the first number of candidate paths. The specific process is as follows: The improved Dijkstra algorithm is also used to integrate multi-dimensional parameter planning, link related paths, and select the best backup path for each candidate path in the entire network. During the planning process, the backup path selection takes into account the network status N, the bit error rate E, and the primary path duplication penalty. ,Link failure probability , define the comprehensive cost function: (16) Among them, the penalty , is the penalty coefficient.

[0034] To ensure the reliability of the backup path, increase the failure probability weight , set to 0.8 to avoid the risk of secondary failure. Calculate the cumulative cost and select the path with the smallest cumulative cost as the alternative path.

[0035] Step S105: obtaining a predefined state space and a predefined action space; Calculate a reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space; At this time, the predefined state space and predefined action space are obtained. It covers network status, candidate path characteristics, terminal energy consumption, business requirements and other information. Network status is the real-time evaluation value N of network status; candidate path characteristics is the minimum candidate path cumulative cost value ; Terminal energy consumption includes battery power , energy consumption rate ; Business requirements include data volume , real-time requirements So the status can be represented as a vector: (17) Predefined action space An optional set of routing paths. Each action Represents the routing from the current node to the next node.

[0036] According to the first number of candidate paths, the predefined state space and the predefined action space, a reward function for each candidate path is constructed, wherein the reward function Taking into account transmission delay, energy consumption and reliability, the specific construction process is as follows: (18) Where delay is the actual delay, max_delay is the maximum delay. energy_consumed is the actual energy consumption, max_energy is the maximum energy consumption. success_count is the number of successfully transmitted packets, and total_count is the total number of transmitted packets. w1, w2, w3 are weight coefficients, satisfying w1+ w2+w3=1.

[0037] In order to obtain the optimal solution for the above weight coefficients, the following process is adopted: Through the weight distribution method, multiple objectives such as transmission delay D, energy consumption E, reliability R, etc. are comprehensively optimized. The multi-objective optimization problem is expressed as: (19) Where x is the decision variable, is the ith objective function.

[0038] By assigning a weight to each goal , after converting the multi-objective optimization problem into a single-objective optimization problem, solve it: (20) By solving the objective function of formula (20), we can obtain the specific values ​​of w1, w2, and w3, and then solve the reward function of formula (18) to calculate the reward value corresponding to each candidate path.

[0039] Step S106: constructing a terminal routing intelligent decision model based on the reward value, the predefined state space and the predefined action space; Solve the terminal routing intelligent decision model to determine the optimal transmission path.

[0040] After the reward value is determined, a terminal routing intelligent decision model is established based on the reward value, the predefined state space and the predefined action space, as follows: (twenty one) Where s is the current state and a is the action taken. r is the reward obtained after taking action a. s′ is the new state. a′ is any action that can be taken from the new state s′. α is the learning rate (0<α<1) that determines the influence of new information on the Q value update. The closer α is to 1, the greater the influence of new information. is a discount factor (0≤γ<1), reflecting the importance the agent places on future rewards, The closer it is to 1, the more it values ​​future rewards.

[0041] At each decision moment, the terminal selects the routing path with the largest Q value in the action space according to the current state as the optimal path for the current transmission. As the network state changes and data is continuously transmitted, the Q value is continuously updated to optimize the decision model.

[0042] In summary, the method proposed in this application perceives the network status through multi-source data fusion, builds an accurate evaluation model, uses intelligent algorithms to optimize routing decisions, optimizes resource utilization, and enhances fault tolerance, thereby achieving terminal energy consumption reduction while ensuring communication quality.

[0043] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-mode satellite terminal adaptive routing intelligent decision method, characterized in that: include: Collect data from multiple sources; Based on the multi-source data, a network status prediction model based on a gradient boosting tree is constructed; Solving the network status prediction model to obtain a real-time evaluation value of the network status; Determining an optimal dynamic resource allocation scheme and a link failure prediction probability based on the real-time evaluation value of the network status; Obtain business requirements; Determine a first number of candidate paths based on the real-time evaluation value of the network status, the service demand, the optimal dynamic resource allocation scheme and the link failure prediction probability; Obtain a predefined state space and a predefined action space; Calculate a reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space; Based on the reward value, the predefined state space and the predefined action space, construct a terminal routing intelligent decision model; Solve the terminal routing intelligent decision model to determine the optimal transmission path.

2. The method according to claim 1, characterized in that The step of constructing a network status prediction model based on a gradient boosting tree based on the multi-source data includes: The principal component analysis method is used to screen and process the multi-source data to determine the main influencing features; Based on the main influencing features, a network status prediction model based on a gradient boosting tree is constructed.

3. The method according to claim 1, characterized in that The determining of the optimal dynamic resource allocation scheme based on the real-time evaluation value of the network status includes: Based on the real-time evaluation value of the network status, construct an objective function for maximizing the overall network benefit; The objective function is solved to determine the optimal dynamic resource allocation solution.

4. The method according to claim 1, characterized in that: The determining the link failure prediction probability based on the real-time evaluation value of the network status includes: Obtain network real-time status parameters and network historical status parameters; Based on the real-time evaluation value of the network status, the real-time network status parameter and the historical network status parameter, a link failure prediction probability is determined.

5. The method according to claim 1, characterized in that The determining a first number of candidate paths based on the real-time evaluation value of the network status, the service demand, the optimal dynamic resource allocation scheme, and the link failure prediction probability includes: Based on the business requirement, determining a second number of candidate paths; Determining a terminal resource allocation vector based on the optimal dynamic resource allocation solution; Determine the comprehensive cost function of each link of the second number of candidate paths based on the real-time evaluation value of the network status, the link failure prediction probability and the terminal resource allocation vector; Based on the comprehensive cost functions of the links, a first number of candidate paths are determined.

6. The method according to claim 5, characterized in that The determining a first number of candidate paths based on the comprehensive cost functions of the links includes: Determining the cumulative cost of each candidate path in the second number of candidate paths based on the comprehensive cost function of each link; The candidate paths of the first number are determined by screening in order of cumulative cost from small to large.

7. The method according to claim 1, characterized in that The method further includes determining an alternative path for each candidate path of the first number of candidate paths.

8. The method according to claim 1, characterized in that: The predefined state space includes the following: The real-time evaluation value of the network status, the characteristics of the candidate paths, the energy consumption of the terminal, and the service requirements; The terminal energy consumption includes battery content and energy consumption rate; The business requirements include data volume and real-time requirements.

9. The method according to claim 1, characterized in that: The calculating the reward value corresponding to each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space includes: constructing a reward function for each candidate path based on the first number of candidate paths, the predefined state space, and the predefined action space; The reward function is solved to calculate the reward value corresponding to each candidate path.

Citation Information

Patent Citations

  • Self-adaptive routing method, system and equipment oriented to high-dynamic network topology

    CN114124823A

  • Adaptive QoS intelligent routing method based on deep reinforcement learning

    CN116614437A

  • Load balancing routing method and system of satellite network

    CN119255230A

  • Satellite network adaptive routing method and system for time delay optimization

    CN119743423A

Cited By

  • Flow transmission method based on SRv6

    CN121367669A

  • A SRv6-based traffic transmission method

    CN121367669B