Intelligent multipath transmission method and system based on time delay synchronization constraint

By building an intelligent decision-making model for multi-path transmission and multi-objective optimization reward function, the problem of delay inconsistency in the intelligent multi-path routing algorithm is solved, path delay consistency and load balancing are achieved in a highly dynamic network environment, and real-time and reliability of information transmission are improved.

CN120281703APending Publication Date: 2025-07-08WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510482670.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing intelligent multipath routing algorithms have failed to effectively solve the problem of delay inconsistency caused by link separation, and the routing transmission decision-making mechanism in a highly dynamic network environment is not high, making it difficult to meet the real-time and intelligent needs of information services.

Method used

Build a multi-path transmission intelligent decision-making model, design multi-dimensional state space and multi-objective optimization reward function, adopt PPO-LSTM architecture and experience playback pool, and guide the agent to select multiple transmission paths with strong delay consistency through multi-objective optimization reward function, reduce path delay and achieve load balancing.

Benefits of technology

By designing multi-objective optimization reward function and multi-dimensional state space, the model training efficiency is improved, path delay consistency and load balancing are achieved in a high dynamic network environment, and the real-time and reliability of information transmission are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281703A_ABST
    Figure CN120281703A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent multipath transmission method and system based on time delay synchronization constraint. The method comprises the following steps: constructing a multipath transmission intelligent decision model; designing a multi-dimensional state space oriented to multi-path transmission; establishing a multi-objective optimization reward function based on time delay synchronization constraint; adopting the designed multi-dimensional state space and the established multi-target optimization reward function to drive intelligent exploration of multiple groups of source-destination node pairs, and generating historical sequence sample data; and training to obtain the optimized multi-path transmission intelligent decision model. According to the method, the reward function of multi-objective optimization is designed, so that multiple paths are allowed to intersect at the node with low congestion degree, and the path delay is effectively reduced while load balancing is realized; in combination with a targeted state space design and a time delay constraint mechanism, an intelligent agent is guided to select multi-path transmission with high time delay consistency; by introducing an experience playback pool and an improved sample playback strategy, the model convergence performance and the model training efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent network communication technology, and in particular to an intelligent multipath transmission method and system based on delay synchronization constraints. Background Art

[0002] With the development of social economy and the advent of the big data era, users' demand for information services is developing in the direction of real-time and intelligentization. The comprehensive use of multiple information service networks to provide transmission services for massive data has become an important direction for the development of current information rapid transmission technology. However, complex and heterogeneous information service networks not only provide multiple transmission links to ensure the rapid transmission of big data, but also the strategy of transmitting data through multiple routing paths, which also brings about inconsistent data transmission delays and causes asynchronous data processing at the receiving end. Therefore, how to achieve rapid information transmission while ensuring the consistency of delays of different transmission paths in a complex multi-link information service network environment is a key issue that needs to be solved in real-time information services.

[0003] The existing routing transmission methods applied to multi-path networks mainly include traditional methods and intelligent methods based on reinforcement learning. Traditional routing algorithms cannot adapt to highly dynamic network environments, while intelligent routing algorithms based on reinforcement learning just make up for this shortcoming. Through the interaction between the agent and the environment, they can self-optimize based on real-time feedback and automatically update routing strategies. Among them, the pheromone-incentivized intelligent multi-path routing scheduling algorithm (reference: Huang Y, Jiang X, Chen S, et al. Pheromone incentivized intelligent multipath traffic scheduling approach for LEO satellite networks. IEEE Transactions on Wireless Communications, 2022, 21(8): 5889-5902.) uses pheromones in the ant colony algorithm (ACO) to record network states, proposes a pheromone-incentivized multi-path discovery algorithm, finds multiple paths and evaluates path performance, selects several routing paths with the best transmission performance from them, and uses the deep deterministic policy gradient algorithm (DDPG) to train the agent for multi-path traffic scheduling. Although the reinforcement learning method is used, it is applied to multi-path traffic scheduling and does not achieve agent path finding, still belonging to the category of traditional algorithms, with low adaptability to dynamic networks. The reinforcement learning-based service-oriented dynamic multi-path routing algorithm (RLMR) (reference: Chen C, Xue F, Lu Z, et al. RLMR: Reinforcement Learning Based Multipath Routing for SDN. Wireless Communications & Mobile Computing, 2022.) formulates a differentiated reward method for each service type through a reinforcement learning model and dynamically assigns routing paths to each service according to the network state. It uses reinforcement learning to achieve routing path allocation rather than the routing discovery process.The Link Disjoint Multipath Routing (LDMR) transmission algorithm (Reference: Huang Y, Yang D, Feng B, et al. AGNN-enabled multipath routing algorithm for spatial-temporal varying LEO satellite networks. IEEE Transactions on Vehicular Technology, 2023.) realizes the traffic allocation among multiple paths based on reinforcement learning. However, the routing discovery simply calls the Dijkstra algorithm multiple times in a loop to find k shortest paths, and its adaptability to the dynamic network environment is also low. Moreover, its link disjoint strategy will increase additional latency, resulting in low service timeliness.

[0004] The Improved Exploration Algorithm for Multipath Routing Based on Q-Learning (Reference: Hassen, Houda, et al. Improved Exploration Strategy for Q-Learning Based Multipath Routing in SDN Networks. Journal of Network and Systems Management, 2024.) proposes an improved Q-Learning algorithm based on congestion avoidance, using a new exploration strategy based on the Max-Boltzman exploration method to select actions. However, due to the limitations of the Q-Learning algorithm, it is difficult to apply to network environments with continuous state spaces. The Intelligent Multipath Routing Algorithm (EMCR) based on Deep Reinforcement Learning (DRL) (Reference: Liu X, Zhou H, Zhang Z, et al. Multipath Cooperative Routing in Ultra-Dense LEO Satellite Networks: A Deep Reinforcement Learning-Based Approach. IEEE Internet of Things Journal, 2024.) estimates the local network state of each network node through deep reinforcement learning and makes independent next-hop forwarding decisions accordingly. However, its multipath routing does not consider the latency inconsistency caused by the large latency gap between paths, which has a great impact on real-time transmission services.

[0005] In summary, the current intelligent multi-path routing algorithms are not designed for the delay inconsistency caused by link separation, and the mechanism of making routing transmission decisions based on the current network state has low performance in a highly dynamic network environment. Therefore, to meet the real-time and intelligent requirements of users for information services, it is necessary to fully consider the impact of network state changes on the performance of routing decisions and the restriction of the delay differences of different transmission paths on the timeliness of services, in order to meet the information service requirements of instant transmission. Summary of the Invention

[0006] This application provides an intelligent multi-path transmission method and system based on delay synchronization constraints, which can solve the technical problems that the existing intelligent multi-path routing algorithms are not designed for the delay inconsistency caused by link separation, and the mechanism of making routing transmission decisions based on the current network state has low performance in a highly dynamic network environment.

[0007] In the first aspect, this application provides an intelligent multi-path transmission method based on delay synchronization constraints, including the following steps:

[0008] Step S1: Construct an intelligent decision-making model for multi-path transmission;

[0009] Step S2: Design a multi-dimensional state space for multi-path transmission;

[0010] Step S3: Establish a multi-objective optimization reward function based on delay synchronization constraints;

[0011] Step S4: Use the designed multi-dimensional state space and the established multi-objective optimization reward function to drive the intelligent exploration of multiple source-destination node pairs, and generate historical sequence sample data including the states, actions, rewards, and next moment states of the source-destination node pairs during the path finding process;

[0012] Step S5: Select some samples from the historical sequence sample data and put them into the experience replay pool of the intelligent decision-making model for multi-path transmission, deploy the trained intelligent decision-making model for multi-path transmission, and generate routing decisions for the multi-path network based on delay synchronization constraints.

[0013] Further, the Step S1: Construct an intelligent decision-making model for multi-path transmission specifically includes the following steps:

[0014] Use the PPO model as the basic model of the multi-path selection agent, add 2 LSTM models respectively before the Actor and Critic networks of the PPO model, and add an experience replay pool suitable for the PPO algorithm to construct an intelligent decision-making model for multi-path transmission based on the PPO-LSTM architecture.

[0015] Further, the Step S2: Design a multi-dimensional state space for multi-path transmission specifically includes the following steps:

[0016] S21: Calculate the minimum hop count from the directly adjacent neighbor nodes of the current node to the destination node using the Dijkstra algorithm as the hop count cost;

[0017] S22: Sum up the transmission delay, propagation delay, and queuing delay to calculate the delay cost of the transmission delay cost between the current node and the neighbor node;

[0018] S23: Evaluate the network congestion degree based on the average queue occupancy rate of the second-order neighbor nodes of the current node;

[0019] S24: Obtain other states including the path loop state, node intersection state, and link intersection state;

[0020] S25: Integrate the hop count cost, delay cost, congestion degree, path loop state, node intersection state, and link intersection state to construct a multi-dimensional state space for multi-path transmission.

[0021] Furthermore, the delay cost D ij has the following specific calculation formula:

[0022]

[0023] In the formula, p ij represents the number of data packets transmitted between node i and adjacent node j, m represents the data packet size, Bw ij represents the link bandwidth between node i and adjacent node j, d ij represents the link distance, c represents the propagation rate of the signal, q ij represents the number of data packets queuing in the queue, and Op represents the speed of the router processing the queuing data packets.

[0024] Furthermore, the step S3: Establish a multi-objective optimization reward function based on delay synchronization constraints specifically includes the following steps:

[0025] S31: Obtain the hop count reward;

[0026] S32: Obtain the delay reward;

[0027] S33: Obtain the congestion degree reward;

[0028] S34: Obtain other case rewards including path loop reward, node intersection reward, and link intersection reward;

[0029] S35: Perform weighted summation on the obtained hop count reward, delay reward, congestion degree reward, path loop reward, node intersection reward, and link intersection reward to obtain the reward function based on multi-objective optimization.

[0030] Further, in step S4: Using the designed multi-dimensional state space and the established multi-objective optimization reward function, drive the intelligent exploration of multiple source-destination node pairs, and generate historical sequence sample data including the states, actions, rewards, and next moment states of the source-destination node pairs during the path finding process, which specifically includes the following steps:

[0031] S41: Randomly select a group of source-destination node pairs, choose one of the source-destination node pairs, start from the source node, obtain the neighbor node states of the current node, input to the intelligent agent to obtain the routing action decision for multi-path transmission, execute the routing action decision, obtain the immediate reward for routing to the next node and the next state, and return them to the intelligent agent until the destination node is found or the maximum step limit is reached. Repeat this process to complete the path finding operation for all source-destination node pairs in this group;

[0032] S42: Repeat step S41 to complete the path finding operations for all groups of source-destination node pairs in the multi-path, and generate historical sequence sample data including the states, actions, rewards, and next moment states of all groups of source-destination node pairs during the path finding process.

[0033] Further, in step S5: Screen some samples from the historical sequence sample data and put them into the experience replay pool of the multi-path transmission intelligent decision-making model. Deploy the trained multi-path transmission intelligent decision-making model to generate routing decisions for the multi-path network based on the time delay synchronization constraint, which specifically includes the following steps:

[0034] S51: Screen samples that meet the screening conditions from the historical sequence sample data and put them into the experience replay pool;

[0035] S52: Perform mini-batch sampling from the experience replay pool, and calculate the ratio of the new and old policy probabilities and the discounted reward using the mini-batch collected data;

[0036] S53: The Critic network uses the discounted reward as the actual state value and calculates the value estimation loss;

[0037] S54: The Actor network uses the ratio of the new and old policy probabilities calculated from the sample data to limit the policy update amplitude output by the basic model;

[0038] S55: According to the obtained ratio of the new and old policy probabilities, the discounted reward, and the value estimation loss, calculate the loss function of the constructed multi-path transmission intelligent decision-making model. Repeat steps S51 - S54. When the convergence condition is met, terminate the training, deploy the trained multi-path transmission intelligent decision-making model, and generate routing decisions for the multi-path network based on the time delay synchronization constraint.

[0039] Further, the sample screening strategy specifically includes:

[0040] Calculate the ratio of the action state probability distribution of each sample in the historical sample sequence to the current policy;

[0041] Set the policy clipping factor ε, and filter the samples whose ratios are within the clipping range.

[0042] In a second aspect, the present application provides an intelligent multipath transmission system based on delay synchronization constraints, including:

[0043] A basic model construction module, configured to construct a multipath transmission intelligent decision-making model;

[0044] A state space design module, configured to design a multi-dimensional state space for multipath transmission;

[0045] A reward function establishment module, configured to establish a multi-objective optimization reward function based on delay synchronization constraints;

[0046] A historical sequence sample data acquisition module, communicatively connected to the state space design module and the reward function establishment module, and configured to drive the intelligent exploration of multiple source-destination node pairs by using the designed multi-dimensional state space and the established multi-objective optimization reward function, and generate historical sequence sample data including the state, action, reward, and next moment state of the source-destination node pairs during the route finding process;

[0047] A model training and routing decision output module, communicatively connected to the model establishment module and the historical sequence sample data acquisition module, and configured to screen out some samples from the historical sequence sample data and put them into the experience replay pool of the multipath transmission intelligent decision-making model, deploy the trained multipath transmission intelligent decision-making model, and generate a routing decision for the multipath network based on delay synchronization constraints.

[0048] Further, the state space design module includes:

[0049] A hop count cost acquisition unit, configured to calculate the minimum hop count from the directly adjacent neighbor nodes of the current node to the destination node as the hop count cost by using the Dijkstra algorithm;

[0050] A node delay cost acquisition unit, configured to sum up the transmission delay, propagation delay, and queuing delay to calculate the delay cost of the transmission delay cost between the current node and the neighbor node;

[0051] A network congestion degree acquisition unit, configured to evaluate the network congestion degree based on the average queue occupancy rate of the second-order neighbor nodes of the current node;

[0052] An other state acquisition unit, configured to acquire other states including path loop state, node intersection state, and link intersection state;

[0053] The multi-dimensional state space construction unit is communicatively connected to the hop count cost acquisition unit, the node delay cost acquisition unit, the network congestion degree acquisition unit, and other state acquisition units, and is used to integrate the hop count cost, the delay cost, the congestion degree, the path loop state, the node intersection state, and the link intersection state to construct a multi-dimensional state space for multi-path transmission.

[0054] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0055] By designing a multi-objective optimization reward function, allowing multi-paths to intersect at nodes with a lower congestion degree, while achieving load balancing, effectively reducing the path delay;

[0056] By designing the state space and the reward function, based on the traditional algorithm, specifically designing the delay constraints of different routing paths to guide the intelligent agent to select multiple transmission paths with strong delay consistency;

[0057] By introducing an experience replay pool and designing a sample replay strategy, the convergence performance of the PPO algorithm is improved, and the model training efficiency is increased. Description of the Drawings

[0058] Figure 1 It is a schematic flowchart of an intelligent multi-path transmission method based on delay synchronization constraints provided by an embodiment of the present application;

[0059] Figure 2 It is a schematic diagram of an intelligent multi-path routing algorithm provided by an embodiment of the present application;

[0060] Figure 3 It is a schematic diagram of the model training process provided by an embodiment of the present application;

[0061] Figure 4 It is a functional module block diagram of an intelligent multi-path transmission system based on delay synchronization constraints provided by an embodiment of the present application. Detailed Embodiments

[0062] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0063] In the description of the specification and claims of this application and the above-mentioned drawings, the terms "comprising", "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. Descriptions such as "first", "second", and "third" are used to distinguish different objects, etc., and do not represent a sequential order, nor do they limit that "first", "second", and "third" are different types.

[0064] In the description of the embodiments of this application, terms such as "exemplary", "for example", or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.

[0065] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0066] In some processes described in the embodiments of this application, a plurality of operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0067] First, some technical terms in this application are explained to facilitate understanding of this application by those skilled in the art.

[0068] Experience Buffer: Experience buffer pool or experience replay pool.

[0069] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the drawings.

[0070] In the first aspect, as Figure 1And Figure 2 As shown, the present application provides an intelligent multipath transmission method based on delay synchronization constraints, including the following steps:

[0071] Step S1: Construct an intelligent decision-making model for multipath transmission;

[0072] Step S2: Design a multi-dimensional state space for multipath transmission;

[0073] Step S3: Establish a multi-objective optimization reward function based on delay synchronization constraints;

[0074] Step S4: Use the designed multi-dimensional state space and the established multi-objective optimization reward function to drive the intelligent exploration of multiple source-destination node pairs, generating historical sequence sample data including the states, actions, rewards, and next moment states of the source-destination node pairs during the path finding process;

[0075] Step S5: Screen out some samples from the historical sequence sample data and put them into the experience replay pool of the intelligent decision-making model for multipath transmission, deploy the trained intelligent decision-making model for multipath transmission, and generate a routing decision for the multipath network based on delay synchronization constraints.

[0076] The present application designs a multi-objective optimization reward function, allowing multipaths to intersect at nodes with a lower congestion level, achieving load balancing while effectively reducing path delay; through the designed state space and reward function, on the basis of traditional algorithms, it specifically designs delay constraints for different routing paths to guide the intelligent agent to select multiple transmission paths with strong delay consistency; by introducing an experience replay pool and designing a sample replay strategy, it improves the convergence performance of the PPO algorithm and improves the model training efficiency; the solution of the present application can intelligently select multiple paths, which are dynamically variable, aiming to break through the bandwidth limitation of a single link through parallel transmission of multiple paths, while improving the transmission rate, ensuring that when multiple links are transmitted, the waiting during in-order reception at the receiving end is avoided.

[0077] In an embodiment, the step S1: Construct an intelligent decision-making model for multipath transmission specifically includes the following steps:

[0078] Step S11: Use the PPO model as the basic model of the multipath selection intelligent agent;

[0079] Step S12: Add 2 LSTM models respectively before the Actor and Critic networks of the PPO model to obtain a reinforcement learning model, which is used to extract the change characteristics of delay cost and congestion level, and input them into the basic model together with other states;

[0080] Step S13: Add an experience replay pool applicable to the PPO algorithm to the basic model for putting in some sample data screened according to rules;

[0081] Based on steps S11, S12, and S13, a multi-path transmission intelligent decision-making model based on the PPO-LSTM architecture is constructed.

[0082] The multi-path selection intelligent agent implemented based on the constructed multi-path transmission intelligent decision-making model dynamically interacts with the network environment. Specifically, according to the network environment state S at the current time t t , calculate the rewards r that can be obtained by taking different actions (in this algorithm, it is routing path selection) for data transmission t , filter out the action a with the maximum reward t , and after applying this action to the environment (perform the next-hop transmission of data based on this path selection), observe the new environment state S at time t+1 t+1 , and loop in sequence until the intelligent agent completes the dynamic path selection decision according to the network environment state.

[0083] In one embodiment, step S2: Design a multi-dimensional state space for multi-path transmission, specifically including the following steps:

[0084] S21: Use the Dijkstra algorithm to calculate the minimum number of hops from a neighbor node to the destination node as the hop cost H j ;

[0085] S22: Sum the transmission delay TD ij , propagation delay PD ij and queuing delay QD ij to calculate the delay cost D between node i and neighbor node j ij , and the calculation formula is shown as follows:

[0086]

[0087] In the formula, p ij represents the number of data packets transmitted between node i and neighbor node j, m represents the data packet size, Bw ij represents the link bandwidth between node i and neighbor node j, d ij represents the link distance between node i and neighbor node j, c represents the signal propagation rate, q ij represents the number of data packets queued in the queue between node i and neighbor node j, and Op represents the speed of the router for processing queued data packets.

[0088] S23: Evaluate the network congestion degree based on the average queue occupancy rate of second-order neighbor nodes; considering that the data is for the next-hop selection after being transmitted from the current node i to the neighbor node j, the congestion degree to be calculated is the congestion degree when the neighbor node j continues to transmit the data to the next-hop node after receiving the data, that is, the congestion degree is represented by the average queue occupancy rate of the neighbor nodes of the next-hop neighbor node j of node i (i.e., the second-order neighbor of node i). For the neighbor node j of node i, the symbol W ij represents the average queue occupancy rate considering the buffer queues in all directions of node j except the direction of node i, and the calculation formula is:

[0089]

[0090] In the formula, R j represents the set of neighbor nodes of node j, q jk represents the number of data packets queued in the queue between node j and its neighbor node k, Q jk represents the queue length, Card(R j ) represents the number of elements in the set R j . When calculating the queue occupancy of the neighbor nodes of node j, since the data does not need to be sent back to node i, the queue occupancy of node i needs to be removed;

[0091] S24: Obtain other states including path loop state, node intersection state, and link intersection state; among them, the path loop state indicates whether the neighbor node j is in the current path, represented by NC j ; the node intersection state indicates whether the neighbor node j is in other paths, represented by NO j ; the link intersection state indicates whether the link between the current node i and the neighbor node j is in other paths, represented by LO ij ; this application proposal allows node intersections and can dynamically give different rewards according to the load of the nodes, enabling the agent to make flexible decisions, improving the flexibility of link routing and the adaptability to dynamic environments;

[0092] S25: Integrate the hop count cost, delay cost, congestion degree, path loop state, node intersection state, and link intersection state to construct a multi-dimensional state space for multi-path transmission as S t ={H j , D ij , W ij , NC j , NO j , LO ij |j∈R i}, R i represents the set of neighbor nodes of node i.

[0093] In one embodiment, step S3: Establish a multi-objective optimization reward function based on time-delay synchronization constraints, which specifically includes the following steps:

[0094] S31: Obtain the hop count reward r H , and the reward function of the hop count reward is:

[0095]

[0096] To ensure that the routing path direction is correct and the path is as short as possible, it is designed that when the next-hop direction is towards the destination node direction (the hop count to the destination node becomes smaller), a positive reward is obtained; when the next-hop direction is opposite to the destination node direction (the hop count to the destination node increases), a negative reward is obtained.

[0097] S32: Obtain the time-delay reward r D , and the reward function of the time-delay reward is:

[0098]

[0099] Among them, DC oj represents the cumulative time delay of the path from the source node o to the node j, which can be obtained by cumulative calculation; D oe represents the estimated minimum time delay from the source node o to the destination node e, which is used as the time-delay threshold and can be estimated according to the shortest path between the source node o and the destination node e; PD represents the estimated time delay of the current path, and the calculation formula is PD = DC oj +H j ×AD, where AD represents the average time delay per hop, which is estimated through historical training data; by comparing the estimated time delay PD of the current path and the time-delay threshold D oe to design the time-delay reward. If the estimated time delay of the current path does not exceed the time-delay threshold, a positive reward is obtained, and the closer to the destination node, the greater the positive reward; when the estimated time delay exceeds the time-delay threshold, a negative reward is obtained.

[0100] S33: Obtain the congestion degree, and the congestion degree reward function is shown in the following formula:

[0101]

[0102] In the formula, represents taking the average queue occupancy rate of the node with the highest congestion degree among all nodes on the current path, and the node with high load will obtain a greater negative congestion degree reward;

[0103] S34: Obtain other situation rewards including the path loop reward r NC , the node intersection reward r NO and the link intersection reward r LO , and the reward functions of the other situation rewards are respectively shown in the following formulas:

[0104] r NC = -NC j Equation (6)

[0105] r NO = -NO j *W ij Equation (7)

[0106] r LO = -LO ij Equation (8)

[0107] S35: Weighted sum the obtained hop count reward, latency reward, congestion degree reward, path loop reward, node intersection reward and link intersection reward to obtain a reward function based on multi-objective optimization, as shown in the following equation:

[0108] r = α1r H + α2r D + α3r W + α4r NC + α5r NO + α6r LO Equation (9)

[0109] In this application, α1 = 0.4, α2 = 0.8, α3 = 0.1, α4 = 1, α5 = 0.1, α6 = 1 are set.

[0110] In one embodiment, the latency cost D ij The specific calculation formula is:

[0111]

[0112] In the formula, p ij represents the number of data packets transmitted between node i and neighbor node j, m represents the data packet size, Bw ij represents the link bandwidth between node i and neighbor node j, d ij represents the link distance between node i and neighbor node j, c represents the signal propagation rate, q ij represents the number of data packets queued in the queue between node i and neighbor node j, and Op represents the speed of the router to process queued data packets.

[0113] In one embodiment, the step S4: Using the designed multi-dimensional state space and the established multi-objective optimization reward function, drive the intelligent exploration of multiple source-destination node pairs to generate historical sequence sample data including the state, action, reward and next moment state of the source-destination node pairs during the path finding process, specifically including the following steps:

[0114] S41: Randomly select a group of source-destination node pairs; specifically, arbitrarily select 3 neighbor nodes of a node as the starting node, and arbitrarily select a node different from the source node as the destination node to form a group of source-destination node pairs;

[0115] Select one of the source-destination node pairs and perform the following operations starting from the source node:

[0116] Obtain the neighbor node status of the current node, including the historical status and the current status, and input it to the intelligent agent for action decision-making.

[0117] Execute the action decision-making, route the data to the next node, and obtain the immediate reward and the next status of the current action decision-making and return them to the intelligent agent.

[0118] Repeat the above operations until the destination node in the current group of source-destination node pairs is found or the maximum step limit is reached. In this method, the maximum step is set to the maximum number of nodes in the network. Obtain the neighbor node status of the current node, input it to the intelligent agent to obtain the routing action decision for multi-path transmission, execute the routing action decision, obtain the immediate reward for routing to the next node and the next status and return them to the intelligent agent until the destination node is found, or the maximum step limit is reached (in this application, the maximum step is set to the maximum number of nodes in the network), and repeat in this way to complete the path-finding operations for all source-destination node pairs in this group;

[0119] S42: Repeat step S41 to complete the path-finding operations for all groups of source-destination node pairs in the multi-path, and generate historical sequence sample data including the status, actions, rewards, and next moment status of all groups of source-destination node pairs during the path-finding process.

[0120] In one embodiment, as Figure 3 shown, the step S5: Screen out some samples from the historical sequence sample data and put them into the experience replay pool of the multi-path transmission intelligent decision-making model, deploy the trained multi-path transmission intelligent decision-making model, and generate routing decisions for the multi-path network based on delay synchronization constraints, which specifically includes the following steps:

[0121] S51: Screen out the samples that meet the screening conditions from the historical sequence sample data and put them into the experience replay pool; further, the sample screening strategy in step S51 specifically includes:

[0122] Calculate the ratio of the action-state probability distribution of each sample in the historical sample sequence to the current policy;

[0123] Set the policy clipping factor ε, and screen out the samples with the ratio within the clipping range and put them into the experience replay pool;

[0124] Considering the proximal policy update idea of the PPO algorithm, the sample screening operation method of the experience replay pool is as follows:

[0125]

[0126] where Trajs = {Traj1, …, Traj n , …, Traj N} represents the historical sample sequence, N represents the number of sample sequences, Traj n = {(S0, a0, r0, S1), …, (S t , a t , r t , S t+1 ), …, (S T-1 , a T-1 , r T-1 , S T )} represents the nth sample sequence, S t , a t , r t respectively represent the state, action, and reward at the t-th moment in the sample sequence; P represents the action-state probability distribution, π θ represents the current policy, ε represents the policy clipping factor, and only when the ratio of the action-state probability distribution of the sample data to the current policy is within the clipping range, the sample is retained in the experience replay pool; in this application, ε takes a value of 0.2; T represents the time length of the sample sequence;

[0127] S52: Perform mini-batch sampling from the experience replay pool and calculate the ratio ratio t (θ) of the new and old policies and the discounted reward R t , as shown in the following formula:

[0128]

[0129] R t = r t+1 + αr t+2 + α 2 r t+3 + … + α T-t r T+1 + α T-t+1 V(S T+1 ) Equation (12)

[0130] where π θ represents the current policy, represents the policy before policy update. α is the discount factor, which takes a value of 0.95 in this algorithm. V is the action value output by the Critic network in the base model;

[0131] S53: The Critic network uses the discounted reward as the actual state value to calculate the value estimation loss, and the calculation formula of the value estimation loss is shown as follows;

[0132] V error = MSE(R t - V(S t )) Equation (13)

[0133] S54: The Actor network uses the ratio of the new and old policy probabilities calculated from the sample data to limit the policy update amplitude output by the base model;

[0134] S55: Repeat the execution of S51 - S54. The total loss function of the base model is the weighted sum of the policy loss, value loss, and entropy reward to obtain the total loss function. Update the parameters of the Actor and Critic networks through backpropagation and an optimizer (such as Adam) to gradually optimize the model, generate a multi-path transmission policy that meets the delay synchronization constraint, deploy the trained multi-path transmission intelligent decision-making model, and generate a routing decision for the multi-path network based on the delay synchronization constraint.

[0135] This application adopts a distributed architecture. Each agent can learn independently and only needs to obtain the state information of the local node and its neighbor nodes to complete the decision-making without global network topology awareness. This design significantly reduces the communication overhead and computational complexity, and improves the feasibility of the system and the practicality of actual deployment.

[0136] In the second aspect, as Figure 4As shown in the figure, the present application provides an intelligent multipath transmission system based on delay synchronization constraints, including: a basic model construction module 100, a state space design module 200, a reward function establishment module 300, a historical sequence sample data acquisition module 400, and a model training and routing decision output module 500; the basic model construction module 100 is used to construct a multipath transmission intelligent decision model; the state space design module 200 is used to design a multi-dimensional state space for multipath transmission; the reward function establishment module 300 is used to establish a multi-objective optimization reward function based on delay synchronization constraints; the historical sequence sample data acquisition module 400 is communicatively connected to the state space design module 200 and the reward function establishment module 300, and is used to drive the intelligent exploration of multiple source-destination node pairs by using the designed multi-dimensional state space and the established multi-objective optimization reward function, and generate historical sequence sample data including the state, action, reward, and next moment state of the source-destination node pair during the path finding process; the model training and routing decision output module 500 is communicatively connected to the model establishment module 100 and the historical sequence sample data acquisition module 400, and is used to screen out some samples from the historical sequence sample data and put them into the experience replay pool of the multipath transmission intelligent decision model, deploy the trained multipath transmission intelligent decision model, and generate a routing decision for the multipath network based on delay synchronization constraints.

[0137] Further, the state space design module includes:

[0138] A hop count cost acquisition unit, configured to calculate the minimum hop count from the directly adjacent neighbor nodes of the current node to the destination node as the hop count cost by using the Dijkstra algorithm;

[0139] A node delay cost acquisition unit, configured to sum the transmission delay, propagation delay, and queuing delay to calculate the delay cost of the transmission delay cost between the current node and the neighbor node;

[0140] A network congestion degree acquisition unit, configured to evaluate the network congestion degree based on the average queue occupancy rate of the second-order neighbor nodes of the current node;

[0141] An other state acquisition unit, configured to acquire other states including the path loop state, node intersection state, and link intersection state;

[0142] A multi-dimensional state space construction unit, communicatively connected to the hop count cost acquisition unit, the node delay cost acquisition unit, the network congestion degree acquisition unit, and the other state acquisition unit, and is used to integrate the hop count cost, delay cost, congestion degree, path loop state, node intersection state, and link intersection state to construct a multi-dimensional state space for multipath transmission.

[0143] Among them, the functional implementation of each module in the above intelligent multipath transmission system based on delay synchronization constraint corresponds to each step in the above embodiment of the intelligent multipath transmission method based on delay synchronization constraint, and its functions and implementation processes will not be elaborated here one by one.

[0144] In a third aspect, an embodiment of the present application provides an intelligent multipath transmission device based on delay synchronization constraint. An intelligent multipath transmission device based on delay synchronization constraint can be a device with data processing functions such as an airborne computer, a personal computer (PC), a laptop computer, a server, an intelligent satellite, etc.

[0145] The communication interface includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing the interconnection of components inside an intelligent multipath transmission device based on delay synchronization constraint, and interfaces for implementing the interconnection between an intelligent multipath transmission device and other devices (such as other computing devices or user devices). The physical interface can be a wireless network interface (including wireless ad hoc network, public mobile communication network, LTE private network, satellite network, etc.), a wired network interface (including Ethernet interface, optical fiber interface, ATM interface, etc.); the user device can be a display screen (Display), a keyboard (Keyboard), etc.

[0146] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0147] The processor can be a general-purpose processor, and the general-purpose processor can call an intelligent multipath transmission program stored in the memory and execute the intelligent multipath transmission method provided by the embodiment of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the intelligent multipath transmission program is called can refer to each embodiment of the intelligent multipath transmission method of the present application, and will not be elaborated here.

[0148] Fourthly, the embodiments of the present application further provide a readable storage medium.

[0149] A smart multipath transmission program based on delay synchronization constraints is stored on the readable storage medium of the present application. When the smart multipath transmission program based on delay synchronization constraints is executed by a processor, the steps of a smart multipath transmission method based on delay synchronization constraints as described above are implemented.

[0150] Among them, the method implemented when the smart multipath transmission program based on delay synchronization constraints is executed can refer to the various embodiments of the smart multipath transmission method based on delay synchronization constraints of the present application, which will not be elaborated here.

[0151] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0153] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. An intelligent multipath transmission method based on time-delay synchronization constraints, characterized in that, It includes the following steps: Step S1: Construct a multi-path transmission intelligent decision-making model; Step S2: Design a multi-dimensional state space for multi-path transmission; Step S3: Establish a multi-objective optimization reward function based on delay synchronization constraints; Step S4: Use the designed multi-dimensional state space and the established multi-objective optimization reward function to drive the intelligent exploration of multiple source-destination node pairs, and generate historical sequence sample data including the states, actions, rewards, and next-moment states of the source-destination node pairs during the path-finding process; Step S5: Screen out some samples from the historical sequence sample data and put them into the experience replay pool of the multi-path transmission intelligent decision-making model, and deploy the trained multi-path transmission intelligent decision-making model to generate a routing decision for the multi-path network based on delay synchronization constraints.

2. The intelligent multipath transmission method based on delay synchronization constraint according to claim 1, characterized in that The said Step S1: Construct a multi-path transmission intelligent decision-making model, which specifically includes the following steps: Use the PPO model as the basic model of the multi-path selection intelligent agent, add 2 LSTM models respectively before the Actor and Critic networks of the PPO model, and add an experience replay pool applicable to the PPO algorithm to construct a multi-path transmission intelligent decision-making model based on the PPO-LSTM architecture.

3. The intelligent multipath transmission method based on delay synchronization constraint according to claim 1, characterized in that The said Step S2: Design a multi-dimensional state space for multi-path transmission, which specifically includes the following steps: S21: Use the Dijkstra algorithm to calculate the minimum number of hops from the directly adjacent neighbor nodes of the current node to the destination node as the hop count cost; S22: Sum the transmission delay, propagation delay, and queuing delay to calculate the delay cost of the transmission delay cost between the current node and the neighbor node; S23: Evaluate the network congestion degree based on the average queue occupancy rate of the second-order neighbor nodes of the current node; S24: Obtain other states including the path loop state, node intersection state, and link intersection state; S25: Integrate the hop count cost, delay cost, congestion degree, path loop state, node intersection state, and link intersection state to construct a multi-dimensional state space for multi-path transmission.

4. The intelligent multipath transmission method based on delay synchronization constraint according to claim 3, wherein The time delay cost D ij The specific calculation formula is as follows: where p ij represents the number of data packets transmitted between node i and adjacent node j, m represents the data packet size, Bw ij represents the link bandwidth between node i and adjacent node j, d ij represents the link distance, c represents the propagation rate of the signal, q ij represents the number of data packets queued in the queue, and Op represents the speed at which the router processes the queued data packets.

5. The intelligent multipath transmission method based on delay synchronization constraint according to claim 1, characterized in that The said Step S3: Establish a multi-objective optimization reward function based on delay synchronization constraints, which specifically includes the following steps: S31: Obtain the hop count reward; S32: Obtain the delay reward; S33: Obtain the congestion degree reward; S34: Obtain other case rewards including the path loop reward, node intersection reward, and link intersection reward; S35: Perform weighted summation on the obtained hop count reward, delay reward, congestion degree reward, path loop reward, node intersection reward, and link intersection reward to obtain a reward function based on multi-objective optimization.

6. The intelligent multipath transmission method based on delay synchronization constraints according to claim 1, wherein, The said Step S4: Use the designed multi-dimensional state space and the established multi-objective optimization reward function to drive the intelligent exploration of multiple source-destination node pairs, and generate historical sequence sample data including the states, actions, rewards, and next-moment states of the source-destination node pairs during the path-finding process, which specifically includes the following steps: S41: Randomly select a group of source-destination node pairs, choose one of the source-destination node pairs, start from the source node, obtain the neighbor node status of the current node, input it to the agent to obtain the routing action decision for multi-path transmission, execute the routing action decision, obtain the immediate reward for routing to the next node and the next state, and return them to the agent until the destination node is found or the maximum step limit is reached. Repeat this process to complete the pathfinding operation for all source-destination node pairs in this group; S42: Repeat step S41 to complete the pathfinding operation for all groups of source-destination node pairs in the multi-path, and generate historical sequence sample data containing the status, actions, rewards, and next moment states of all groups of source-destination node pairs during the pathfinding process.

7. The intelligent multipath transmission method based on delay synchronization constraint according to claim 1, wherein The step S5: Screen out some samples from the historical sequence sample data and put them into the experience replay pool of the multi-path transmission intelligent decision-making model. Deploy the trained multi-path transmission intelligent decision-making model to generate routing decisions for the multi-path network based on delay synchronization constraints, which specifically includes the following steps: S51: Screen out the samples that meet the screening conditions from the historical sequence sample data and put them into the experience replay pool; S52: Conduct mini-batch sampling from the experience replay pool, and calculate the ratio of the new and old policy probabilities and the discounted reward using the mini-batch collected data; S53: The Critic network uses the discounted reward as the actual state value and calculates the value estimation loss; S54: The Actor network uses the ratio of the new and old policy probabilities calculated from the sample data to limit the policy update amplitude output by the basic model; S55: According to the obtained ratio of the new and old policy probabilities, the discounted reward, and the value estimation loss, calculate the loss function of the constructed multi-path transmission intelligent decision-making model. Repeat steps S51 - S54. When the convergence condition is met, terminate the training, deploy the trained multi-path transmission intelligent decision-making model, and generate routing decisions for the multi-path network based on delay synchronization constraints.

8. The intelligent multipath transmission method based on delay synchronization constraints according to claim 7, characterized in that, The sample screening strategy in step S51 specifically includes: Calculate the ratio of the action-state probability distribution of each sample in the historical sample sequence to the current policy; Set the policy clipping factor ε and screen out the samples whose ratios are within the clipping range.

9. An intelligent multipath transmission system based on delay synchronization constraints, characterized in that, Including: The basic model construction module is used to construct the multi-path transmission intelligent decision-making model; The state space design module is used to design a multi-dimensional state space for multi-path transmission; The reward function establishment module is used to establish a multi-objective optimization reward function based on delay synchronization constraints; The historical sequence sample data acquisition module is communicatively connected to the state space design module and the reward function establishment module, and is used to drive the intelligent exploration of multiple groups of source-destination node pairs by using the designed multi-dimensional state space and the established multi-objective optimization reward function, and generate historical sequence sample data containing the status, actions, rewards, and next moment states of the source-destination node pairs during the pathfinding process. The model training and routing decision output module, which is communicatively connected to the model establishment module and the historical sequence sample data acquisition module, is configured to screen out some samples from the historical sequence sample data and put them into the experience replay pool of the multi-path transmission intelligent decision model, deploy the trained multi-path transmission intelligent decision model, and generate routing decisions for the multi-path network based on delay synchronization constraints.

10. The intelligent multipath transmission system based on delay synchronization constraint according to claim 9, characterized in that, The state space design module includes: The hop count cost acquisition unit is configured to calculate the minimum hop count from the directly adjacent neighbor nodes of the current node to the destination node as the hop count cost by using the Dijkstra algorithm; The node delay cost acquisition unit is configured to calculate the sum of the transmission delay, propagation delay, and queuing delay to obtain the delay cost of the transmission delay cost between the current node and the neighbor node; The network congestion degree acquisition unit is configured to evaluate the network congestion degree based on the average queue occupancy rate of the second-order neighbor nodes of the current node; The other state acquisition unit is configured to acquire other states including the path loop state, node intersection state, and link intersection state; The multi-dimensional state space construction unit, which is communicatively connected to the hop count cost acquisition unit, the node delay cost acquisition unit, the network congestion degree acquisition unit, and the other state acquisition unit, is configured to integrate the hop count cost, delay cost, congestion degree, path loop state, node intersection state, and link intersection state to construct a multi-dimensional state space for multi-path transmission.