Unmanned aerial vehicle task network message transmission route planning system and method based on deep reinforcement learning
Through the routing planning system of deep reinforcement learning, combined with mission intent and situational awareness, the DQN algorithm is used to optimize the routing planning of the UAV mission network, which solves the problems of dynamic adaptability and computational complexity in the UAV mission network and realizes efficient and reliable message transmission.
Patent Information
- Application Number
- CN202511026371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies lack dynamic adaptability and computational complexity in drone mission networks, making it difficult for traditional routing algorithms to adjust paths in real time, affecting the reliability and efficiency of message transmission. Traditional reinforcement learning methods are not effective in high-dimensional decision-making problems.
A routing planning system based on deep reinforcement learning is adopted. Through the mission intention translation module, situation awareness module and routing planning module, the DQN algorithm is used to perform routing planning in the UAV mission network, combining QoS requirements and link situation information to achieve dynamic adjustment and optimization.
It improves the message transmission efficiency of the UAV mission network, reduces latency and cost, enhances the adaptability and reliability of the network, and can adaptively adjust the transmission path in a dynamic environment.
Smart Images

Figure CN120639690A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of UAV mission network communication technology, and specifically relates to a UAV mission network message transmission routing planning system and method based on deep reinforcement learning. Background Art
[0002] The rapid development of aviation, electronics, automatic control, and artificial intelligence technologies has fueled the rapid growth of drone technology. Drones are widely used in various civil and military applications, such as real-time surveillance, search and rescue, military reconnaissance, and hazardous site inspections. With the advancement of drone technology and the emergence of numerous miniaturized drones, drone applications are increasingly embracing multi-drone or swarm operations. Collaborative drones communicate using wireless communication devices and dynamically form networks through self-organization. With increasingly stringent mission requirements and increasingly complex environments, the number of nodes required for self-organizing drone networks has increased significantly, resulting in large-scale and highly dynamic networks. This is primarily manifested in the dramatic increase in the number of messages transmitted between nodes, which not only increases the network's communication burden but also significantly increases routing overhead. This undoubtedly impacts the efficiency and accuracy of drone missions and places higher demands on the efficiency and dynamic adaptability of routing algorithms to adapt to rapidly changing environments.
[0003] Through the above analysis, the problems and defects of the existing technology are as follows:
[0004] (1) Existing technology: Route planning is performed through traditional routing algorithms based on static network topology. Traditional routing algorithms mainly rely on pre-built network topology maps and perform routing selection based on network status parameters. They can effectively select the optimal transmission path for mission messages and have high reliability and efficiency in a stable network environment. However, traditional routing algorithms have significant limitations and are not able to adapt to complex environmental changes. In the scenario of UAV mission networks, the network topology and communication links change all the time. Traditional routing algorithms find it difficult to effectively capture these changes and cannot achieve dynamic adaptation, which in turn affects the reliability and timeliness of mission message transmission.
[0005] (2) In the second existing technology, the network behavior and response pattern are continuously learned through the reinforcement learning algorithm to carry out message transmission route planning. When encountering dynamic changes in the network or link failures, the message transmission path can be adjusted and optimized in real time to ensure that the mission message can be transmitted reliably and timely, and reduce the impact of network state changes on message transmission. However, the limitation of this technology is that the high dynamics of the UAV mission network leads to the high complexity of its state, and it is difficult to solve its high-dimensional decision-making problem using traditional reinforcement learning methods. The strong dependence on a large amount of real-time data and the complex calculations lead to poor model training results and the inability to guarantee the accuracy of route planning, thereby affecting the overall mission efficiency.
[0006] Through the above analysis, the problems and defects of the existing technology are as follows:
[0007] 1. Lack of dynamic adaptability: Existing technologies rely on pre-built network topologies, which exhibit poor adaptability to the dynamic changes in UAV mission networks and make it difficult to adjust and optimize paths in real time. Dynamic network changes and link failures can significantly affect the effectiveness of the original message transmission path, thereby affecting the efficient and accurate transmission of mission messages.
[0008] 2. Computational complexity and poor model training: Existing technology relies heavily on real-time data and complex routing models. This reliance makes the system prone to errors when data is insufficient or of poor quality, leading to incorrect planning of message transmission paths. In highly dynamic or complex environments, model training suffers from the "curse of dimensionality," which can reduce training effectiveness and impact routing decisions. This makes it unsuitable for the efficient and accurate requirements of UAV mission networks. Summary of the Invention
[0009] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a system and method for routing planning of drone mission network message transmission based on deep reinforcement learning, using DQN in deep reinforcement learning for routing planning to solve the "dimensionality curse" problem caused by the high dynamics of drone mission networks. It realizes dual drive of intention and situation. When network resources are limited, in order to ensure the delivery of high-value mission messages, by analyzing the QoS requirements and link situation information in the task intentions of different drone mission networks, DQN is used to obtain a transmission path suitable for network messages from the source node to the destination node, so as to improve the efficiency of network message transmission and reduce delays and costs. At the same time, when the network situation or node status changes, adjustments and optimizations are made to ensure that the mission messages are transmitted in a timely and reliable manner.
[0010] In order to achieve the above object, the technical solution adopted by the present invention is:
[0011] A UAV mission network message transmission routing planning system based on deep reinforcement learning, including a mission intent translation module, a situational awareness module, and a routing planning module;
[0012] The mission intent translation module obtains key mission intent information from the drone business or drone mission input by the user through knowledge extraction technology, and combines the strategy knowledge graph and network situation knowledge graph to perform network demand mapping to obtain the QoS requirements of related mission messages;
[0013] The situation awareness module collects key data of network situation information and centrally processes and manages the collected link delay, packet loss rate, bandwidth and load-related data, comprehensively evaluates the situation of each path in the network, and forms a resource situation database to provide a basis for the subsequent routing planning module;
[0014] The routing planning module analyzes the QoS requirements and network link status information in different network task intentions, and uses deep reinforcement learning to obtain a transmission path suitable for network messages from the source node to the destination node, thereby improving network message transmission efficiency, reducing latency and costs. At the same time, it adjusts and optimizes when the network situation or node status changes to ensure that task messages are transmitted in a timely and reliable manner.
[0015] Furthermore, the task intent translation module is divided into two submodules: a knowledge extraction module and a requirement mapping module; the knowledge extraction module extracts key knowledge from the input intent;
[0016] The input intent can be processed through named entity recognition to extract key entity information from irregular Chinese intents and output translated intent tuples to achieve standardized representation.
[0017] The demand mapping module combines the extracted intention tuples with the strategy knowledge graph and the network situation knowledge graph to perform network demand mapping, obtains relevant task message requirements, and sends the results to the routing planning module to provide a basis for it.
[0018] Furthermore, the situation awareness module includes six links: data collection, data processing and detailed representation, information transmission, information fusion, and generation and maintenance of situation database;
[0019] The situation awareness module first realizes the collection of network situation-related data (network bandwidth, latency, packet loss rate, channel load and other parameters) during network operation to achieve data collection;
[0020] The collected data is then processed and refined, and written into data packets for information transmission, sharing, and information fusion to form link situation information, with the goal of comprehensively evaluating the situation of each path.
[0021] Finally, a situation database is generated and dynamically maintained, and the link situation information in the database is the data basis for the subsequent routing planning module.
[0022] Furthermore, in the routing planning module, the node set in the network, the intent tuple obtained by the task intent translation module, the task message transmission QoS requirement information and the link situation information obtained by the situation awareness module are used as initial inputs of the routing planning module;
[0023] The routing planning module uses the deep reinforcement learning algorithm to find multiple reachable paths P(s,d)={P1,P2,P3,...,P n} to find the optimal transmission path, and the obtained path meets the QoS requirements of the task message. When the link situation changes, the algorithm will promptly re-plan the transmission path that does not meet the QoS requirements to ensure the real-time and accurate transmission of the task message.
[0024] Furthermore, the deep reinforcement learning algorithm in the routing planning module uses DQN for routing planning, selecting transmission efficiency and communication quality as planning objectives. Transmission efficiency is measured by message transmission delay, packet loss rate and bandwidth, and communication quality is measured by link delay, bandwidth, packet loss rate, load and jitter.
[0025] The state space of the deep reinforcement learning algorithm contains link status information and mission message requirements, while the action space contains the next-hop node for message transmission. The intelligent agent autonomously perceives the surrounding environment to achieve situational awareness and adaptively generates an effective transmission path based on planning decisions and model training. To obtain the optimal transmission path and ensure that the mission intent requirements are met, the system needs to dynamically interact with the network environment. When acquiring the state space, a set of state information is generated at each sampling interval. Due to the high dynamic mobility of drones, the amount of state information is enormous. Therefore, a nonlinear approximation method using a neural network is used to approximate the Q function, thereby addressing the curse of dimensionality. Each intelligent node has a convolutional neural network (CNN) with the same structure. The CNN is trained to fit the state-action values under high-dimensional continuous states.
[0026] Different experiences in the experience pool of the algorithm have different effects on the agent's strategy learning. In order to make full use of high-quality experience to improve the efficiency of model convergence, a priority experience replay mechanism is introduced. The priority of each experience sample is determined according to the state estimation error (temporal-difference, TD error) of the sample at each time step in the DQN algorithm. In the early stage of the training process, a transmission path is randomly selected to fill the experience pool with data. Afterwards, the agent uses a greedy strategy to select actions based on the information in the state space, and continuously stores experience samples of the agent's interaction with the environment in the experience pool, so that it can fully explore the surrounding environment. The online decision result of CNN is the optimal transmission path for the task message available at the current moment. During the learning stage, the priority weight is randomly selected after segmenting according to the number of experience samples, and the experience corresponding to the weight is selected.
[0027] A method for planning routing of message transmission in a UAV mission network based on deep reinforcement learning, comprising the following steps:
[0028] S101, the user inputs the drone mission and sends it to the mission intention translation module;
[0029] S102: The UAV mission is translated into a mission intent by the mission intent translation module, and the knowledge extraction module extracts knowledge to achieve a standardized representation. The strategy mapping module outputs the intent as a relevant mission message requirement, and the result is sent to the routing planning module to provide a basis for it.
[0030] S103, the situation awareness module collects key data of network situation information and processes and manages it to form a resource situation database, providing a data basis for the routing planning module;
[0031] S104, the routing planning module uses a deep reinforcement learning algorithm to find multiple reachable paths P(s,d)={P1,P2,P3,...,P n} to find the optimal transmission path, and the obtained path meets the QoS requirements of the task message;
[0032] S105 , if the network situation changes or a link fails, adjustments and optimizations are performed to reduce the impact of the network dynamic changes or link failures on network message transmission.
[0033] In S101, the user inputs the drone mission or business through the front-end interface, and then sends it to the mission intention translation module for translation processing.
[0034] The S102 is specifically as follows:
[0035] The user inputs the task intention in natural language;
[0036] In the first stage, the knowledge extraction module extracts key knowledge from the input intent. If the given intent is complete and valid, named entity recognition is performed to extract key entity information from the irregular Chinese intent and output the translated intent tuple to achieve standardized representation.
[0037] In the second stage, the strategy mapping module combines the extracted intent tuples with the strategy knowledge graph and the network status knowledge graph to map the network requirements, obtain relevant intent requirements, and send the results to the routing planning module in S104 for further configuration.
[0038] The S103 is specifically as follows:
[0039] The situation awareness module includes six links, namely data collection, data processing and detailed representation, information transmission, information fusion, and generation and maintenance of situation database;
[0040] The situation awareness module first collects network situation-related data (network bandwidth, latency, packet loss rate, and channel load parameters) during network operation to achieve data collection. It then processes and refines the collected data and writes it into data packets for information transmission, sharing, and fusion processing to form link situation information.
[0041] The information fusion process is as follows:
[0042] Assume G = (V, E) to represent the network topology model, where V represents the set of all network nodes in the network model, E represents the set of links between network nodes, and each link represents a direct path between two adjacent network nodes; Assume the source node is s (s∈V), the destination node is d (d∈V), and there are multiple transmission paths P (s, d) = {P1, P2, P3, ..., P n}, where P i ={e i1 ,e i2 ,...,e im For any transmission path in the network model, four QoS metric functions are given, namely delay (P), bandwidth (P), packet loss rate (P), and load (P).
[0043] Definition 1 (Delay): For any transmission path P from source node s to destination node d i , the delay calculation formula is as follows:
[0044]
[0045] Among them, delay(e ik ) is the transmission path P i Each link e ik Delay;
[0046] Definition 2 (Bandwidth): For any transmission path P from source node s to destination node d i , the bandwidth calculation formula is as follows:
[0047] bandwidth(P i )=min{bandwidth(e ik ),e ik ∈P i},
[0048] Among them, bandwidth(e ik ) is the transmission path P i Each link e ik bandwidth;
[0049] Definition 3 (Packet Loss Rate): For any transmission path P from source node s to destination node d i , the packet loss rate calculation formula is as follows:
[0050]
[0051] Among them, loss(e ik ) is the transmission path P i Each link e ik Packet loss rate;
[0052] Definition 4 (Load): For any transmission path P from source node s to destination node d i , the load calculation formula is as follows:
[0053]
[0054] Among them, Hop(P i ) is the path P i The number of hops, load(e ik ) is the transmission path P i Each link e ik load;
[0055] Finally, a situation database is generated and dynamically maintained, and the link situation information in the database is the data basis for the subsequent routing planning module.
[0056] The S104 is specifically as follows:
[0057] The Deep Q-Network (DQN) in deep reinforcement learning is used for routing planning. Considering the difficulty of routing planning caused by the multiple QoS requirements of the UAV mission network task messages, the transmission efficiency Lr and communication quality Lq are selected as the routing planning objectives of the UAV mission network message transmission. Among them, the transmission efficiency Lr is measured by the delay, packet loss rate and bandwidth of the message transmission path, and the communication quality Lq is measured by the delay, bandwidth, packet loss rate and load of the link. The objective function L is calculated. r and L q :
[0058]
[0059] L q =w1*delay(P i )+w2*bandwidth(P i )+w3*loss(P i )+w4*load(P i ),
[0060] w1+w2+w3+w4=1,
[0061] Among them, loss(P i ) represents the packet loss rate of the transmission path, bandwidth (P i ) represents the bandwidth of the transmission path, delay(P i ) represents the delay of the transmission path, load(P i ) represents the load of the transmission path, and w1, w2, w3, and w4 represent the weights of the link quality.
[0062] The state space contains link status information and message transmission requirements, which is defined as O = [o1, o2, ..., o k ] T , where o n =[delay(e ik ),bandwidth(e ik ),loss(e ik ),load(e ik ),D,B,Loss,Load] T is the transmission path P i Each link e ik The action space contains the next hop node information of the message transmission, which is defined as a n =(next_node n ); In order to improve the transmission efficiency and communication quality of task message transmission, the reward function R(s t ,a t ) is defined as:
[0063] R(s t ,a t )=α*L r +β*L q ,
[0064] α+β=1,
[0065] Among them, α and β are the transmission efficiency L r and communication quality L q The weight of
[0066] The intelligent agent autonomously perceives the surrounding environment to achieve situational awareness and adaptively generates an effective transmission path based on planning decisions and model training. To obtain the optimal transmission path and ensure that the mission intent is met, the system needs to dynamically interact with the network environment. When acquiring the state space, a set of state information is generated at each sampling interval. A nonlinear approximation method using a neural network is used to approximate the Q function, thereby resolving the curse of dimensionality. Each intelligent node has a convolutional neural network (CNN) with the same structure. By training the CNN, the state action value under high-dimensional continuous state is fitted.
[0067] The Q function of node n at time t is expressed as Indicates that node n is in state The maximum long-term accumulated reward value that can be obtained by performing action a under the condition of , the Q value update equation is expressed as
[0068]
[0069] Among them, α is the learning rate and γ is the discount factor;
[0070] To minimize the gap between the estimated Q value and the target Q value, the loss function of DQN is defined as
[0071]
[0072] in, is the weight coefficient of the predicted CNN of node n;
[0073] The target Q value is
[0074]
[0075] in, is the weight coefficient of the target CNN model;
[0076] The Q value at each moment is updated using a CNN consisting of an input layer, a convolutional layer (Conv), a flattening layer, a fully connected layer (FC), and an output layer. Conv has 16 filters of size 10×4, a stride of 2, and uses tanh as the activation function; FC has 512 units and a sigmoid activation function; the activation function of the output layer is a linear activation function. The CNN training process uses the gradient descent method, and the gradient of the loss function is expressed as
[0077]
[0078] Different experiences in the DQN experience pool have different effects on the agent's strategy learning. In order to make full use of high-quality experience to improve the model convergence efficiency, a priority experience replay mechanism is introduced. The priority pj of each experience sample j is calculated based on the state estimation error (temporal-difference, TD error) δ at each time step. j OK. Experiences with large TD error will be replayed more frequently. j If it is 0, the experience sample j cannot be sampled, and a small positive number ζ is added to it.
[0079] p j =|δ j |+ζ
[0080] The replay probability p(j) of experience j is:
[0081]
[0082] Among them, α is the parameter that controls the priority. When α is 0, the priority experience replay becomes random sampling. To alleviate the loss of sampling diversity, random priority sampling is used. Since biased sampling introduces bias, to ensure that the biased sampling learning strategy is the same as the uniform sampling strategy, the importance sampling weight w is used. j Correct the deviation; considering stability, use the maximum weight value for normalization; β is the annealing factor, which will anneal from the initial value to 1 at the end of learning, and work together with α to correct the deviation, and n is the number of samples.
[0083]
[0084] During the training process, transmission paths are randomly selected in the first K moments to fill the experience pool with data. Afterwards, the agent uses a greedy strategy to select actions based on the information in the state space and continuously stores experience samples of the agent's interaction with the environment in the experience pool, allowing the agent to fully explore the surrounding environment. The CNN's online decision result is the optimal transmission path for the task message available at the current moment. During the learning phase, priority weights are randomly selected after being segmented according to the number of experience samples, and the experience corresponding to the weight is selected.
[0085] The specific steps of S105 are:
[0086] The network status is monitored in real time. When the network situation changes or a link fails, causing the transmission path to no longer meet the QoS requirements of task message transmission, DQN is used to re-plan the route and adjust and optimize the transmission path based on the current network situation information to reduce the impact of dynamic network changes or link failures, enabling the network to have real-time decision-making and optimization capabilities.
[0087] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps in a method for planning network message transmission routes for unmanned aerial vehicle missions based on deep reinforcement learning are implemented.
[0088] An information data processing system provides a machine learning environment with the ability of self-learning and self-generation, providing an intelligent foundation for the above method.
[0089] Beneficial effects of the present invention:
[0090] This invention can achieve a deep understanding of drone mission intent, automatically plan message transmission paths on demand, and achieve accurate and efficient distribution of drone mission messages. It can also adaptively adjust drone mission message distribution paths based on real-time changes in the network environment, improving the efficiency of the drone mission network and meeting the high-performance requirements of missions in dynamic environments. It also helps managers or users adjust network resource deployment in a timely manner, thereby improving the adaptability and reliability of the drone mission network.
[0091] The present invention introduces a priority experience replay mechanism. The priority of each experience sample is determined according to the state estimation error (temporal-difference, TD error) at each time step. In the early stage of the training process, a transmission path is randomly selected to fill the experience pool with data. Afterwards, the agent uses a greedy strategy to select actions based on the information in the state space, and continuously stores experience samples of the agent's interaction with the environment in the experience pool, so that it can fully explore the surrounding environment. The online decision result of CNN is the optimal transmission path for the task message available at the current moment. During the learning stage, the priority weight is randomly selected after segmenting according to the number of experience samples, and the experience corresponding to the weight is selected. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 It is a flow chart of the system of the present invention.
[0093] Figure 2 These are the specific implementation steps of the method of the present invention.
[0094] Figure 3 It is a specific implementation flow chart of the task intention translation module of the present invention.
[0095] Figure 4 It is a specific implementation flow chart of the situation awareness module of the present invention.
[0096] Figure 5 This is a specific implementation flow chart of the deep reinforcement learning algorithm of the present invention.
[0097] Figure 6 This is a schematic diagram comparing the present invention with AODV and OLSR routing protocols. DETAILED DESCRIPTION
[0098] The present invention will be described in further detail below with reference to the accompanying drawings.
[0099] In response to the problems existing in the prior art, the present invention provides a system and method for planning network message transmission routing for UAV missions based on deep reinforcement learning. The present invention is described in detail below with reference to the accompanying drawings.
[0100] like Figure 1 As shown in the figure, it includes a task intent translation module, a situation awareness module, and a routing planning module. The task intent translation module obtains key information of the task intent through knowledge extraction technology, and combines the policy knowledge graph and the network situation knowledge graph to map the network requirements and obtain the relevant task message requirements. The situation awareness module collects key data of network situation information, and centrally processes and manages the collected data, comprehensively evaluates the situation of each path in the network, and forms a resource situation database, thereby providing a basis for the subsequent routing planning module. The routing planning module analyzes the QoS requirements and link situation information in different network task intents, and uses a deep reinforcement learning algorithm to obtain a transmission path suitable for network messages from the source node to the destination node, thereby improving the efficiency of network message transmission and reducing latency and cost. At the same time, adjustments and optimizations are made when the network situation or node status changes to ensure the timely and reliable transmission of task messages.
[0101] like Figure 2 As shown, the method for planning the routing of network message transmission for UAV missions based on deep reinforcement learning provided by an embodiment of the present invention specifically includes the following steps:
[0102] S101: The user inputs a drone mission and sends it to the mission intent translation module.
[0103] S102, the UAV mission is standardized through knowledge extraction through the mission intention translation module, and the intent is output as relevant mission message requirements through demand mapping, and the result is sent to the routing planning module to provide a basis for it.
[0104] S103, the situation awareness module collects key data of network situation information and processes and manages it to form a resource situation database, providing a data basis for the routing planning module.
[0105] S104, the routing planning module uses a deep reinforcement learning algorithm to find multiple reachable paths P(s,d)={P1,P2,P3,...,P n} to find the optimal transmission path, and the obtained path meets the QoS requirements of the task message.
[0106] S105 , if the network situation changes or a link fails, adjustments and optimizations are performed to reduce the impact of the network dynamic changes or link failures on network message transmission.
[0107] like Figure 3 As shown, the specific implementation process of the task intention translation module in S101 provided by the embodiment of the present invention includes:
[0108] The user inputs the task intent in natural language. In the first stage, the knowledge extraction module extracts key knowledge from the input intent. If the given intent is complete and valid, it can use named entity recognition to extract the key entity information in the irregular Chinese intent and output the translated intent tuple to achieve standardized representation. In the second stage, the strategy mapping module combines the extracted intent tuple with the strategy knowledge graph and the network status knowledge graph to map the network requirements, obtain the relevant intent requirements, and send the results to the routing planning module in S104 for further configuration.
[0109] like Figure 4 As shown, the specific implementation process of the situation awareness module in S103 provided in the embodiment of the present invention includes:
[0110] The six steps are data collection, data processing and detailed representation, information transmission, information fusion, and generation and maintenance of a situation database. The situation awareness module first collects network situation-related data (network bandwidth, latency, packet loss rate, and channel load parameters) during network operation. It then refines and represents the collected data and writes it into data packets for transmission, sharing, and fusion processing to form link situation information.
[0111] The centralized fusion processing of situation information is as follows: Let G = (V, E) represent the network topology model, where V represents the set of all network nodes in the network model, E represents the set of links between network nodes, and each link represents a direct path between two adjacent network nodes. Let the source node be s (s∈V), the destination node be d (d∈V), and there are multiple transmission paths P (s, d) = {P1, P2, P3, ..., P n}, where P i ={e i1 ,e i2 ,...,e im For any transmission path in the network model, four QoS measurement functions are given, namely delay function delay(P), bandwidth(P), packet loss rate loss(P), and load(P).
[0112] Definition 1 (Delay): For any transmission path P from source node s to destination node d i , the delay calculation formula is as follows:
[0113]
[0114] Definition 2 (Bandwidth): For any transmission path P from source node s to destination node d i , the bandwidth calculation formula is as follows:
[0115] bandwidth(P i )=min{bandwidth(e ik ),e ik ∈P i}
[0116] Definition 3 (Packet Loss Rate): For any transmission path P from source node s to destination node d i , the packet loss rate calculation formula is as follows:
[0117]
[0118] Definition 4 (Load): For any transmission path P from source node s to destination node d i , the number of hops of the path Hop (P), the load calculation formula is as follows:
[0119]
[0120] Finally, a situation database is generated and dynamically maintained, and the link situation information in the database is the data basis for the subsequent routing planning module.
[0121] like Figure 5 As shown, the specific implementation process of the S104 deep reinforcement learning algorithm provided by the embodiment of the present invention includes:
[0122] The Deep Q-Network (DQN) in deep reinforcement learning is used for route planning. Considering the difficulty of route planning caused by the multiple QoS requirements of the UAV mission network task messages, the transmission efficiency L is selected. r and communication quality L q As the UAV mission network message transmission routing planning target. Among them, the transmission efficiency L r The communication quality L is measured by the message transmission delay, packet loss rate and bandwidth. q It is measured by link delay, bandwidth, packet loss rate and load. Calculate the objective function L r and L q :
[0123]
[0124] Lq=w1*delay(P i)+w2*bandwidth(P i )+w3*loss(P i )+w4*load(P i )
[0125] w1+w2+w3+w4=1
[0126] The state space contains link status information and message transmission requirements, which is defined as O = [o1, o2, ..., o k ] T , where o n =[delay(e ik ),bandwidth(e ik ),loss(e ik ),load(e ik ),D,B,Loss,Load] T is the transmission path P i Each link e ik The action space contains the next hop node information of the message transmission, which is defined as a n =(next_node n ). In order to improve the transmission efficiency and communication quality of task message transmission, the reward function is defined as:
[0127] R(s t ,a t )=α*L r +β*L q
[0128] α+β=1
[0129] The intelligent agent autonomously perceives the surrounding environment to achieve situational awareness and adaptively generates an effective transmission path based on planning decisions and model training. To obtain the optimal transmission path and ensure that the mission intent is met, the system needs to dynamically interact with the network environment. When acquiring the state space, a set of state information is generated at each sampling interval. A nonlinear neural network approximation method is used to approximate the Q function, thereby resolving the curse of dimensionality. Each intelligent node has a convolutional neural network (CNN) with the same structure. The CNN is trained to fit the state-action values of high-dimensional continuous states.
[0130] The Q function of node n at time t is expressed as Indicates that the node is in state The maximum long-term accumulated reward value that can be obtained by performing action a under the Q value update equation is expressed as
[0131]
[0132] Among them, α is the learning rate and γ is the discount factor.
[0133] To minimize the gap between the estimated Q value and the target Q value, the loss function of DQN is defined as
[0134]
[0135] in, is the weight coefficient of the predicted CNN for node n.
[0136] The target Q value is
[0137]
[0138] in, is the weight coefficient of the target CNN model.
[0139] To update the Q value at each moment, a CNN consisting of an input layer, a convolutional layer (Conv), a flattening layer, a fully connected layer (FC), and an output layer is used. Conv has 16 filters of size 10×4, a stride of 2, and uses tanh as the activation function; FC has 512 units and a sigmoid activation function; the output layer's activation function is a linear activation function. The CNN training process uses the gradient descent method, and the gradient of the loss function is expressed as
[0140]
[0141] Different experiences in the DQN experience pool have different effects on the agent's strategy learning. In order to make full use of high-quality experience to improve the model convergence efficiency, a priority experience replay mechanism is introduced. The priority p of each experience sample j j According to the state estimation error (temporal-difference, TD error) δ at each time step j OK. Experiences with large TD error will be replayed more frequently. j If it is 0, the experience sample j cannot be sampled, and a small positive number ζ is added to it.
[0142] p j =|δ j |+ζ
[0143] The replay probability p(j) of experience j is:
[0144]
[0145] Among them, α is a parameter that controls the priority. When α is 0, the priority experience replay becomes random sampling. To alleviate the loss of sampling diversity, random priority sampling is used. Since biased sampling introduces bias, to ensure that the biased sampling learning strategy is the same as the uniform sampling strategy, the importance sampling weight w is used. j Correcting bias. Considering stability, normalization is performed using the maximum weight value. β is the annealing factor, which anneals from the initial value to 1 at the end of learning and works together with α to correct bias. n is the number of samples.
[0146]
[0147] During training, transmission paths are randomly selected for the first K moments to populate the experience pool. The agent then uses a greedy strategy to select actions based on information in the state space and continuously stores experience samples of the agent's interactions with the environment in the experience pool, allowing the agent to fully explore its surroundings. The CNN's online decision result is the optimal transmission path for the task message available at the current moment. During the learning phase, priority weights are randomly selected based on the number of experience samples, and the experience corresponding to each weight is selected.
[0148] The technical effects of the present invention are described in detail below with reference to simulations.
[0149] This simulation adopts a modular design to build a UAV mission network simulation platform. Simulations are performed in QT 5.15.2. PyCharm is used to implement intelligent routing planning based on deep reinforcement learning algorithms. Through the interaction between the two, dynamic and adaptive routing planning is achieved by fully considering the QoS requirements of mission messages and network status information. The results are compared with the AODV and OLSR routing protocols.
[0150] like Figure 6 As shown in the figure, the average route generation time for different mission messages in the same network scenario is statistically analyzed. The average route generation time of the routing planning method of the present invention is relatively stable, generally controlled within 1.2 seconds. This shows that this method can quickly plan routing strategies in a dynamic UAV mission network environment and has good real-time performance. In contrast, the average route generation time of the AODV and OLSR routing protocols fluctuates greatly.
[0151] 1. To address the issue of poor dynamic adaptability, this invention introduces an intent-driven network. Based on a mission intent translation module and a situational awareness module, this system translates drone missions into message QoS requirements and generates routing strategies. The network status is then continuously monitored in real time. If changes in the network status or link failures cause the generated routing strategy to no longer meet message QoS requirements, routing planning is re-performed to adjust and optimize the routing strategy. This reduces the impact of dynamic network changes or link failures and improves the dynamic adaptability of routing planning.
[0152] 2. To address the problems of complex calculations and poor model training effects, this paper uses deep reinforcement learning for intelligent route planning, merges neural networks into intelligent agents in the reinforcement learning framework, solves the high-dimensional decision-making problems of drone mission networks, ensures the accuracy of route planning, improves network message transmission efficiency, and reduces latency and costs.
[0153] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0154] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A UAV mission network message transmission routing planning system based on deep reinforcement learning, characterized by: It includes mission intention translation module, situational awareness module and routing planning module; The mission intent translation module obtains key mission intent information from the drone business or drone mission input by the user through knowledge extraction technology, and combines the strategy knowledge graph and network situation knowledge graph to perform network demand mapping to obtain the QoS requirements of related mission messages; The situation awareness module collects key data of network situation information and centrally processes and manages the collected link delay, packet loss rate, bandwidth and load-related data, comprehensively evaluates the situation of each path in the network, and forms a resource situation database to provide a basis for the subsequent routing planning module; The routing planning module analyzes the QoS requirements and network link status information in different network task intentions, and uses deep reinforcement learning to obtain a transmission path suitable for network messages from the source node to the destination node, thereby improving network message transmission efficiency, reducing latency and costs; at the same time, it adjusts and optimizes when the network situation or node status changes to ensure the transmission of task messages.
2. The UAV mission network message transmission routing planning system based on deep reinforcement learning according to claim 1 is characterized in that: The task intention translation module is divided into two submodules: knowledge extraction module and requirement mapping module; the knowledge extraction module extracts key knowledge from the input intention; The input intent undergoes named entity recognition to extract key entity information from irregular Chinese intents and outputs the translated intent tuple to achieve standardized representation. The demand mapping module combines the extracted intention tuples with the strategy knowledge graph and the network situation knowledge graph to perform network demand mapping, obtains relevant task message requirements, and sends the results to the routing planning module to provide a basis for it.
3. The UAV mission network message transmission routing planning system based on deep reinforcement learning according to claim 1 is characterized in that: The situation awareness module performs six steps: data collection, data processing and detailed representation, information transmission, information fusion, and generation and maintenance of the situation database; The situational awareness module first realizes the collection of data related to link delay, packet loss rate, bandwidth and load during network operation to achieve data collection; The collected data is then processed and refined, and written into data packets for information transmission, sharing, and fusion processing to form link situation information, with the goal of comprehensively evaluating the situation of each path. Finally, a situation database is generated and dynamically maintained, and the link situation information in the database is the data basis for the subsequent routing planning module.
4. The UAV mission network message transmission routing planning system based on deep reinforcement learning according to claim 1 is characterized in that: In the routing planning module, the node set in the network, the intention tuple obtained by the task intention translation module, the task message transmission QoS requirement information and the link situation information obtained by the situation awareness module are used as the initial input of the routing planning module; The routing planning module uses the deep reinforcement learning algorithm to find multiple reachable paths P(s,d)={P1,P2,P3,...,P n } to find the optimal transmission path, and the obtained path meets the QoS requirements of the task message. When the link situation changes, the transmission path that does not meet the QoS requirements is replanned to ensure real-time and accurate transmission of the task message.
5. A method for planning network message transmission routing for UAV missions based on deep reinforcement learning according to any one of claims 1 to 4, characterized in that: The following steps are included: S101, the user inputs the drone mission and sends it to the mission intention translation module; S102: The UAV mission is translated into a mission intent by the mission intent translation module, and the knowledge extraction module extracts knowledge to achieve a standardized representation. The strategy mapping module outputs the intent as a relevant mission message requirement, and the result is sent to the routing planning module to provide a basis for it. S103, the situation awareness module collects key data of network situation information and processes and manages it to form a resource situation database, providing a data basis for the routing planning module; S104, the routing planning module uses a deep reinforcement learning algorithm to find multiple reachable paths P(s,d)={P1,P2,P3,...,P n } to find the optimal transmission path, and the obtained path meets the QoS requirements of the task message; S105 , if the network situation changes or a link fails, adjustments and optimizations are performed to reduce the impact of the network dynamic changes or link failures on network message transmission.
6. A method for planning network message transmission routes for UAV missions based on deep reinforcement learning according to claim 5, characterized in that: In S101, the user inputs the drone mission or business through the front-end interface, and then sends it to the mission intention translation module for translation processing.
7. The method for planning the routing of UAV mission network message transmission based on deep reinforcement learning according to claim 5, characterized in that: The S102 is specifically as follows: The user inputs the task intention in natural language; In the first stage, the knowledge extraction module extracts key knowledge from the input intent. If the given intent is complete and valid, named entity recognition is performed to extract key entity information from the irregular Chinese intent and output the translated intent tuple to achieve standardized representation. In the second stage, the strategy mapping module combines the extracted intent tuples with the strategy knowledge graph and the network status knowledge graph to map the network requirements, obtain relevant intent requirements, and send the results to the routing planning module in S104 for further configuration.
8. The method for planning the routing of UAV mission network message transmission based on deep reinforcement learning according to claim 5, characterized in that: The S103 is specifically as follows: The situation awareness module includes six links, namely data collection, data processing and detailed representation, information transmission, information fusion, and generation and maintenance of situation database; The situation awareness module first collects data related to the network situation during network operation, realizes data collection, and then processes and refines the collected data and writes it into data packets for information transmission, sharing and fusion processing to form link situation information; The information fusion process is as follows: Assume G = (V, E) to represent the network topology model, where V represents the set of all network nodes in the network model, E represents the set of links between network nodes, and each link represents a direct path between two adjacent network nodes; Assume the source node is s (s∈V), the destination node is d (d∈V), and there are multiple transmission paths P (s, d) = {P1, P2, P3, ..., P n }, where P i ={e i1 ,e i2 ,...,e im For any transmission path in the network model, four QoS metric functions are given, namely delay (P), bandwidth (P), packet loss rate (P), and load (P). Delay Definition 1: For any transmission path P from source node s to destination node d i , the delay calculation formula is as follows: Among them, delay(e ik ) is the transmission path P i Each link e ik Delay; Bandwidth Definition 2: For any transmission path P from source node s to destination node d i , the bandwidth calculation formula is as follows: bandwidth(P i )=min{bandwidth(e ik ),e ik ∈P i }, Among them, bandwidth(e ik ) is the transmission path P i Each link e ik bandwidth; Packet loss rate definition 3: For any transmission path P from source node s to destination node d i , the packet loss rate calculation formula is as follows: Among them, loss(e ik ) is the transmission path P i Each link e ik Packet loss rate; Load Definition 4: For any transmission path P from source node s to destination node d i , the load calculation formula is as follows: Among them, Hop(P i ) is the path P i The number of hops, load(e ik ) is the transmission path P i Each link e ik load; Finally, a situation database is generated and dynamically maintained, and the link situation information in the database is the data basis for the subsequent routing planning module.
9. A method for planning routing for network message transmission in a UAV mission based on deep reinforcement learning according to claim 8, characterized in that: The S104 is specifically as follows: Use the deep Q network in deep reinforcement learning for route planning and select the transmission efficiency L r and communication quality L q As the UAV mission network message transmission routing planning target; among them, the transmission efficiency L r The communication quality L is measured by the delay, packet loss rate and bandwidth of the message transmission path. q Measured by link delay, bandwidth, packet loss rate and load; calculate the objective function L r and L q : L q =w1*delay(P i )+w2*bandwidth(P i )+w3*loss(P i )+w4*load(P i ), w1+w2+w3+w4=1, Among them, loss(P i ) represents the packet loss rate of the transmission path, bandwidth (P i ) represents the bandwidth of the transmission path, delay(P i ) represents the delay of the transmission path, load(P i ) represents the load of the transmission path, w1, w2, w3, w4 represent the weights of the link quality; The state space contains link status information and message transmission requirements, which is defined as O = [o1, o2, ..., o k ] T , where o n =[delay(e ik ),bandwidth(e ik ),loss(e ik ),load(e ik ),D,B,Loss,Load] T is the transmission path P i Each link e ik The action space contains the next hop node information of the message transmission, which is defined as a n =(next_node n ); reward function R(s t ,a t ) is defined as: R(s t ,a t )=α*L r +β*L q , α+β=1, Among them, α and β are the transmission efficiency L r and communication quality L q The weight of The intelligent agent autonomously perceives the surrounding environment to achieve situational awareness and adaptively generates effective transmission paths based on planning decisions and model training. When acquiring the state space, each sampling interval generates a set of state information. A nonlinear approximation method of a neural network is selected to approximate the Q function. Each intelligent node has a convolutional neural network with the same structure. By training the CNN, the state action value under the high-dimensional continuous state is fitted. The Q function of node n at time t is expressed as Indicates that node n is in state The maximum long-term accumulated reward value that can be obtained by performing action a under the condition of , the Q value update equation is expressed as Among them, α is the learning rate and γ is the discount factor; To minimize the gap between the estimated Q value and the target Q value, the loss function of DQN is defined as in, is the weight coefficient of the predicted CNN of node n; The target Q value is in, is the weight coefficient of the target CNN model; The CNN consisting of input layer, convolution layer (Conv), flattening layer, fully connected layer (FC) and output layer is used to update the Q value at each moment. Tanh is used as the activation function; the FC activation function is sigmoid; the activation function of the output layer is a linear activation function; the CNN training process uses the gradient descent method, and the gradient of the loss function is expressed as The priority pj of each experience sample j is calculated based on the state estimation error δ at each time step. j Determine that the experience with large state estimation error will be replayed more frequently; add a small positive number ζ to avoid j If it is 0, the experience sample j cannot be sampled. p j =|δ j |+g The replay probability p(j) of experience j is: Among them, α is a parameter that controls priority. When α is 0, the priority experience replay becomes random sampling. The random priority sampling method is used to alleviate the loss of sampling diversity. The importance sampling weight w is used. j Correct the deviation; use the maximum weight value for normalization; β is the annealing factor, which will anneal from the initial value to 1 at the end of learning, and work together with α to correct the deviation, and n is the number of samples; During the training process, transmission paths are randomly selected in the first K moments to fill the experience pool with data. Afterwards, the agent uses a greedy strategy to select actions based on the information in the state space and continuously stores experience samples of the agent's interaction with the environment in the experience pool, allowing the agent to fully explore the surrounding environment. The CNN's online decision result is the optimal transmission path for the task message available at the current moment. During the learning phase, priority weights are randomly selected after being segmented according to the number of experience samples, and the experience corresponding to the weight is selected.
10. A method for planning network message transmission routes for UAV missions based on deep reinforcement learning according to claim 9, characterized in that: The specific steps of S105 are: The network status is monitored in real time. When the network situation changes or a link fails, causing the transmission path to no longer meet the QoS requirements of task message transmission, DQN is used to re-plan the route and adjust and optimize the transmission path based on the current network situation information to reduce the impact of dynamic network changes or link failures, enabling the network to have real-time decision-making and optimization capabilities.
Citation Information
Cited By
Explanatable reinforcement learning decision-making system and method based on large language model enhancement
CN120722758A
Cross-domain multi-modal flow security interaction method based on intranet AI engine
CN121077801A
5G-A low-altitude communication link quality optimization and adaptive adjustment method based on AI prediction
CN121510061A