Adaptive network shunting method and system based on deep reinforcement learning

By adopting an adaptive network shunt method based on deep reinforcement learning in heterogeneous network environments, the problems of multi-protocol differences, multi-objective policy conflicts and insufficient adaptability of dynamic environments are solved, and efficient network shunt, stable service quality and strong security protection capabilities are achieved.

CN120128534AActive Publication Date: 2025-06-10SHANGHAI KEGUANG COMM TECH CO LTD

Patent Information

Application Number
CN202510352358.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-10
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In existing heterogeneous network environments, network shunt efficiency, large service quality fluctuations and lag in security protection caused by multi-protocol differences, multi-objective policy conflicts and insufficient adaptability to dynamic environments.

Method used

Adaptive network shunt method based on deep reinforcement learning is adopted to achieve efficient analysis and dynamic adaptation of cross-network protocols by building a protocol fingerprint library and a unified feature encoding mechanism. Based on the multi-agent differential game framework, coordinated the coordinated decision-making of traffic control, service quality optimization and security protection, and introduced a meta-learning-driven dynamic weight adjustment mechanism to enhance the system's adaptability to complex network environments.

Benefits of technology

Significantly improve heterogeneous data processing capabilities, effectively resolve conflicts in traditional single agent strategy, synchronously optimize network throughput and transmission stability, reduce the risk of service quality violations, realize efficient inference and real-time response of edge devices, and break through the response delay bottleneck of traditional solutions in burst traffic processing and security defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128534A_ABST
    Figure CN120128534A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of adaptive network shunting, and discloses an adaptive network shunting method and system based on deep reinforcement learning, and the method comprises the steps: obtaining the flow data of a heterogeneous network, building a protocol fingerprint database, and converting the data of different protocols into a unified network state feature vector; constructing a multi-agent reinforcement learning framework, including defining a state space, defining an action space and defining a reward function, and optimizing a strategy of each agent based on a differential game; adjusting the weight of each agent reward function by using a meta-learning model, training a multi-agent reinforcement learning framework by using deep reinforcement learning, and learning an optimal shunting strategy under different heterogeneous network conditions; and deploying the trained strategy to a heterogeneous network environment, and dynamically selecting an optimal network path according to the current network state. And the heterogeneous data processing capability is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of adaptive network traffic diversion, and in particular to an adaptive network traffic diversion method and system based on deep reinforcement learning. Background Art

[0002] In recent years, multi-agent reinforcement learning (MARL) and meta-learning technology have provided new ideas for solving complex network optimization problems, but they still face significant challenges in actual deployment. Although existing MARL solutions alleviate single-agent policy conflicts through distributed decision-making, they lack quantitative modeling of the game relationship between agents, resulting in low efficiency in solving Nash equilibrium. For example, in the scenario of traffic control and security agent collaboration, the convergence time of traditional collaborative training methods is more than 3 times longer than that of single-agent solutions because they do not consider the policy coupling effect. At the protocol processing level, existing studies have attempted to extract cross-protocol commonalities (such as delay-sensitive service identifiers) through feature engineering, but manually designed features are difficult to cover dynamic protocol interaction patterns. In addition, existing dynamic weight adjustment methods rely on manually set thresholds and are difficult to adapt to nonlinear changes in the network environment. In edge computing scenarios, the high computational complexity of deep reinforcement learning models and the resource limitations of edge devices form a sharp contradiction, forcing existing solutions to compromise between model accuracy and deployment costs, resulting in a decrease in the performance of diversion strategies. At the same time, the dynamic topology of heterogeneous networks (such as intermittent connections of low-orbit satellite nodes) places higher demands on the rapid migration capabilities of strategies, while traditional methods require retraining models for each topology, which takes several weeks. The above technical bottlenecks have seriously restricted the practical application of adaptive traffic diversion systems, and there is an urgent need for an innovative solution that can uniformly characterize multi-protocol features, dynamically coordinate multi-objective conflicts, and adapt to edge deployment. Summary of the invention

[0003] In view of the above-mentioned problems, the present invention is proposed.

[0004] Therefore, the technical problem solved by the present invention is: the technical problems of low network diversion efficiency, large fluctuations in service quality and lagging security protection caused by multi-protocol differences, multi-target strategy conflicts and insufficient adaptability to dynamic environments in the existing heterogeneous network environment.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: an adaptive network offloading method based on deep reinforcement learning, comprising:

[0006] Obtain traffic data from heterogeneous networks, build a protocol fingerprint library, and convert data from different protocols into a unified network status feature vector;

[0007] Construct a multi-agent reinforcement learning framework, including defining the state space, action space, and reward function, and optimize the strategies of each agent based on differential games;

[0008] A meta-learning model is used to adjust the weight of each agent's reward function, and deep reinforcement learning is used to train a multi-agent reinforcement learning framework to learn the optimal diversion strategy under different heterogeneous network conditions;

[0009] The trained strategy is deployed to a heterogeneous network environment, and the optimal network path is dynamically selected based on the current network status.

[0010] As a preferred solution of the adaptive network traffic diversion method based on deep reinforcement learning described in the present invention, wherein: the traffic data of the heterogeneous network includes: protocol type identification, message header information, timestamp, timing mode, signal strength and quality index, bandwidth utilization, packet loss rate, data packet length, traffic size, traffic mode, abnormal traffic mode and authentication identification;

[0011] Establishing a protocol fingerprint library includes preprocessing the acquired traffic data, extracting key features from the preprocessed traffic data, establishing a binary coding template, encoding the key features, obtaining the protocol fingerprint, and storing it in the protocol fingerprint library in an index structure;

[0012] The key features include data features and multi-agent network status features; the multi-agent network status features include flow control agent status features, QoS management agent status features and security agent status features;

[0013] Converting data of different protocols into a unified network status feature vector includes: converting the protocol fingerprint into a numerical vector, normalizing continuous numerical values ​​in the numerical vector, and concatenating all normalized continuous numerical values ​​into the network status feature vector.

[0014] As a preferred solution of the adaptive network traffic diversion method based on deep reinforcement learning described in the present invention, wherein: the multi-agent reinforcement learning framework includes a traffic control agent, a QoS management agent and a security agent;

[0015] Defining the state space is to define the observable variables of each agent based on the state characteristics of the multi-agent network to obtain the state space;

[0016] The state space includes: the state space of the flow control agent, the state space of the QoS management agent and the state space of the security agent; the defined action space includes: the action space of the flow control agent, the action space of the QoS management agent and the action space of the security agent;

[0017] The reward functions defined include: the reward function of the traffic control agent, the reward function of the QoS management agent, and the reward function of the security agent;

[0018] The reward function of the traffic control agent is to maximize network throughput and minimize load imbalance; the reward function of the QoS management agent is to ensure that the delay, packet loss rate and jitter of the service flow are maintained within the optimal range; the reward function of the security agent is to minimize the proportion of abnormal traffic while reducing false alarms.

[0019] As a preferred solution of the adaptive network traffic diversion method based on deep reinforcement learning described in the present invention, wherein: optimizing the decision strategy of each intelligent agent based on differential game includes constructing a state equation of differential game based on the state space of multi-agent reinforcement learning framework to describe the dynamic process of network state changing over time;

[0020] Solve the Nash equilibrium of the differential game to ensure that the overall strategy of the multi-agent reinforcement learning framework remains optimal when the network state is constantly changing.

[0021] As a preferred solution of the adaptive network diversion method based on deep reinforcement learning described in the present invention, the meta-learning model is used to adjust the weight of each agent's reward function, including analyzing the network state of the past N time steps based on the attention mechanism, calculating the weight of each agent's network state characteristics, and obtaining each agent's state characteristic weight vector; using the weighted summation method, and combining the state characteristic weight vector of each agent, dynamically adjusting the weight of each agent's reward function.

[0022] As a preferred solution of the adaptive network diversion method based on deep reinforcement learning described in the present invention, wherein: using deep reinforcement learning to train a multi-agent reinforcement learning framework includes constructing a deep reinforcement learning strategy network, setting an exploration and utilization mechanism, initializing an experience pool, and training a multi-agent reinforcement learning framework.

[0023] As a preferred solution of the adaptive network diversion method based on deep reinforcement learning described in the present invention, wherein: constructing a deep reinforcement learning strategy network includes: an input layer, a hidden layer and an output layer;

[0024] The input layer includes inputting the state space into the input layer to obtain an input vector;

[0025] The hidden layer includes a fully connected layer, a feature mapping layer and a strategy generation layer;

[0026] The feature mapping layer sets the attention score, calculates the weight of the internal features of the state space of each agent, and performs weighted summation to obtain the final state feature vector of each agent; establishes a differential game equation to describe the impact of different agents' strategies on other agents; solves the Nash equilibrium to ensure that the multi-agent strategy remains optimal in a dynamic environment, and obtains the game impact between agents;

[0027] Combine the final state feature vector of each agent and the game influence between agents to generate the final strategy input vector of each agent;

[0028] The strategy generation layer includes calculating the Q value of different actions of each agent using the final strategy input vector and action space of each agent, selecting the action with the largest Q value as the optimal strategy of each agent; and updating the Q value using the Q-Learning method;

[0029] The output layer includes outputting the optimal strategy of each agent.

[0030] An adaptive network traffic diversion system based on deep reinforcement learning, wherein:

[0031] The protocol fingerprint library module obtains the traffic data of heterogeneous networks, establishes a protocol fingerprint library, and converts the data of different protocols into a unified network status feature vector;

[0032] Multi-agent reinforcement learning module, which builds a multi-agent reinforcement learning framework, including defining the state space, action space, and reward function, and optimizing the strategies of each agent based on differential games;

[0033] The strategy module uses a meta-learning model to adjust the weight of each agent's reward function, and uses deep reinforcement learning to train a multi-agent reinforcement learning framework to learn the optimal diversion strategy under different heterogeneous network conditions;

[0034] The deployment module deploys the trained strategy to a heterogeneous network environment and dynamically selects the optimal network path based on the current network status.

[0035] A computer device comprises: a memory and a processor; the memory stores a computer program, wherein the processor implements the steps of any one of the methods of the present invention when executing the computer program.

[0036] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the methods of the present invention.

[0037] Beneficial effects of the present invention: The adaptive network diversion method based on deep reinforcement learning provided by the present invention realizes efficient parsing and dynamic adaptation across network protocols by constructing a protocol fingerprint library and a unified feature coding mechanism, and significantly improves the heterogeneous data processing capability. Based on the multi-agent differential game framework, the collaborative decision-making of flow control, service quality optimization and security protection is coordinated to effectively resolve the traditional single-agent strategy conflicts and simultaneously optimize network throughput and transmission stability. The dynamic weight adjustment mechanism driven by meta-learning is introduced to enhance the system's adaptive ability to complex network environments and greatly reduce the risk of service quality violations. Combining deep reinforcement learning with lightweight deployment technology, the model resource consumption is significantly compressed, and efficient reasoning and real-time response of edge devices are achieved, breaking through the response delay bottleneck of traditional solutions in burst traffic processing and security defense. Finally, a set of intelligent diversion systems that support multi-protocol compatibility, multi-objective dynamic balance, and low resource occupancy are formed, providing flexible and reliable adaptive scheduling capabilities for heterogeneous network environments, and meeting the stringent requirements of new network applications for real-time and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0039] Figure 1 An overall flow chart of an adaptive network traffic diversion method based on deep reinforcement learning provided for the first embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0041] Example 1, reference Figure 1 , is an embodiment of the present invention, and provides an adaptive network offloading method based on deep reinforcement learning, comprising:

[0042] S1: Obtain traffic data of heterogeneous networks, establish a protocol fingerprint library, and convert data from different protocols into a unified network status feature vector.

[0043] The traffic data of the heterogeneous network includes: protocol type identification, message header information, timestamp, timing mode, signal strength and quality indicators, bandwidth utilization, packet loss rate, data packet length, traffic size, traffic mode, abnormal traffic mode and authentication identification.

[0044] Establishing a protocol fingerprint library includes preprocessing the acquired traffic data, extracting key features from the preprocessed traffic data, establishing a binary coding template, encoding the key features, obtaining the protocol fingerprint, and storing it in the protocol fingerprint library in an index structure.

[0045] The key features include data features and multi-agent network status features. The multi-agent network status features include flow control agent status features, QoS management agent status features and security agent status features. Data features are used for flow statistics, and multi-agent status features are used for reinforcement learning agent decision making. Multi-agent status features include status features of flow control agents, status features of QoS management agents and status features of security agents.

[0046] Converting data of different protocols into a unified network status feature vector includes: converting the protocol fingerprint into a numerical vector, normalizing continuous numerical values ​​in the numerical vector, and concatenating all normalized continuous numerical values ​​into the network status feature vector.

[0047] Furthermore, by establishing a protocol fingerprint library and converting data from different protocols into a unified network status feature vector, traffic data in heterogeneous network environments can be standardized. By extracting key features such as protocol type, message header information, bandwidth utilization, packet loss rate, and using binary encoding to store protocol fingerprints, data retrieval efficiency and adaptability can be improved, providing a consistent input data format for subsequent multi-agent reinforcement learning. At the same time, the normalization method is used to standardize values ​​in different ranges, making the data distribution more balanced, which is conducive to the stable training of reinforcement learning algorithms and improving the generalization ability of the model for complex network environments.

[0048] Furthermore, data features are used for traffic statistics, and multi-agent state features are used for agent decision-making, so that traffic control, QoS management, and security agents can optimize strategies based on a unified state feature vector. By building a unified network state feature vector, the agent can effectively perceive the traffic characteristics under different network protocols, realize cross-protocol intelligent diversion, QoS optimization, and security protection, and improve the overall network's adaptability and decision-making efficiency.

[0049] S2: Construct a multi-agent reinforcement learning framework, including defining the state space, defining the action space, and defining the reward function, and optimizing the strategy of each agent based on differential games.

[0050] The multi-agent reinforcement learning framework includes traffic control agent, QoS management agent and security agent.

[0051] The state space is defined by defining observable variables of each agent based on the state characteristics of the multi-agent network to obtain the state space.

[0052] The state space includes: the state space of the flow control agent, the state space of the QoS management agent and the state space of the security agent. The defined action space includes: the action space of the flow control agent, the action space of the QoS management agent and the action space of the security agent.

[0053] State space of the traffic control agent: network load, bandwidth utilization, packet loss rate.

[0054] The state space of the QoS management agent: service type, QoS priority, and current network delay.

[0055] The state space of the security agent: abnormal traffic detection results, attack types, and historical security records.

[0056] Action space of the traffic control agent: select the traffic distribution method (5G, WiFi6, satellite network).

[0057] Action space of QoS management agent: adjust bandwidth priority of video, voice, data and other service flows.

[0058] Action space of security agents: enable or adjust attack defense strategies (such as DDoS filtering, access control).

[0059] The reward functions defined include: reward function of traffic control agent, reward function of QoS management agent and reward function of security agent.

[0060] The reward function of the traffic control agent is to maximize the network throughput and minimize the load imbalance, which is expressed as follows:

[0061] R T =α 1 B-α 2 L

[0062] Among them, R T represents the reward of the traffic control agent, B represents the network throughput, L represents the degree of network load imbalance, and α 1 ,α 2 Indicates the importance of determining each indicator.

[0063] The reward function of the QoS management agent is to ensure that the delay, packet loss rate and jitter of the service flow are maintained in the optimal range, which is expressed as follows:

[0064] R Q =-β 1 D-β 2 P-β 3 J

[0065] Among them, R Q : represents the reward of the QoS management agent, D represents the delay, P represents the packet loss rate, J represents the jitter, β 1 ,β 2 ,β 3 Represents the weight coefficient.

[0066] The reward function of the security agent is to minimize the proportion of abnormal traffic while reducing false positives, and the formula is expressed as:

[0067] R S =-γ 1 A+γ 2 (1-F)

[0068] Among them, R S represents the reward of the security agent, A represents the abnormal traffic ratio, F represents the false alarm rate, and the value range is 0 to 1, γ 1 ,γ 2 Represents the weight coefficient.

[0069] Optimizing the decision-making strategies of each intelligent agent based on differential games includes constructing the state equation of the differential game based on the state space of the multi-agent reinforcement learning framework to describe the dynamic process of network state changing over time.

[0070] Solve the Nash equilibrium of the differential game to ensure that the overall strategy of the multi-agent reinforcement learning framework remains optimal when the network state is constantly changing. The formula is expressed as:

[0071]

[0072] Among them, R i represents the reward function of the ith agent, indicating the overall benefit in a given state, The optimal strategy of the i-th agent, The optimal strategy for all agents except the i-th agent is: represents the action set of the ith agent, Represents all possible actions for the i-th agent.

[0073] Furthermore, by constructing a multi-agent reinforcement learning framework, the flow control agent, QoS management agent, and security agent can adaptively optimize the diversion strategy in a dynamic network environment. By reasonably defining the state space and action space, the agent can accurately perceive the network status and make autonomous decisions based on reinforcement learning, making the allocation of network resources more reasonable. Among them, the flow control agent can dynamically adjust the distribution of data flows, improve network throughput and balance the load. The QoS management agent optimizes traffic scheduling according to the service type and network conditions to ensure that the delay, packet loss rate, and jitter are in the optimal range. The security agent effectively reduces the impact of malicious attacks through anomaly detection and defense strategy adjustment, while reducing the false alarm rate and improving network security.

[0074] Furthermore, differential games are used to optimize the decision-making strategies of each intelligent agent. By constructing differential game state equations to describe changes in network states and solving Nash equilibrium, it is ensured that each intelligent agent reaches the global optimal decision result under mutual influence. Compared with the traditional independent training reinforcement learning method, this method can dynamically adapt to the interaction relationship between different intelligent agents in a complex heterogeneous network environment, improve the overall network performance, service quality and security protection capabilities, and realize intelligent and autonomously optimized network diversion strategies.

[0075] S3: A meta-learning model is used to adjust the weight of each agent’s reward function, and deep reinforcement learning is used to train a multi-agent reinforcement learning framework to learn the optimal diversion strategy under different heterogeneous network conditions.

[0076] The meta-learning model is used to adjust the weight of each agent's reward function, including analyzing the network state of the past N time steps based on the attention mechanism, calculating the weight of each agent's network state feature, and obtaining each agent's state feature weight vector. The weighted summation method is used, and the weight of each agent's reward function is dynamically adjusted in combination with the state feature weight vector of each agent.

[0077] Using deep reinforcement learning to train a multi-agent reinforcement learning framework includes building a deep reinforcement learning strategy network, setting up exploration and utilization mechanisms, initializing the experience pool, and training the multi-agent reinforcement learning framework.

[0078] Constructing a deep reinforcement learning strategy network includes: input layer, hidden layer and output layer.

[0079] The input layer includes inputting the state space into the input layer to obtain an input vector.

[0080] The hidden layer includes a fully connected layer, a feature mapping layer and a strategy generation layer.

[0081] The feature mapping layer sets the attention score, calculates the weight of the internal features of the state space of each agent, and performs weighted summation to obtain the final state feature vector of each agent. Establish differential game equations to describe the impact of different agent strategies on other agents. Solve the Nash equilibrium to ensure that the multi-agent strategy remains optimal in a dynamic environment and obtain the game impact between agents.

[0082] The final state feature vector of each agent and the game influence between agents are combined to generate the final strategy input vector of each agent.

[0083] The strategy generation layer includes calculating the Q value of different actions of each agent using the final strategy input vector and action space of each agent, selecting the action with the largest Q value as the optimal strategy of each agent, and updating the Q value using the Q-Learning method.

[0084] The output layer includes outputting the optimal strategy of each agent.

[0085] Furthermore, by dynamically adjusting the reward function weights of each agent through the meta-learning model, the reinforcement learning system can adaptively optimize the learning direction of each agent according to the historical network state and environmental changes. By introducing the attention mechanism, the agent can focus on the most important historical state features and improve the accuracy and generalization ability of policy decisions. In addition, combined with the weighted summation method, it can ensure that each agent can flexibly adjust the optimization target under different network conditions, avoid fixed strategies, and improve adaptability to heterogeneous network environments.

[0086] Furthermore, by using a deep reinforcement learning policy network to train a multi-agent reinforcement learning framework, and using a fully connected layer, feature mapping layer, and policy generation layer to hierarchically process state information, it is possible to efficiently extract key features in complex network environments. Combined with differential game optimization, it ensures that the strategies of each agent reach a global equilibrium and avoids system performance degradation caused by over-optimization of a single agent. By updating the Q value through the Q-Learning algorithm, the decision-making quality of the agent can be continuously improved, the dynamic adjustment of the optimal diversion strategy can be achieved, the network throughput can be improved, the QoS experience can be optimized, and the stability and reliability of the system in complex environments can be enhanced.

[0087] S4: Deploy the trained strategy to a heterogeneous network environment and dynamically select the optimal network path based on the current network status.

[0088] By observing the reward curve, we can ensure that the agent strategy converges and no longer fluctuates drastically.

[0089] The trained strategy is deployed to the network controller for actual traffic scheduling.

[0090] Real-time monitoring of the network environment and dynamic adjustment of strategies: When burst traffic increases, the traffic control agent re-optimizes the diversion. When business needs change, the QoS management agent adjusts the priority. When abnormal traffic is found, the security agent takes the initiative to defend.

[0091] Furthermore, by deploying the trained agent strategy to a heterogeneous network environment, it is possible to perceive changes in network status in real time and dynamically adjust network traffic diversion strategies. When burst traffic increases, the traffic control agent can automatically optimize traffic scheduling to avoid network congestion. When business needs change, the QoS management agent can adjust traffic priorities in a timely manner to ensure the quality of service for key business flows. This adaptive optimization mechanism can improve bandwidth utilization, reduce latency and packet loss, and thus enhance overall network performance.

[0092] Furthermore, through the continuous monitoring and dynamic decision-making of the intelligent agent, it can respond quickly when abnormal traffic or malicious attacks occur. The security intelligent agent can actively detect and intercept abnormal traffic, effectively reducing the impact of attacks on the network. This dynamic adjustment mechanism not only improves network security defense capabilities, but also reduces false alarm rates and ensures that normal traffic is not disturbed. In addition, with the continuous learning and optimization of the intelligent agent strategy, the system can maintain a stable optimal strategy in long-term operation, reduce frequent manual intervention, and reduce operation and maintenance costs.

[0093] Embodiment 2 is an embodiment of the present invention, which provides an adaptive network traffic distribution system based on deep reinforcement learning, including:

[0094] The protocol fingerprint library module obtains the traffic data of heterogeneous networks, establishes a protocol fingerprint library, and converts the data of different protocols into a unified network status feature vector.

[0095] The multi-agent reinforcement learning module builds a multi-agent reinforcement learning framework, including defining the state space, defining the action space, and defining the reward function, and optimizing the strategy of each agent based on differential games.

[0096] The strategy module adopts a meta-learning model to adjust the weight of each agent's reward function, and uses deep reinforcement learning to train a multi-agent reinforcement learning framework to learn the optimal diversion strategy under different heterogeneous network conditions.

[0097] The deployment module deploys the trained strategy to a heterogeneous network environment and dynamically selects the optimal network path based on the current network status.

[0098] Embodiment 3, an embodiment of the present invention, is different from the first two embodiments in that:

[0099] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0100] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0101] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory,

[0102] (EPROM or flash memory), fiber optic devices, and portable CD-ROM

[0103] (CDROM). In addition, the computer readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting or processing in another suitable manner as necessary, and then stored in a computer memory.

[0104] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0105] Example 4 is an embodiment of the present invention, which provides an adaptive network diversion method and system based on deep reinforcement learning. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.

[0106] In order to verify the applicability of the adaptive network offloading method based on deep reinforcement learning in a complex heterogeneous network environment, the experiment was conducted in a real network test environment including 5G base stations, WiFi6 access points and low-orbit satellite links. The test network includes a 5G base station that supports a maximum throughput of 1Gbps, a WiFi6 access point that supports a throughput of 600Mbps, and a low-orbit satellite link that supports a throughput of 200Mbps. Business traffic types include high-definition video, VoIP calls, IoT data transmission, and web browsing, and simulated DDoS attack traffic is added to test the security and adaptability of the network offloading method.

[0107] The experiment first collects data to obtain traffic data of heterogeneous networks, mainly including key features such as protocol type, message header information, timestamp, traffic pattern, bandwidth utilization, and packet loss rate. The collected data is preprocessed to extract key features and establish a protocol fingerprint library. All protocol data are converted into a unified network state feature vector to be input into the reinforcement learning framework for intelligent decision-making.

[0108] The training of the multi-agent reinforcement learning framework is divided into three stages. The first stage is the initialization of the agent, which constructs the flow control agent, QoS management agent and security agent respectively. The flow control agent is responsible for optimizing the network traffic distribution, the QoS management agent ensures that the latency, packet loss rate and jitter of different business traffic are maintained in the optimal range, and the security agent detects and defends against abnormal traffic. In the second stage, reinforcement learning training is carried out. The agents continuously optimize the decision-making strategy through interaction, and optimize the strategy of each agent based on differential game, so that the overall network diversion strategy is optimal. During the training process, the agent learns under different network conditions, and uses meta-learning methods to dynamically adjust the reward weights so that it can adapt to the complex and changing network environment. In the third stage, after the training is completed, the agent strategy is deployed to the heterogeneous network environment, and real-time strategy adjustment is performed to dynamically select the optimal network path.

[0109] The experiment sets up different network scenarios, including high-load environments, sudden traffic changes, malicious attack traffic penetration, etc., to test whether the method can balance traffic, optimize QoS, and improve network security in different scenarios. During the test, the agent monitors the network status in real time, the traffic control agent optimizes network traffic when sudden traffic increases, the QoS management agent adjusts traffic priority according to business needs, and the security agent takes the initiative to take defense strategies after detecting abnormal traffic. The entire experimental process lasts for several hours to ensure stability and optimization effects in multiple scenarios.

[0110] In the 5G network environment, the throughput before optimization was 818.14Mbps, and it increased to 948.76Mbps after optimization, with a significant improvement in throughput. In the WiFi6 network environment, the throughput before optimization was 505.58Mbps, and it reached 580.11Mbps after optimization, with a relatively stable increase in throughput after optimization. In the satellite network environment, the throughput before optimization was 163.41Mbps, and it reached 197.11Mbps after optimization, which improved the transmission efficiency compared to before optimization.

[0111] In terms of QoS indicators, the average latency before optimization was 51.81ms, which was reduced to 38.27ms after optimization. The latency was significantly reduced after optimization, which improved the network response speed. In terms of packet loss rate, the packet loss rate before optimization was 3.86%, which was reduced to 1.85% after optimization, which reduced the lost data packets and improved the stability of transmission. The jitter value before optimization was 10.64ms, which was reduced to 8.24ms after optimization. The data fluctuation tended to be more stable, which improved the experience of real-time services.

[0112] In terms of abnormal traffic detection and protection, the abnormal traffic accounted for 24.88% before optimization and dropped to 6.78% after optimization, indicating that the optimization method can effectively reduce the impact of abnormal traffic. At the same time, the false alarm rate was 7.87% before optimization and dropped to 3.30% after optimization, indicating that this method can effectively detect and defend abnormal traffic while reducing misjudgment, ensuring the stability and security of the system. Experimental data analysis shows that the adaptive network diversion method based on deep reinforcement learning has significant advantages over traditional methods in multiple key performance indicators. First, in terms of throughput optimization, this method can adjust the traffic allocation strategy according to the real-time network status, thereby improving link utilization. The throughput improvement effect in the 5G network environment is particularly significant, and the throughput optimization of WiFi6 and satellite networks also shows a stable improvement trend. Compared with the traditional method that adopts a fixed traffic allocation strategy, this method can dynamically adjust the network resource allocation so that the throughput under different network types is kept in the optimal range.

[0113] Secondly, in terms of QoS optimization, this method can dynamically adjust traffic priority according to changes in business needs to ensure the service quality of high-priority services. In high-definition video traffic scenarios, this method can effectively reduce latency, greatly improving the smoothness of video playback, while traditional methods are prone to video freezes when the network load is high due to the lack of dynamic optimization capabilities. In voice call scenarios, this method can reduce packet loss rate, ensure stable transmission of voice data, and improve call quality, while traditional methods are prone to call quality degradation when burst traffic increases.

[0114] In terms of abnormal traffic suppression, this method can quickly detect and isolate abnormal traffic through security agents trained by reinforcement learning. After a DDoS attack occurs, this method identifies attack traffic in a very short time and reduces its impact on normal business through an intelligent traffic adjustment mechanism, so that the stability of normal business traffic can be guaranteed. Compared with traditional methods that rely on fixed thresholds for anomaly detection, this method can improve detection accuracy and reduce false alarms by dynamically optimizing detection strategies. Experimental data show that this method significantly improves the detection rate of abnormal traffic, while the false alarm rate is greatly reduced compared with traditional methods.

[0115] In addition, in terms of computational efficiency and adaptability, the agent training time of this method is much shorter than that of traditional methods. At the same time, after actual deployment, it can quickly respond to changes in network status and achieve real-time diversion optimization. Traditional methods require manual intervention to adjust strategies when facing dynamically changing network environments, while this method can automatically optimize strategies through reinforcement learning, reduce manual intervention, and improve the level of intelligent network management.

[0116] In summary, the adaptive network traffic diversion method based on deep reinforcement learning is superior to traditional methods in terms of throughput optimization, QoS guarantee, network security protection, and computational efficiency. It solves the problem that existing traffic diversion strategies are difficult to adapt to complex network environments and has significant technical innovation and practical application value. This method can not only improve the utilization of network resources, but also perform intelligent optimization and adjustment according to changes in network status, and has broad application potential in heterogeneous network environments.

[0117] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An adaptive network offloading method based on deep reinforcement learning, characterized in that: include: Obtain traffic data from heterogeneous networks, build a protocol fingerprint library, and convert data from different protocols into a unified network status feature vector; Construct a multi-agent reinforcement learning framework, including defining the state space, action space, and reward function, and optimize the strategies of each agent based on differential games; A meta-learning model is used to adjust the weight of each agent's reward function, and deep reinforcement learning is used to train a multi-agent reinforcement learning framework to learn the optimal diversion strategy under different heterogeneous network conditions; The trained strategy is deployed to a heterogeneous network environment, and the optimal network path is dynamically selected based on the current network status.

2. The method for adaptive network offloading based on deep reinforcement learning according to claim 1, characterized in that: The traffic data of the heterogeneous network includes: protocol type identification, message header information, timestamp, timing mode, signal strength and quality indicators, bandwidth utilization, packet loss rate, data packet length, traffic size, traffic mode, abnormal traffic mode and authentication identification; Establishing a protocol fingerprint library includes preprocessing the acquired traffic data, extracting key features from the preprocessed traffic data, establishing a binary coding template, encoding the key features, obtaining the protocol fingerprint, and storing it in the protocol fingerprint library in an index structure; The key features include data features and multi-agent network status features; the multi-agent network status features include flow control agent status features, QoS management agent status features and security agent status features; Converting data of different protocols into a unified network status feature vector includes: converting the protocol fingerprint into a numerical vector, normalizing continuous numerical values ​​in the numerical vector, and concatenating all normalized continuous numerical values ​​into the network status feature vector.

3. The adaptive network offloading method based on deep reinforcement learning according to claim 2, characterized in that: The multi-agent reinforcement learning framework includes traffic control agent, QoS management agent and security agent; Defining the state space is to define the observable variables of each agent based on the state characteristics of the multi-agent network to obtain the state space; The state space includes: the state space of the flow control agent, the state space of the QoS management agent and the state space of the security agent; the defined action space includes: the action space of the flow control agent, the action space of the QoS management agent and the action space of the security agent; The reward functions defined include: the reward function of the traffic control agent, the reward function of the QoS management agent, and the reward function of the security agent; The reward function of the traffic control agent is to maximize network throughput and minimize load imbalance; the reward function of the QoS management agent is to ensure that the delay, packet loss rate and jitter of the service flow are maintained within the optimal range; the reward function of the security agent is to minimize the proportion of abnormal traffic while reducing false alarms.

4. The method for adaptive network offloading based on deep reinforcement learning according to claim 3, characterized in that: The decision-making strategy of each agent is optimized based on differential games, including constructing the state equation of differential games based on the state space of multi-agent reinforcement learning framework to describe the dynamic process of network state changing over time; Solve the Nash equilibrium of the differential game to ensure that the overall strategy of the multi-agent reinforcement learning framework remains optimal when the network state is constantly changing.

5. The method for adaptive network offloading based on deep reinforcement learning according to claim 4, characterized in that: The meta-learning model is used to adjust the weight of each agent's reward function, including analyzing the network status of the past N time steps based on the attention mechanism, calculating the weight of each agent's network state characteristics, and obtaining each agent's state characteristic weight vector; using the weighted summation method, combined with the state characteristic weight vector of each agent, dynamically adjusting the weight of each agent's reward function.

6. The method for adaptive network offloading based on deep reinforcement learning according to claim 5, characterized in that: Using deep reinforcement learning to train a multi-agent reinforcement learning framework includes building a deep reinforcement learning strategy network, setting up exploration and utilization mechanisms, initializing the experience pool, and training the multi-agent reinforcement learning framework.

7. The method for adaptive network offloading based on deep reinforcement learning according to claim 6, characterized in that: Constructing a deep reinforcement learning strategy network includes: input layer, hidden layer and output layer; The input layer includes inputting the state space into the input layer to obtain an input vector; The hidden layer includes a fully connected layer, a feature mapping layer and a strategy generation layer; The feature mapping layer sets the attention score, calculates the weight of the internal features of the state space of each agent, and performs weighted summation to obtain the final state feature vector of each agent; establishes a differential game equation to describe the impact of different agents' strategies on other agents; solves the Nash equilibrium to ensure that the multi-agent strategy remains optimal in a dynamic environment, and obtains the game impact between agents; Combine the final state feature vector of each agent and the game influence between agents to generate the final strategy input vector of each agent; The strategy generation layer includes calculating the Q value of different actions of each agent using the final strategy input vector and action space of each agent, selecting the action with the largest Q value as the optimal strategy of each agent; and updating the Q value using the Q-Learning method; The output layer includes outputting the optimal strategy of each agent.

8. An adaptive network traffic distribution system based on deep reinforcement learning using the method according to any one of claims 1 to 7, characterized in that: The protocol fingerprint library module obtains the traffic data of heterogeneous networks, establishes a protocol fingerprint library, and converts the data of different protocols into a unified network status feature vector; Multi-agent reinforcement learning module, which builds a multi-agent reinforcement learning framework, including defining the state space, action space, and reward function, and optimizing the strategies of each agent based on differential games; The strategy module uses a meta-learning model to adjust the weight of each agent's reward function, and uses deep reinforcement learning to train a multi-agent reinforcement learning framework to learn the optimal diversion strategy under different heterogeneous network conditions; The deployment module deploys the trained strategy to a heterogeneous network environment and dynamically selects the optimal network path based on the current network status.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the adaptive network diversion method based on deep reinforcement learning described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the adaptive network diversion method based on deep reinforcement learning described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Intelligent network path optimization method and system based on deep reinforcement learning

    CN116527567A

  • Intention-driven intelligent routing method for multi-domain heterogeneous data link network

    CN118612138A

  • Intelligent path optimization method and system based on link state perception enhancement

    CN119011463A

  • Dynamic multi-link intelligent management and scheduling system

    CN119420691A

  • Autonomous traffic (self-driving) network with traffic classes and passive and active learning

    US20230145097A1

Cited By

  • Intelligent warehouse management system and method based on Internet of Things

    CN120338670A

  • Dynamic routing collaborative optimization method for heterogeneous network based on reinforcement learning

    CN120729774A

  • Invisible structure parameter optimization method and system based on reinforcement learning and layering strategy

    CN121302941A

  • Covert communication flow identification method and system based on reinforcement learning dynamic feature screening

    CN121598147A

  • Multi-agent dynamic defense game method and system based on federal reinforcement learning

    CN121864500A