Intelligent multicast routing optimization method based on deep reinforcement learning and game theory

Through the intelligent multicast routing optimization method based on deep reinforcement learning and game theory, the problems of redundant transmission and resource waste in traditional unicast routing strategies are solved, and efficient and low-cost data transmission and network resource management are achieved.

CN119945960AInactive Publication Date: 2025-05-06HUZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093856.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional unicast routing strategy has redundant transmission problems when transmitting large amounts of data in one-to-many scenarios, resulting in the risk of waste of network resources and congestion. The existing intelligent optimization algorithms lack flexibility and adaptability in multi-dimensional service quality, and have high energy consumption and low path construction efficiency in large-scale networks.

Method used

The intelligent multicast routing optimization method based on deep reinforcement learning and game theory is adopted. By abstracting edge devices into an undirected graph network topology, link costs and unicast paths are set, and multicast paths are built using potential game relationships to optimize resource allocation and data transmission.

Benefits of technology

It significantly reduces transmission costs, improves transmission efficiency, enhances the dynamic adaptability of the network, and has higher resource utilization and path construction efficiency than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945960A_ABST
    Figure CN119945960A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of routing optimization, in particular to an intelligent multicast routing optimization method based on deep reinforcement learning and the game theory, and the method comprises the steps: abstracting edge equipment in a real scene into an undirected graph network topology structure, and reading a network state information data set obtained through a software defined network technology; setting link cost for each edge according to topological information and link information in the data set; according to a potential game relationship existing between unicast paths, establishing an incentive mechanism to prepare for subsequent reinforcement learning training; constructing a multicast path in a manner of gradually selecting a unicast path to each destination node according to a multicast demand; a total reward value is given after unicast paths of destination nodes are sequentially added into multicast paths, and a Nash equilibrium state is achieved through training. According to the invention, by optimizing the selection of the multicast routing path and utilizing the game relationship between unicast paths to share resources, the transmission cost is obviously reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of routing optimization, and in particular to an intelligent multicast routing optimization method based on deep reinforcement learning and game theory. Background Art

[0002] With the widespread application of the Internet of Things (IoT) and one-to-many information transmission applications such as video conferencing and cloud storage, traditional unicast routing strategies have the problem of redundant transmission when transmitting large amounts of data in one-to-many scenarios, resulting in network resource waste and congestion risks.

[0003] Multicast technology sends data to multiple destination nodes at one time through the source node, significantly reducing redundant transmission, network traffic and energy consumption, and has obvious advantages in scenarios such as the Internet of Things. In order to meet the quality of service requirements, multicast routing is usually implemented by constructing a Steiner tree. However, the construction of a Steiner tree is an NP-complete problem, and commonly used methods include heuristic algorithms and intelligent optimization algorithms. Although these algorithms can approximate the solution, they lack flexibility and adaptability in multi-dimensional service quality, and still face problems such as high energy consumption and low path construction efficiency in large-scale networks. In order to solve these problems, researchers have proposed intelligent optimization algorithms, such as ant colony algorithms and genetic algorithms, which can find the global optimal solution in some cases, but there are still limitations in convergence speed and computational complexity.

[0004] In addition, the Nash equilibrium theory in game theory is also applied to the construction of multicast trees. By sharing resources between multiple unicast paths, the total cost of the multicast tree can be minimized. In practical applications, how to effectively select shared paths to meet traffic with different service quality requirements and improve resource utilization is still a problem that needs to be solved. Summary of the invention

[0005] The purpose of the present invention is to provide an intelligent multicast routing optimization method based on deep reinforcement learning and game theory, which can achieve optimal resource allocation and efficient data transmission in a dynamic network environment by optimizing the selection of multicast routing paths.

[0006] To achieve the above object, the present invention provides an intelligent multicast routing optimization method based on deep reinforcement learning and game theory, comprising the following steps:

[0007] Step 1: Abstract the edge devices in the real scene into an undirected graph network topology structure, read the network status information data set obtained by using software-defined network technology; obtain multicast requirements, source node s, and destination node;

[0008] Step 2: Set the link cost for each edge according to the network status information of the topology, and select k relatively short unicast paths to each destination node based on the number of hops;

[0009] Step 3: Based on the potential game relationship between unicast paths, establish a model that incentivizes unicast paths to select more shared edges when building multicast paths;

[0010] Step 4: Set the state space, action space and reward function in deep reinforcement learning, and construct the multicast path by gradually selecting the unicast path to each destination node;

[0011] Step 5: The unicast path of the destination node is added to the multicast path in turn and the total reward value is given. The Nash equilibrium state is reached through training, that is, the reward value cannot be increased by changing any thin path.

[0012] Optionally, in step 1, in the undirected graph network G, G = (V, E), V is a node in G, E is a set of links in G, and e ij ∈E represents the link between node i and node j. Any node can be used as a source node or a destination node. The source node is represented by s, and D = {d1, d2...d n} represents the target node set; the minimum Steiner tree from the source node to multiple destination nodes is defined as the minimum cost tree from s to D in G, which is obtained through reinforcement learning training. The minimum cost tree is the multicast tree with the highest reward value obtained using the reinforcement learning algorithm.

[0013] Optionally, the network status information data set includes pickle files of network status information of different topologies, which are obtained from a traffic generation tool to simulate traffic conditions for 24 hours a day, and the network status information includes the remaining bandwidth of the link and the link, the average transmission delay and the link packet loss rate;

[0014] During the execution of step 2, the remaining bandwidth, average transmission delay, and link packet loss rate in the link information are normalized and weighted to obtain the edge cost between any two nodes i and j. The definitions are as follows:

[0015]

[0016] τ1+τ2+τ3=1

[0017] in, The total cost, α, β, γ are the normalized results of the remaining bandwidth, average propagation delay and average packet loss rate between two nodes, respectively, and τ is an adjustable weight coefficient.

[0018] Optionally, in step 3, a multicast tree with minimum cost is constructed by elaborating the potential game relationship between paths. In the multicast tree constructed according to the link cost information, any unicast path cannot be changed to make the constructed multicast tree have a lower cost, so that the multicast tree maintains a Nash equilibrium state.

[0019] Optionally, the execution process of step 4 is specifically a process of constructing a multicast tree using a reinforcement learning method, including the following steps:

[0020] The execution process of step 4 is specifically the process of constructing a multicast tree using a reinforcement learning method, and includes the following steps:

[0021] Step 4.1: Use the collected network link information to train the agent offline, learn how to build the optimal multicast path, update the network parameters, and store the trained routing strategy in the experience replay buffer;

[0022] Step 4.2: The dual-depth dual-Q network reinforcement learning algorithm uses two networks to separate action selection and action evaluation, where the Q network is a deep Q network, which is a neural network used to approximate the Q-value function in reinforcement learning. The Q-value function represents the cumulative expected return / reward value that the agent can obtain after performing an action a in a certain state s.

[0023] Step 4.3: Update the parameters of the Q network using the loss function based on temporal differences;

[0024] Step 4.4: The design of the state space uses five different matrices, including the topological connection matrix, the selected path matrix, and three link information matrices, which are expressed as follows:

[0025] s t =[M t ,M p ,M b ,M d ,M l ]

[0026] Among them, the topological connection matrix M t The connection between all nodes in the network is recorded and displayed in the form of an adjacency matrix. Each element indicates whether there is a direct physical connection between two nodes. The selected path matrix M p It is used to track the decision-making process in the multicast tree construction process, mark the selected path, avoid repeated selection, and assist in subsequent decision-making; the three link information matrices M b ,M d ,M l The remaining bandwidth, delay, and packet loss rate of the link are recorded respectively;

[0027] Step 4.5: Set the reward function, d = (d1, d2...d n ) represents the target node set, each target node d i There is at least one simple path to get data from the source node, a finite set Represents the distance from the source node to any target node d i All paths of p di Represents a finite set A unicast path selected in;

[0028] Define the path cost function l(p di ) is used to quantify the cost of the unicast path to a single destination node, which is calculated as follows:

[0029]

[0030] Among them, len(p di ) is the selected path p di The length of m is the number of edges in the multicast tree, and h is a tuning parameter used to prevent the path length from having too much influence on the path selection;

[0031] According to the unicast path cost l(p di ) assigns reward values ​​and obtains the single-step reward value r of path selection step (p di ) can be expressed as:

[0032] r step (p di )=-l(p di ).

[0033] Optionally, the comprehensive reward function r of the entire multicast path in step 5 is whole (p d ) is expressed as:

[0034]

[0035] Among them, S e is the set of all destination nodes that use the edge to e, c di (p) represents the unicast path cost to reach the destination node di. The greater the cost consumed, the smaller the reward value.

[0036] Among them, the piecewise function To adjust the distribution of rewards, the specific form is as follows:

[0037]

[0038] x is an evaluation metric used to measure the relative cost of the target node on the path.

[0039] The present invention provides an intelligent multicast routing optimization method based on deep reinforcement learning and game theory, which abstracts the edge devices in the real scene into an undirected graph network topology structure, reads the network status information data set obtained by using software-defined network technology; sets the link cost for each edge according to the topology information and link information in the data set; establishes an incentive mechanism according to the potential game relationship between unicast paths to prepare for subsequent reinforcement learning training; and then sets the state space, action space, and reward function in deep reinforcement learning according to multicast requirements. The multicast path is constructed by gradually selecting the unicast path to each destination node; the unicast path of the destination node is added to the multicast path in turn and the total reward value is given, and the Nash equilibrium state is reached through training. The present invention significantly reduces the transmission cost by optimizing the selection of multicast routing paths and using the game relationship between unicast paths to share resources. Compared with the existing methods, the present invention has significant advantages in improving transmission efficiency, reducing network costs and dynamic adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 It is a schematic diagram of the steps of an intelligent multicast routing optimization method based on deep reinforcement learning and game theory of the present invention.

[0042] Figure 2 It is a schematic diagram of the connection between the real scene and the network scene in the application scene of the present invention.

[0043] Figure 3 It is a schematic diagram of the potential game relationship between multiple unicast paths in the network topology.

[0044] Figure 4 It is a flowchart of the reward function design of the deep reinforcement learning method of the present invention.

[0045] Figure 5 It is a state space matrix diagram in the deep reinforcement learning method of the present invention. DETAILED DESCRIPTION

[0046] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0047] See also Figure 1 The present invention provides an intelligent multicast routing optimization method based on deep reinforcement learning and game theory, comprising the following steps:

[0048] S1: Abstract the edge devices in the real scene into an undirected graph network topology structure, read the network status information data set obtained by software-defined network technology; obtain multicast requirements, source node s, and destination node;

[0049] S2: Set the link cost for each edge according to the link information of the topology, and select k unicast paths to each destination node based on the number of hops;

[0050] S3: Based on the potential game relationship between unicast paths, a model is established to incentivize unicast paths to select more shared edges when building multicast paths;

[0051] S4: Set the state space, action space and reward function in deep reinforcement learning, and construct the multicast path by gradually selecting the unicast path to each destination node;

[0052] S5: The unicast path of the destination node is added to the multicast path in turn to give a total reward value, and the Nash equilibrium state is reached through training, that is, the reward value cannot be increased by changing any thin path.

[0053] The following is further described in conjunction with specific embodiments and execution steps:

[0054] See also Figures 2 to 5 ,like Figure 2 As shown, in step S1, the edge devices such as sensor nodes and mobile terminals in the real scene are abstracted into an undirected graph network topology structure. Software-defined networking is a way to implement network virtualization, which realizes flexible control of network traffic, can simulate real network information for experimental reference, and provides a good platform for innovation in network scheduling and application. The method is implemented using the known topology and the obtained link information. First, the pickle file storing the network status information of different topologies is read.

[0055] Specifically, in a given network G, G = (V, E), V is the node in G, E is the set of links in G, e ij∈E represents the link between node i and node j. Any node can be used as a source node or a destination node. The source node is represented by s, and D = {d1, d2...d n} represents the target node set; the minimum Steiner tree from the source node to multiple destination nodes is defined as finding the minimum cost tree from s to D in G.

[0056] Step S2: Set the link cost for each edge according to the link bandwidth delay information at different times, and select k unicast paths to each destination node based on the number of hops to facilitate subsequent reinforcement learning method training.

[0057] The edge cost between any two nodes i and j The definitions are as follows:

[0058]

[0059] τ1+τ2+τ3=1

[0060] in, The total cost, α, β, γ, depends on the normalized results of the remaining bandwidth, average propagation delay and average packet loss rate between the two nodes, respectively. τ is the weight coefficient. According to the needs, these weights can be adjusted to balance the remaining bandwidth, delay and packet loss rate.

[0061] The normalization method used for α(e) is defined as follows:

[0062]

[0063] where bw e is the residual bandwidth of edge e between nodes i and j, R τ is the reference bandwidth used to normalize the bandwidth. e The larger the value of , the higher the link transmission efficiency and the higher the throughput. The value of α(e) is related to bw e When the value of α(e) is 1, it means that the current link is unavailable, and the corresponding cost value will also increase accordingly.

[0064] The normalized calculation method for β(e) is as follows:

[0065]

[0066] Among them, delay e represents the average transmission delay between nodes i and j at the current time, w δ It is the delay used for normalization in the network structure. The present invention sets it here as the maximum transmission delay measured in all links at the current moment. The delay of the current link e is proportional to β(e), and the corresponding cost value will also be affected.

[0067] The normalized calculation method for γ(e) is as follows:

[0068]

[0069] Among them, loss e represents the packet loss rate of the information sent and received in edge e between nodes i and j, T β It is the maximum packet loss rate of all links in the network structure at the current moment. The packet loss rate of the current link e is proportional to β(e).

[0070] Step S3: Based on the potential game relationship between the unicast paths reaching each destination node, an edge model is established to incentivize unicast paths to choose sharing, and the pros and cons of sharing and not participating in sharing are weighed to improve network resource utilization.

[0071] There is only one path to some destination nodes, which means that there are no more choices for this path, and the network resources consumed are fixed. However, there may be multiple paths from the source node to some other destination nodes, so there are multiple path choices. Therefore, the multicast route to multiple destination nodes has multiple path components, such as Figure 3 The multicast paths reaching the four destination nodes are split into two typical forms. Their weights represent the path cost, that is, how much network resources are consumed. It can be calculated that the two cost values ​​are 23 and 24 respectively. Among the numerous path compositions, there may be multicast trees smaller than 23. Therefore, the purpose of the present invention is to find the minimum cost multicast tree. No target node can make the cost of the multicast tree smaller by changing its path, so that its multiple paths are in an NE state, reaching a global relatively optimal state.

[0072] Nash equilibrium in the path selection game (PSGame), if a path selection configuration If it is a Nash equilibrium (NE), it means that no destination node can further reduce its cost by unilaterally changing its path, that is:

[0073]

[0074] in, Indicates arrival i The optimal path of To reach d i The optimal path to the destination node is from a single source node s to multiple destination nodes D = {d1, d2...d n The multiple paths to a destination node in One of the multiple paths is selected as When a path is determined, other paths except this path are expressed as

[0075] Best Path (BP): In the path selection game, given the path selection configuration p of all receiving nodes except i -i , the optimal path for node i is the path that maximizes its reward relative to the path selection of other nodes and is defined as:

[0076]

[0077] In this definition, the optimal path is the best strategy for an individual node based on the choices of other nodes, and the Nash equilibrium is the stable state of the entire system.

[0078] In step S4, the deep reinforcement learning algorithm is used to set the state space, action space, and reward function in deep reinforcement learning. The overall flow chart is as follows: Figure 4 As shown. The multicast path is constructed by gradually selecting the unicast path to each destination node. Each action selection, that is, the insertion of a unicast path, gives a corresponding reward value. The unicast path added later is affected by the path added earlier. The maximum reward value for the cooperation between the paths is found.

[0079] S4.1 uses the collected network link information (NLIs) to train the agent offline, learn how to build the optimal multicast path, update the network parameters, and store the trained routing strategy in the experience replay buffer.

[0080] S4.2 This method is based on the dual deep double Q network (D3QN) reinforcement learning algorithm, which uses two networks to separate action selection and action evaluation in order to solve the problem of large errors caused by overestimation of Q value. The Q network is a deep Q network, which is a neural network used to approximate the Q value function in reinforcement learning. The Q value function represents the cumulative expected return (reward value) that the agent can obtain after performing an action a in a certain state s. The core idea is to use the main Q network Q main Select actions and use the target Q network Q target Evaluate these actions. Here is the target Q value update formula:

[0081]

[0082] Among them, the role of the main Q network is to select the optimal action, that is, to select the action a' that maximizes the Q value in state s'.

[0083] S4.3 uses a loss function based on temporal difference (TD) to update the parameters of the main Q network. Its expression is:

[0084]

[0085] Among them, y is the target Q value, and the loss function aims to minimize the Q network output Q(s,a;θ main ) and the target value y, thereby optimizing the parameters θ of the Q network main .

[0086] The design of S4.4 state space uses five different matrices such as Figure 5 As shown in the figure, the purpose is to help the intelligent agent make the best decision in a complex network environment. It includes the topological connection matrix, the selected path matrix and three link information matrices.

[0087] s t =[M t ,M p ,M b ,M d ,M l ]

[0088] Topological connection matrix M t The connection between all nodes in the network is recorded and displayed in the form of an adjacency matrix. Each element indicates whether there is a direct physical connection between two nodes. p It is used to track the decision-making process in the multicast tree construction process, mark the selected path, avoid repeated selection, and assist in subsequent decision-making. b ,M d ,M l The remaining bandwidth, delay and packet loss rate of the link are recorded respectively.

[0089] S4.5 For each destination node d i k candidate paths are provided, and k relatively short paths are selected based on the shortest path algorithm. These path sets represent multiple candidate paths to each destination node. Each destination node has multiple candidate paths, and a multicast tree is constructed by selecting candidate paths to each destination node in turn.

[0090] S4.6 The reward function evaluates the action selection and determines the learning direction of the agent. The reward function flow chart is as follows: Figure 4 As shown in the figure, the goal is to let the agent learn the optimal path selection strategy and set the optimal strategy for the path to a specific target node through the reward value. n ) represents the target node set. Each target node d i There is at least one simple path to get data from the source node. Finite set Represents the distance from the source node to any target node d iAll paths of .

[0091] As more destination nodes choose the same edge e, the cost of the edge The cost is proportionally distributed among all the destination nodes of the edge, that is, the cost shared by each destination node decreases as the number of nodes participating in path sharing increases. For each destination node di, the cost calculation formula using edge e is:

[0092]

[0093] Among them, |S e | is the number of destination nodes using edge e in the current multicast path selection. As the multicast tree is constructed, nodes will share the link cost according to the optimal path to ensure the efficiency of path construction.

[0094] In multicast path construction, the total cost of the path selected by each target node is the sum of the cost of all edges on the path. That is, for a path pi occupied by a target node di, its cost can be expressed as:

[0095]

[0096] in, is the cost sharing value of the target node di on edge e. Further expansion, the path cost of each target node di can also be expressed according to the sharing of each edge in the path. The cost of each edge e By using the target node S of this edge e If the cost is evenly distributed, the cost of each path is:

[0097]

[0098] Among them, S e is the number of destination nodes using edge e in the current path. This formula shows that as more destination nodes share the same path, the cost shared by a single node on the path will decrease. Through this sharing mechanism, the cost of multi-destination paths can be effectively balanced, ensuring that the path selected in the multicast tree is efficient and fair.

[0099] After selecting a unicast path to a single destination node, the evaluation of the path is crucial. The present invention defines the path cost function l(p di ) to quantify the cost of this path, which is calculated as follows:

[0100]

[0101] Here, m is the number of edges in the multicast tree, and h is a tuning parameter used to prevent the path length from having too much influence on the path selection and ensure that other network performance is not negatively affected.

[0102] The present invention is based on l(p di ) Design the path selection allocation reward value under the current state. Specifically, the single-step reward value r of path selection step (p di ) can be expressed as:

[0103] r step (p di )=-l(p di ).

[0104] Step S5: The unicast path of the destination node is added to the multicast path in turn to give a total reward value, and the Nash equilibrium state is reached through training, that is, the reward value cannot be increased by changing one path. Determine the multicast path.

[0105] After the multicast tree is constructed, the present invention needs to conduct a comprehensive evaluation of the entire multicast tree. Through the following formula, the present invention can obtain the evaluation index x, which represents the relative value of the target node cost:

[0106]

[0107] Next, define a piecewise function φ(x) to adjust the distribution of rewards. The purpose of designing the piecewise function is to optimize the cost sharing during path selection, so that the cost sharing value of the target node with a larger weight on the edge e is further increased, while the cost sharing value of the target node with a smaller weight is reduced. In this way, the agent can be guided to give priority to those paths that have a greater impact on the overall performance of the network, ensuring that the critical path obtains more resources and the burden of the secondary path is reduced, thereby optimizing the overall efficiency of the multicast tree. The specific form of the piecewise function is as follows:

[0108]

[0109] Among them, x is an evaluation index used to measure the relative cost of the target node on the path. The function φ(x) is used to enhance the polarization effect. This design enables the system to allocate resources more efficiently and optimize network performance during path sharing. Finally, the comprehensive reward function r of the entire multicast path is whole (p d ) is expressed as:

[0110]

[0111] The reward function design mechanism ensures that in the process of multicast tree construction, path selection not only considers the cost and hop count, but also integrates the overall network performance. By introducing a dynamic reward mechanism and path utility calculation, the present invention can promote a more intelligent and efficient path selection strategy, so that the constructed multicast tree achieves a better balance between performance and cost.

[0112] What is disclosed above is only a preferred embodiment of the present invention, and it certainly cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. An intelligent multicast routing optimization method based on deep reinforcement learning and game theory, characterized in that: The following steps are involved: Step 1: Abstract the edge devices in the real scene into an undirected graph network topology structure, read the network status information data set obtained by using software-defined network technology; obtain multicast requirements, source node s, and destination node; Step 2: Set the link cost for each edge according to the network status information of the topology, and select k relatively short unicast paths to each destination node based on the number of hops; Step 3: Based on the potential game relationship between unicast paths, establish a model that incentivizes unicast paths to select more shared edges when building multicast paths; Step 4: Set the state space, action space and reward function in deep reinforcement learning, and construct the multicast path by gradually selecting the unicast path to each destination node; Step 5: The unicast path of the destination node is added to the multicast path in turn and the total reward value is given. The Nash equilibrium state is reached through training, that is, the reward value cannot be increased by changing any thin path.

2. The intelligent multicast routing optimization method based on deep reinforcement learning and game theory as claimed in claim 1, characterized in that: In step 1, in the undirected graph network G, G = (V, E), V is a node in G, E is a set of links in G, e ij ∈E represents the link between node i and node j. Any node can be used as a source node or a destination node. The source node is represented by s, and D = {d1, d2...d n } represents the target node set; the minimum Steiner tree from the source node to multiple destination nodes is defined as the minimum cost tree from s to D in G. The minimum cost tree is the multicast tree with the highest reward value obtained using the reinforcement learning algorithm.

3. The intelligent multicast routing optimization method based on deep reinforcement learning and game theory as claimed in claim 1, characterized in that: The network status information dataset includes pickle files of network status information of different topological structures, which are from traffic generation tools to simulate traffic conditions for 24 hours a day. The network status information includes the remaining bandwidth of the link and the link, the average transmission delay and the link packet loss rate; During the execution of step 2, the remaining bandwidth, average transmission delay, and link packet loss rate in the link information are normalized and weighted to obtain the edge cost between any two nodes i and j. The definitions are as follows: τ1+τ2+τ3=1 in, The total cost, α, β, γ are the normalized results of the remaining bandwidth, average propagation delay and average packet loss rate between two nodes, respectively, and τ is an adjustable weight coefficient.

4. The intelligent multicast routing optimization method based on deep reinforcement learning and game theory as claimed in claim 1, characterized in that: In step 3, the multicast tree with the minimum cost is constructed by explaining the potential game relationship between the paths. In the multicast tree constructed according to the link cost information, any unicast path cannot be changed to make the constructed multicast tree have a lower cost, so that the multicast tree maintains a Nash equilibrium state.

5. The intelligent multicast routing optimization method based on deep reinforcement learning and game theory as claimed in claim 1, characterized in that: The execution process of step 4 is specifically the process of constructing a multicast tree using a reinforcement learning method, and includes the following steps: Step 4.1: Use the collected network link information to train the agent offline, learn how to build the optimal multicast path, update the network parameters, and store the trained routing strategy in the experience replay buffer; Step 4.2: Based on the dual-depth dual-Q network reinforcement learning algorithm, two networks are used to separate action selection and action evaluation. The Q network is a deep Q network, which is a neural network used to approximate the Q-value function in reinforcement learning. The Q-value function represents the cumulative expected return / reward value that the agent can obtain after performing an action a in a certain state s. Step 4.3: Update the parameters of the Q network using the loss function based on temporal differences; Step 4.4: The design of the state space uses five different matrices, including the topological connection matrix, the selected path matrix, and three link information matrices, which are expressed as follows: s t =[M t ,M p ,M b ,M d ,M l ] Among them, the topological connection matrix M t The connection between all nodes in the network is recorded and displayed in the form of an adjacency matrix. Each element indicates whether there is a direct physical connection between two nodes. The selected path matrix M p It is used to track the decision-making process in the multicast tree construction process, mark the selected path, avoid repeated selection, and assist in subsequent decision-making; the three link information matrices M b ,M d ,M l The remaining bandwidth, delay, and packet loss rate of the link are recorded respectively; Step 4.5: Set the reward function, d = (d1, d2...d n ) represents the target node set, each target node d i There is at least one simple path to get data from the source node, a finite set Represents the distance from the source node to any target node d i All paths of p di Represents a finite set A unicast path selected in; Define the path cost function l(p di ) is used to quantify the cost of the unicast path to a single destination node, which is calculated as follows: Among them, len(p di ) is the selected path p di length; m is the number of edges in the multicast tree, and h is a tuning parameter used to prevent the path length from having too much influence on the path selection; According to the unicast path cost l(p di ) assigns reward values ​​and obtains the single-step reward value r of path selection step (p di ) can be expressed as: r step (p di )=-l(p di )。 6. The intelligent multicast routing optimization method based on deep reinforcement learning and game theory as claimed in claim 1, characterized in that: The comprehensive reward function r of the entire multicast path in step 5 is whole (p d ) is expressed as: Among them, S e is the set of all destination nodes that use the edge to e, c di (p) represents the unicast path cost to reach the destination node di. The greater the cost consumed, the smaller the reward value. Among them, the piecewise function To adjust the distribution of rewards, the specific form is as follows: x is an evaluation metric used to measure the relative cost of the target node on the path.