Multi-modal network global routing optimization method and device, electronic equipment and medium

By building a hierarchical structure and global routing optimization method in a multimodal network, the problems of insufficient coordination and low global characteristics in a multi-domain data environment are solved, and network transmission efficiency and cross-domain collaboration capabilities are improved.

CN120358190AActive Publication Date: 2025-07-22WUHAN UNIV

Patent Information

Application Number
CN202510860413.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Traditional network optimization methods lack coordination and low global feature utilization in multimodal network environments, resulting in insufficient network transmission efficiency and cross-domain collaboration capabilities.

Method used

By building a hierarchical structure of a software-defined network, dividing the multimodal network topology into a logical subnet, using the global policy model and path search algorithm to generate the global route of the multimodal network, combining dynamic weight allocation and federated learning to optimize local policy parameters, an efficient global routing solution is generated.

Benefits of technology

It significantly improves the overall transmission efficiency and cross-domain collaboration capabilities of the network, solves the problem of insufficient collaboration in multi-domain data environment, and realizes multi-dimensional QoS optimization with low latency, low packet loss rate and high resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358190A_ABST
    Figure CN120358190A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network communication and intelligent optimization, in particular to a multi-mode network global routing optimization method and device, electronic equipment and a medium. Comprising the following steps: constructing a data plane and a control plane according to a hierarchical structure of the software defined network, performing multi-modal analysis on an actual network topology of the data plane according to a modal type supported by a switch, and dividing the actual network topology into logic subnets of different modals based on a multi-modal analysis result; generating a set of N paths satisfying a preset condition between the source switch and the destination switch by using a preset path searching algorithm based on a network topology relationship between the logic subnets of different modalities; and based on the N path sets and the global controller of the control plane, generating a global route of the multi-modal network by using a preset global strategy model. Therefore, the problems that a traditional network optimization method is insufficient in collaboration and low in global characteristic utilization rate in a multi-domain data environment are solved, and the network transmission efficiency and the cross-domain collaboration capability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network communication and intelligent optimization, and particularly to an optimization method, device, electronic device and medium for global routing of a multi-modal network. Background Art

[0002] With the continuous expansion of the network scale and the surge in cross-domain data transmission requirements, many bottlenecks have emerged in the traditional network architecture in terms of routing optimization. In a multi-modal network, there are completely different identification methods and addressing mechanisms among different modalities (such as traditional IP (Internet Protocol) networks, content-driven networks, location-based networks, and emerging FlexIP (Flexible Variable Length Internet Protocol Addressing) networks). Traditional routing protocols are difficult to adapt in such heterogeneous networks. At the same time, seamless switching and data intercommunication in the heterogeneous identification space have become key challenges, because the cooperation efficiency between networks directly determines the performance of cross-domain path optimization. The diversity and dynamic changes of multi-modal traffic have put forward higher requirements for the real-time performance, privacy protection, and resource scheduling capabilities of routing decisions.

[0003] In related technologies, a deep deterministic policy gradient algorithm is introduced to make optimal routing selections according to the current state of the network. Combined with asynchronous federated learning, the parameters uploaded by deep reinforcement learning agents in each domain are fused to obtain a deep reinforcement agent with stronger decision-making capabilities, and this agent is applied to the decision-making process of cross-domain routing. Each domain forwards cross-domain data according to the inter-domain routing and the intra-domain routing policy.

[0004] However, this method still has problems such as insufficient cooperation and low utilization rate of global characteristics in a multi-domain data environment, which need to be solved urgently. Summary of the Invention

[0005] The present invention provides an optimization method, device, electronic device and medium for global routing of a multi-modal network to solve the problems of insufficient cooperation and low utilization rate of global characteristics in a multi-domain data environment in traditional network optimization methods, and significantly improve the overall transmission efficiency and cross-domain cooperation ability of the network.

[0006] To achieve the above object, the first aspect embodiment of the present invention proposes an optimization method for global routing of a multi-modal network, including the following steps: Construct a data plane and a control plane according to the hierarchical structure of the software-defined network, where the data plane includes the actual network topology and the routing environment, and the control plane includes at least one domain controller and a global controller; Perform multimodal analysis on the actual network topology according to the modal types supported by the switch, and divide the actual network topology into logical subnets of different modalities based on the multimodal analysis results; Based on the network topology relationship between the logical subnets of different modalities, use a preset path finding algorithm to generate a set of N paths that meet the preset conditions between the source switch and the destination switch, where N is a positive integer; Based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller, use a preset global policy model to generate the global routing of the multimodal network.

[0007] According to an embodiment of the present invention, the generating the global routing of the multimodal network based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller using a preset global policy model includes: Obtain the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information of the multimodal network based on the set of N paths, the actual network topology, and the routing environment; Based on the boundary nodes, the boundary link information, the transmission parameters of each domain controller, and the routing link information, use the global controller and the preset global policy model to generate the global routing of the multimodal network.

[0008] According to an embodiment of the present invention, before generating the global routing of the multimodal network based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller using a preset global policy model, it further includes: Obtain the performance parameters of the intra-domain links based on the set of N paths, train a preset local model based on the performance parameters of the intra-domain links, and obtain the local policy parameters of each domain based on the preset local model; Based on the dynamic weight allocation mechanism, use the global controller to fuse the local policy parameters of each domain, and generate the preset global policy model based on the fused parameters.

[0009] According to an embodiment of the present invention, the obtaining the performance parameters of the intra-domain links based on the set of N paths and training a preset local model based on the performance parameters of the intra-domain links includes: Based on the set of N paths, use in-band network telemetry technology to obtain the performance parameters of the intra-domain links in real time, aggregate the performance parameters according to a preset time window, and construct at least one short-term traffic state matrix; Input each short-term traffic status matrix into the current Actor network to obtain an initial path selection corresponding to each current short-term traffic status; Input each of the short-term traffic statuses and the initial path selections corresponding to each of the current short-term traffic statuses into the current Critic network to obtain an expected return corresponding to each of the current short-term traffic statuses; Input the initial path selections corresponding to each of the current short-term traffic statuses and the expected returns corresponding to each of the current short-term traffic statuses into the target network for training to obtain an optimized Actor network and an optimized Critic network; Using a preset feedback mechanism, obtain the preset local model based on the optimized Actor network and the optimized Critic network.

[0010] According to an embodiment of the present invention, after fusing the local policy parameters of each domain using the global controller based on the dynamic weight allocation mechanism and generating the preset global policy model based on the fused parameters, it further includes: Update the preset local model based on the model parameters of the preset global policy model.

[0011] According to an embodiment of the present invention, after performing multimodal analysis on the actual network topology according to the modal types supported by the switch and dividing the actual network topology into logical subnets of different modalities based on the multimodal analysis results, it further includes: Based on the service requirements and resource requirements of the logical subnets of different modalities, determine the bandwidth resources of each logical subnet; Perform bandwidth allocation for each logical subnet according to the bandwidth resources of each logical subnet.

[0012] According to an embodiment of the present invention, the generating, based on the network topology relationship between the logical subnets of different modalities, a set of N paths that meet preset conditions between the source switch and the destination switch by using a preset path finding algorithm includes: Determine the source node and the target node; Based on the network topology relationship between the logical subnets of different modalities, use the preset path finding algorithm to calculate the path planning situation from the source node to the target node; Based on each path in the path planning situation, determine at least one parent node and the length of each path; Based on the length of each path, start from the target node and trace back in reverse according to at least one parent node of each path until returning to the source node to obtain a set of N paths that meet preset conditions between the source switch and the destination switch.

[0013] According to the optimization method for global routing of a multimodal network proposed by an embodiment of the present invention, based on the hierarchical structure of a software-defined network, a data plane and a control plane can be constructed. Then, according to the modal types supported by switches, a multimodal analysis is performed on the actual network topology in the data plane, and based on the results of the multimodal analysis, the actual network topology is divided into logical subnets of different modalities. Based on the network topology relationships between the logical subnets of different modalities, a set of N paths that meet preset conditions between a source switch and a destination switch is generated by using a preset path finding algorithm; based on the set of N paths and the global controller in the control plane, a global routing of the multimodal network is generated by using a preset global policy model. Thereby, the problems of insufficient collaboration and low utilization rate of global characteristics in traditional network optimization methods in a multi-domain data environment are solved, and the overall transmission efficiency and cross-domain collaboration ability of the network are significantly improved.

[0014] To achieve the above object, an embodiment of the second aspect of the present invention proposes an optimization device for global routing of a multimodal network, including: A construction module, configured to construct a data plane and a control plane according to the hierarchical structure of a software-defined network, where the data plane includes an actual network topology and a routing environment, and the control plane includes at least one domain controller and a global controller; A division module, configured to perform a multimodal analysis on the actual network topology according to the modal types supported by switches, and divide the actual network topology into logical subnets of different modalities based on the results of the multimodal analysis; A first generation module, configured to generate a set of N paths that meet preset conditions between a source switch and a destination switch by using a preset path finding algorithm based on the network topology relationships between the logical subnets of different modalities, where N is a positive integer; A second generation module, configured to generate a global routing of the multimodal network by using a preset global policy model based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller.

[0015] According to an embodiment of the present invention, the second generation module is specifically configured to: Obtain boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information of the multimodal network based on the set of N paths, the actual network topology, and the routing environment; Generate the global routing of the multimodal network by using the global controller and the preset global policy model based on the boundary nodes, the boundary link information, the transmission parameters of each domain controller, and the routing link information.

[0016] According to an embodiment of the present invention, before generating the global routing of the multimodal network by using a preset global policy model based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller, the second generation module further includes: An obtaining unit, configured to obtain performance parameters of intra-domain links based on the set of N paths, train a preset local model based on the performance parameters of the intra-domain links, and obtain local policy parameters of each domain based on the preset local model; A generating unit, configured to fuse the local policy parameters of each domain by using the global controller based on a dynamic weight allocation mechanism, and generate the preset global policy model based on the fused parameters.

[0017] According to an embodiment of the present invention, the obtaining unit is specifically configured to: Based on the set of N paths, use in-band network telemetry technology to obtain performance parameters of intra-domain links in real time, aggregate the performance parameters according to a preset time window, and construct at least one short-term traffic state matrix; Input each short-term traffic state matrix into the current Actor network to obtain an initial path selection corresponding to each current short-term traffic state; Input each short-term traffic state and the initial path selection corresponding to each current short-term traffic state into the current Critic network to obtain an expected return corresponding to each current short-term traffic state; Input the initial path selection corresponding to each current short-term traffic state and the expected return corresponding to each current short-term traffic state into the target network for training to obtain an optimized Actor network and an optimized Critic network; Use a preset feedback mechanism to obtain the preset local model based on the optimized Actor network and the optimized Critic network.

[0018] According to an embodiment of the present invention, after fusing the local policy parameters of each domain by using the global controller based on the dynamic weight allocation mechanism and generating the preset global policy model based on the fused parameters, the generating unit is further configured to: Update the preset local model based on the model parameters of the preset global policy model.

[0019] According to an embodiment of the present invention, after performing multimodal analysis on the actual network topology according to the modal types supported by the switch and dividing the actual network topology into logical subnets of different modalities based on the multimodal analysis result, the dividing module is further configured to: Determine the bandwidth resources of each logical subnet based on the service requirements and resource requirements of the logical subnets of the different modalities. Perform bandwidth allocation for each logical subnet according to the bandwidth resources of each logical subnet.

[0020] According to an embodiment of the present invention, the first generation module is specifically configured to: Determine a source node and a destination node; Based on the network topology relationship between the logical subnets of the different modalities, use the preset path finding algorithm to calculate the path planning situation from the source node to the destination node; Determine at least one parent node and the length of each path based on each path in the path planning situation; Based on the length of each path, start from the destination node and trace back in reverse according to at least one parent node of each path until returning to the source node, to obtain an N-path set that meets the preset conditions between the source switch and the destination switch.

[0021] The optimization device for global routing of a multi-modal network proposed in the embodiment of the present invention can construct a data plane and a control plane according to the hierarchical structure of the software-defined network. Then, perform multi-modal analysis on the actual network topology in the data plane according to the modal types supported by the switch, and divide the actual network topology into logical subnets of different modalities based on the results of the multi-modal analysis. Based on the network topology relationship between the logical subnets of different modalities, use the preset path finding algorithm to generate an N-path set that meets the preset conditions between the source switch and the destination switch; based on the N-path set and the global controller of the control plane, use the preset global policy model to generate the global routing of the multi-modal network. Thus, the problems of insufficient collaboration and low utilization rate of global characteristics in the traditional network optimization method in a multi-domain data environment are solved, and the overall transmission efficiency and cross-domain collaboration ability of the network are significantly improved.

[0022] To achieve the above object, an embodiment of the third aspect of the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the optimization method for global routing of a multi-modal network as described in the above embodiment.

[0023] To achieve the above object, an embodiment of the fourth aspect of the present invention proposes a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to be used to implement the optimization method for global routing of a multi-modal network as described in the above embodiment.

[0024] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0025] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the drawings, in which: Figure 1 FIG. is a flowchart of an optimization method for multi-modal network global routing according to an embodiment of the present invention; Figure 2 FIG. is a schematic diagram of an architecture for optimizing multi-modal network global routing based on federated reinforcement learning according to an embodiment of the present invention; Figure 3 FIG. is a schematic diagram of multi-modal network slice resource partitioning based on an SDN (Software-Defined Networking controller) controller according to an embodiment of the present invention; Figure 4 FIG. is a schematic diagram of a data packet format for network data measurement based on INT (Inband Network Telemetry) according to an embodiment of the present invention; Figure 5 FIG. is a schematic diagram of the principle of reinforcement learning according to an embodiment of the present invention; Figure 6 FIG. is a block schematic diagram of an optimization device for multi-modal network global routing according to an embodiment of the present invention; Figure 7 FIG. is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Embodiments

[0026] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0027] The following describes an optimization method, device, electronic device, and medium for multi-modal network global routing according to embodiments of the present invention with reference to the drawings.

[0028] Figure 1 FIG. is a flowchart of an optimization method for multi-modal network global routing according to an embodiment of the present invention.

[0029] Exemplarily, as Figure 1 shown, the optimization method for multi-modal network global routing includes the following steps: In step S101, according to the hierarchical structure of the software-defined network, a data plane and a control plane are constructed. The data plane includes the actual network topology and the routing environment, and the control plane includes at least one domain controller and a global controller.

[0030] It can be understood that the software-defined network is a new type of network architecture that centrally controls the behavior and traffic of the network through software, making the network more flexible and programmable.

[0031] Specifically, as Figure 2 shown, according to the hierarchical structure of the software-defined network, the network can be divided into two main planes: the data plane and the control plane. The data plane is mainly responsible for forwarding network traffic and includes various physical or virtual devices in the network, such as switches, routers, etc. These devices constitute the actual network topology structure and execute routing decisions to determine how data packets move in the network. The data plane can use the Mininet simulation platform to deploy the actual network topology and routing environment, and reasonably divide the network into three domains according to geographical distribution and traffic characteristics and deploy them in three virtual machines. The control plane is responsible for the intelligent decision-making and management of the entire network and consists of an ONOS (Open Network Operating System) domain controller and an ONOS global controller to form a distributed hierarchical management structure. Each domain is managed and optimized within the domain by a domain controller. The ONOS domain controller can master the complete network topology and performance data within a specific domain. The global controller is used to coordinate across domains for each domain controller and path selection to ensure the collaborative work of the entire network. The control plane determines the configuration and policies of the network, is responsible for formulating routing rules and policies, and then distributes them to the data plane for execution.

[0032] For example, the actual network topology of the data plane is represented by , where is the set of nodes, and is the set of links. The node set of each domain is , and the link set is . The control plane consists of three domain controllers and one global controller to form a distributed hierarchical control architecture. Each domain controller is responsible for topology discovery, performance monitoring, and optimization within the domain, and masters the complete information of the node set and the link set within the domain. The global controller masters the information of the border switches across domains and collaborates with the domain controller through the federated learning mechanism for cross-domain path optimization.

[0033] By separating the control layer and the data forwarding layer of network devices, flexible management and optimization of the network can be achieved.

[0034] In step S102, perform multimodal analysis on the actual network topology according to the modal types supported by the switches, and divide the actual network topology into logical subnets of different modalities based on the results of the multimodal analysis.

[0035] It can be understood that a switch is a network device used to connect computers and other network devices in a network and can forward data packets to the correct ports according to the destination information of the data packets. Different switches support different modalities, and the modal types include, for example, IP, Geo (the geographical location mode of the switch), ID (Virtual Local Area Network Identifier, virtual local area network identifier (identity identification network)), MF (Multi-Factor Authentication, multi-factor authentication network), NDN (Named Data Networking, named data network), etc.

[0036] Specifically, when conducting network design and optimization, an important operation is to identify the different modal types supported by the switches in the actual network topology, perform multimodal analysis on the actual network topology structure (physical network), and based on the results of this multimodal analysis, the actual network topology can be further divided into multiple logical subnets, each corresponding to a specific modality, such as Figure 3 shown. This division helps to better manage and optimize network resources, ensure that network traffic of different modalities can be transmitted efficiently and securely, and is also conducive to the expansion and maintenance of the network.

[0037] In step S103, based on the network topology relationship between the logical subnets of different modalities, use a preset path finding algorithm to generate a set of N paths that meet the preset conditions between the source switch and the destination switch, where N is a positive integer.

[0038] That is to say, according to the network topology relationship between the logical subnets of different modalities, a preset path finding algorithm can be used to generate a set of N paths that meet the preset conditions (the shortest) between the source switch and the destination switch, and this set is used as the alternative path set. Here, N is a positive integer, referring to the number of paths.

[0039] For ease of understanding, the following details how to obtain a set of N paths that meet the preset conditions.

[0040] As a possible implementation method, in some embodiments, based on the network topology relationship between logical subnets of different modalities, an N-path set that meets preset conditions is generated between the source switch and the destination switch by using a preset path finding algorithm, including: determining a source node and a target node; calculating the path planning situation from the source node to the target node by using the preset path finding algorithm based on the network topology relationship between logical subnets of different modalities; determining at least one parent node and the length of each path based on each path in the path planning situation; and backtracking from the target node according to at least one parent node of each path based on the length of each path until returning to the source node, to obtain an N-path set that meets preset conditions between the source switch and the destination switch.

[0041] Specifically, in the embodiments of the present invention, the INT technology can be first used to monitor key performance parameters such as traffic and delay of switch ports in real time. Then, an adjacency matrix of a local network can be generated based on the collected performance data. The adjacency matrix is a two-dimensional array, where each element represents the link weight between two nodes, and the link weight can be delay, bandwidth utilization rate, or other performance metrics. Finally, the connection information between different local networks is integrated to construct a global network topology map spanning multiple domains (i.e., the network topology relationship between logical subnets of different modalities).

[0042] Further, based on the global network topology map, all the shortest paths between the source node and the target node are calculated by using a preset path finding algorithm (such as the Dijkstra algorithm, which is a classic shortest path algorithm and can find the path with the minimum total weight (such as total delay) among the paths passing between two points in the network). The Dijkstra algorithm can record the parent node and path length of each node in each path during the calculation process. By backtracking the parent node information, N shortest paths are extracted from the target node back to the source node and stored in the alternative path pool to provide an alternative path set for route optimization. This alternative path set can be used to quickly select the optimal path for data transmission when the network fails or route optimization is required, thereby improving the reliability and efficiency of the network.

[0043] The following details how to obtain the alternative path set by backtracking the parent node information.

[0044] Specifically, the actual network topology of the data plane is represented by where is the node set, is the link set, considering the weighted directed graph Each edge has a weight . To extract from the source node s to the target nodet N shortest paths. First, a source node is given s and a standard Dijkstra's algorithm is run once. The role of Dijkstra's algorithm is to calculate the shortest paths from the source node s to all other nodes. During this process, the parent node and path length of each node are recorded, and the path can be extracted by backtracking the parent node information. Then, the shortest path from the source node s to the target node t is extracted and recorded. To avoid reselecting the same path, the selected path can be removed from the graph, and a large value is added to the weight of each edge in the selected path to ensure that this path will not be selected again. Then, Dijkstra's algorithm is run again on the modified graph to calculate the new shortest paths, and the new paths are extracted and recorded. At this time, since the weights of the selected paths have been adjusted, the calculation results no longer contain the selected paths, thus achieving path deduplication. This process is continuously repeated until the N shortest paths from the source node s to the target node t are all extracted. Finally, these N shortest paths are returned to complete the extraction of multiple shortest paths from the source node s to the target node t .

[0045] In step S104, based on the set of N paths, the actual network topology, the routing environment, at least one domain controller, and the global controller, a global routing for the multi-modal network is generated using a preset global policy model.

[0046] Among them, the preset global policy model is pre-trained to guide the global controller on how to generate and adjust the routing. The global routing refers to a routing scheme generated within the entire network scope. It is not for a single path but a routing strategy that covers the entire network.

[0047] That is to say, after obtaining the set of N paths that meet the preset conditions, based on this set of N paths, the actual network topology, the routing environment, at least one domain controller, and the global controller of the control plane, a preset global policy model can be used to determine how to generate and optimize the routing in the network. Finally, through this mechanism, a global routing for the multi-modal network can be generated, that is, a routing scheme that can adapt to various network situations and requirements.

[0048] As a possible implementation, in some embodiments, based on the set of N paths, the actual network topology, the routing environment, at least one domain controller, and the global controller, the global routing of the multi-modal network is generated by using a preset global policy model, including: obtaining the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information of the multi-modal network based on the set of N paths, the actual network topology, and the routing environment; generating the global routing of the multi-modal network by using the global controller and the preset global policy model based on the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information.

[0049] It can be understood that the boundary nodes refer to the nodes connecting different domains, and these nodes play a bridging role in the multi-modal network. The boundary links refer to the links connecting different domains, and these links are the channels for cross-domain data transmission. The transmission parameters may include link bandwidth, delay, packet loss rate, queue depth, etc., and these parameters reflect the performance and transmission capacity of the links. The routing links refer to the links used for data transmission in the network, including intra-domain links and cross-domain links.

[0050] Specifically, by analyzing the structure of the actual network topology, the boundary nodes, boundary link information, and routing link information connecting different domains can be identified; by using the INT technology to monitor the routing environment (i.e., link performance) in real time, data such as link utilization rate, packet loss rate, and time delay (i.e., the transmission parameters of each domain controller) can be collected. The global controller can generate a complete network view by collecting and integrating the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information of the multi-modal network. Based on this network view, by using the preset global policy model for comprehensive analysis and intelligent decision-making, an efficient and reliable global routing scheme for the multi-modal network can be generated, thereby optimizing the performance of the entire network.

[0051] Next, how to obtain the preset global policy model will be described in detail.

[0052] As a possible implementation, in some embodiments, before generating the global routing of the multi-modal network by using the preset global policy model based on the set of N paths, the actual network topology, the routing environment, at least one domain controller, and the global controller, it further includes: obtaining the performance parameters of the intra-domain links based on the set of N paths, training a preset local model based on the performance parameters of the intra-domain links, and obtaining the local policy parameters of each domain based on the preset local model; fusing the local policy parameters of each domain by using the global controller based on the dynamic weight allocation mechanism, and generating the preset global policy model based on the fused parameters.

[0053] Specifically, the embodiments of the present invention can use INT technology to monitor the performance parameters of intra-domain links in real time, and use these performance parameter data to train a preset local model. After obtaining the preset local model, the local policy parameters of each domain output by the preset local model can be further obtained. The global controller can regularly receive the local policy parameters uploaded by each domain. and fuse the local policy parameters uploaded by each domain based on the dynamic weight allocation mechanism. Considering factors such as intra-domain data quality and model convergence speed, global policy parameters are generated. Based on the global policy parameters, a preset global policy model can be obtained. By distributing the preset global policy model, the global controller can assist the domain controller to complete the collaborative routing optimization within the domain and across domains. In addition, using the federated learning algorithm, the preset global policy model can be optimized based on the local policy parameters of each regional controller, thereby optimizing the cross-domain routing policy; through the global routing policy generated by the global controller, cross-regional routing selection and performance optimization are performed, thereby improving the overall transmission efficiency of the network.

[0054] It should be noted that in view of the high feature similarity between different agents, the embodiments of the present invention adopt the strategy of horizontal federated learning and deploy a set of additional evaluation neural networks. As the core model of federated learning, this network is located on the federated learning server. During the training process, each time the agent's network is trained, first, the network parameters of are copied, and then the agent uses the sliding average update method: to adjust its own parameters. .

[0055] At the th time iteration, the state information is input into the agent's Actor network. Combining the noise generated by the OU (Ornstein-Uhlenbeck, a stochastic process) process, a deterministic action can be output. Through the federated learning mechanism, the model is aggregated and the joint action is executed in the environment. . After the overall routing policy of the environment is updated, a simulation communication process will be carried out. After the simulation communication is completed, the reinforcement learning agent obtains a new local observation result by perceiving the environmental state. and calculates the corresponding reward function according to the feedback of the simulation communication. The agent forms a quadruple consisting of the current state, action, next state, and reward value. Stored in the experience replay pool, it provides sample data for subsequent strategy updates and training, and then conducts further training.

[0056] When the intelligent When the amount of data M in its own experience pool is larger than the preset batch size, the agent can extract experience samples of the corresponding batch size from it for training. Can be input to the target policy network , to select the next action The Q target value of the state-action pair is further obtained from the Bellman equation (dynamic programming equation): :

[0057] in, is the immediate reward function, is the discount factor, is the target Critic network, is the target Actor network, and are the parameters of the target critic and target actor networks respectively.

[0058] By learning the impact of different path selections on the end-to-end performance of the network, evaluation feedback can be provided, and the network parameters can be updated using the minimized loss function. :

[0059] Loss function for evaluating network parameters Take the derivative:

[0060] After each domain reinforcement learning agent has completed its local update, it can send the updated parameters and reward information to the federated learning global server The server first evaluates the Normalize , and then judge the quality of the parameters based on the normalized results, and make more updates in the direction of better results. These normalized values are used as aggregation weights to update the federated global network: , and then sent to each domain reinforcement learning agent.

[0061] When the number of iterations Training frequency for Actor network In order to enable the agent to select the optimal path under different network conditions, the server can use the policy gradient method to adjust the policy network Perform an update:

[0062] That is:

[0063] When the number of iterations reaches the target network update frequency At this time, the soft update method can also be used to update the target evaluation network and the target policy network. The Actor network is used to select paths, and the Critic network is used to evaluate the quality of path selection. The two cooperate and optimize, enabling the intelligent agent to learn the optimal policy for multi-dimensional QoS (Quality of Service) performance indicators.

[0064] During the entire training process, the target network will be updated regularly at a frequency To help stabilize the training process. Both the Critic network and the Actor network can adjust their parameters according to the current state and actions, so as to achieve the optimization of routing. Through the global aggregation of federated learning and the reinforcement learning feedback mechanism, it is ensured that each domain intelligent agent can continuously adjust its strategy without sharing the original data, and improve the overall routing decision-making ability through gradient aggregation, so as to achieve the multi-dimensional Qos target optimization of low latency, low packet loss rate and high resource utilization.

[0065] Furthermore, in some embodiments, after fusing the local policy parameters of each domain by using the global controller based on the dynamic weight allocation mechanism and generating a preset global policy model based on the fused parameters, it further includes: updating the preset local model based on the model parameters of the preset global policy model.

[0066] It can be understood that when the network environment changes, such as a change in the link state or traffic fluctuation, each domain controller can update the preset local model by using the model parameters of the preset global policy model. Through such an update, the domain controller can adjust the routing policy in real time, that is, the rules for determining how data is transmitted in the network, so as to cope with the dynamically changing network conditions and further ensure the stability and efficiency of communication between different domains in the network.

[0067] Next, elaborate on how to obtain the preset local model, that is, the DDPG (Deep Deterministic Policy Gradient) model.

[0068] As a possible implementation method, in some embodiments, obtaining performance parameters of intra-domain links based on an N-path set, and training a preset local model based on the performance parameters of intra-domain links, includes: based on the N-path set, using in-band network telemetry technology to obtain the performance parameters of intra-domain links in real time, and aggregating the performance parameters according to a preset time window to construct at least one short-term traffic state matrix; inputting each short-term traffic state matrix into the current Actor network to obtain an initial path selection corresponding to each current short-term traffic state; inputting each short-term traffic state and the initial path selection corresponding to each current short-term traffic state into the current Critic network to obtain an expected return corresponding to each current short-term traffic state; inputting the initial path selection corresponding to each current short-term traffic state and the expected return corresponding to each current short-term traffic state into the target network for training to obtain an optimized Actor network and an optimized Critic network; using a preset feedback mechanism to obtain a preset local model based on the optimized Actor network and the optimized Critic network.

[0069] It can be understood that the INT (in-band network telemetry) mechanism realizes real-time, packet-level, end-to-end monitoring of the network state by embedding telemetry information in each data packet. First, it is necessary to change the form of the data packet to implement the INT measurement scheme, such as Figure 4 As shown, when performing network monitoring, an INT header can be implanted in the data packet structure. The INT header specifically includes a telemetry identification field If_probe and a type identification field EtherType, as well as a series of fields for recording network state data. The If_probe identification indicates whether the data packet has undergone the tagging process, enabling the switch to check this field to determine whether data telemetry is required. When If_probe == 1, it means the data packet has been tagged and the switch needs to append telemetry metadata to the INT header; conversely, if If_probe == 0, the data packet does not participate in the telemetry process at the current switch and no telemetry data will be appended. The original data packet type is reserved by the EtherType field. If the routing selection of the data packet is controlled by the source routing policy, then a bos (bottom of stack) field can also be included in the INT header to indicate whether the current switch is the first node in the path. In this way, each node in the network can automatically decide whether to perform telemetry tagging on the data packets flowing through it without the intervention of the control plane, thus ensuring the continuity of data measurement and the transparency of data communication.

[0070] Run the send.py file on the host terminal h1. This file can use the Scapy tool to send INT packets to simulate the transmission of multi-modal traffic flows. In addition, the send.py file also sets an appropriate sending rate to ensure that the packet sending not only meets the requirements of the simulated traffic flow but also does not impose too much burden on the network. Run the receive packet file receive.py on the receiving host terminal h2. This file can receive and parse INT packets to parse out the original measurement data, as shown in Table 1 for specific details.

[0071]

[0072] The original packet contains three layers, namely Ethernet, Probe, and ProbeData layers. The data in Tables 1 - 3 is to transmit INT packets from the host switch to the telemetry server. After receiving the INT packets, the telemetry server parses the telemetry data, enabling it to measure and analyze the packets of all switches on the path. Through this complete packet parsing, the telemetry server can obtain end-to-end delay parameters.

[0073] In the setting of the embodiment of the present invention, the control plane consists of a domain controller and a global controller to form a distributed hierarchical management structure. Each domain is managed and optimized within the domain by a domain controller. Based on the multi-domain distribution characteristics of the above INT data, that is, the controller of each domain can only obtain the INT data within the domain (such as traffic matrix, delay, packet loss rate, and link load, etc.), and cannot directly access the INT data of other domains. The global controller is used for cross-domain coordination and path selection. Multiple domain controllers execute the global routing decisions of the global controller while monitoring the network state.

[0074] By using INT technology to obtain the performance parameters of the intra-domain links in real time, including link delay (end-to-end transmission delay), queue depth (the cache queue status of the switch port), etc., these performance parameters can be aggregated according to a preset time window (such as 1s) to construct a short-term traffic state matrix , whose dimension is . Among them, is the number of switch nodes in the network; is the number of links on each switch node; is the number of performance metrics for each link (such as delay, bandwidth utilization, packet loss rate). Each element of the matrix represents the real-time performance of a certain link, such as link utilization, packet loss rate, and delay, etc. These performance data are collected at a specific moment (such as every second). Therefore, it can reflect the traffic state of each link in the network in a short period.

[0075] After obtaining the short-term traffic state matrix After that, the model parameter configuration stage is entered. First, local agents are deployed for the 3 domain controllers, and each agent is initialized. The neural network of includes the Critic network (i.e., the evaluation network) and the Actor network (i.e., the policy network). The global controller core evaluation network is located on the server and serves as the benchmark reference for the federated model. Set the learning rate of the Critic network to , the learning rate of the Actor network to , and at the same time configure the discount factor , the experience pool capacity M and the training batch size batch size. The Actor network training frequency is , and the target network update frequency is .

[0076] The training process unfolds with the time step as the iteration unit. As shown in Figure 5 , for each agent , the short-term traffic state matrix is input into the Actor network. These data represent the "observations" of the current network environment, reflecting the traffic states of network nodes and each link in the network in a short period of time. The agent can decide the next action based on this information. The state space as a set can be expressed as: . Among them, is the traffic state vector at time , from the switch node to the node , is the bandwidth utilization rate, is the link delay, is the packet loss rate, is the link load (such as queue depth), is the adjacency matrix, including the connection methods of network nodes and the link topology structure.

[0077] After the short-term traffic state matrix is input into the Actor network, the Actor network can output an action according to the current short-term traffic state , that is, select a path from the set of alternative paths (i.e., initial path selection) . Subsequently, the current traffic state and the action Input to the Critic network, and a Q-value is output. This Q-value is used to evaluate the long-term expected return (i.e., expected reward) under a given state and action combination. The estimation of the Q-value can guide the Actor network to adjust the policy so as to select the optimal path and optimize the overall network performance.

[0078] To further improve the decision-making ability of the agent and accelerate convergence, the embodiment of the present invention proposes a reward function. The calculation method of the reward function can be based on real-time INT data to more precisely measure the impact of the current policy on network performance. The agent is in the current traffic state Select an action and then receives an immediate reward feedback from the environment . This immediate reward function combines three key factors: network delay (Delay), packet loss rate (Packet Loss), and resource utilization (Utilization). The specific form of the immediate reward function is as follows:

[0079] Where, is the immediate reward function, , , are coefficients for adjusting the weights of each item, is the network delay under the current traffic state and action , is the packet loss rate under the current traffic state and action , is the resource utilization under the current traffic state and action .

[0080] This immediate reward function not only ensures low latency and low packet loss rate, but also considers multi-modal traffic management. Combining with the bandwidth allocation optimization goal, it makes the resource utilization optimal and avoids over-occupation of a single path.

[0081] During the training process, use the target network to stabilize the training. The target network can be updated regularly to help stabilize the training process. The Critic network and the Actor network can be based on the current traffic state and the corresponding action Adjust the parameters to obtain the optimized Actor network and the optimized Critic network. Based on the optimized Actor network and the optimized Critic network, through the feedback mechanism, the agent continuously adjusts the strategy to make the routing decision achieve multi-objective optimization as much as possible, that is, minimize latency, improve bandwidth utilization, reduce packet loss, etc. That is to say, during the training process, the preset local model can continuously optimize the parameters of the Actor network, and these parameters represent local policy decisions (i.e., local policy parameters), and the finally output local policy parameters can be used as the basic data for subsequent federated learning.

[0082] In addition, in some embodiments, after performing multi-modal analysis on the actual network topology according to the modal types supported by the switch and dividing the actual network topology into logical subnets of different modalities based on the multi-modal analysis results, it further includes: determining the bandwidth resources of each logical subnet based on the service requirements and resource requirements of the logical subnets of different modalities; and performing bandwidth allocation for each logical subnet according to the bandwidth resources of each logical subnet.

[0083] Specifically, each logical subnet of each modality has independent resources (such as bandwidth, computing power, storage, etc.). The ONOS domain controller within the domain schedules the bandwidth resources of the physical network, makes decisions according to the traffic rate of the logical subnet of the current modality in combination with the set control policy, and sends it to the bottom layer to update the bandwidth of each logical subnet, so as to solve the competition problem of resources such as bandwidth.

[0084] The resource allocation model is as follows: The pre-allocated bandwidth according to the statistical demand result of the logical subnet is , and the current logical subnet has a bandwidth of . The resource utilization rate of the logical subnet of the current modality is:

[0085] To ensure the operation of each modality service, when ( represents the upper limit of the resource utilization rate), that is, the logical subnet of this modality is in a congested state, more bandwidth resources can be allocated to it; and when , that is, in a waste state, the bandwidth can be adjusted to reduce the amount of bandwidth resources allocated to the logical subnet of this modality; when ( represents the lower limit of the resource utilization rate), it indicates that the network is in a feasible state and there is no need to adjust the resource allocation.

[0086] When the resource utilization rate is outside the normal range, a resource scheduling request can be sent to the ONOS domain controller within the domain. Considering the number of network slices is , when network resources need to be adjusted, use denotes the adjusted bandwidth, and is used with to denote the total amount, is the traffic of the current subnet, then the resource utilization rate of the logical subnet is:

[0087] The objective function is to maximize the resource utilization rate after bandwidth resource allocation, that is:

[0088] At the same time, the adjustment of network resources will inevitably cause losses, such as the delay caused by updating the logical subnet. At this time, use to represent the adjustment of the logical subnet the cost per unit bandwidth, and the objective function is to minimize the cost of bandwidth adjustment, that is:

[0089] Combining the above two objectives, the final objective function can be obtained as:

[0090] The constraint conditions are:

[0091] It can be seen that through resource partitioning, each logical subnet can be ensured to operate independently and meet its specific requirements, thus providing a basis for subsequent routing optimization and federated learning.

[0092] In summary, the optimization method for multi-modal network global routing proposed in the embodiments of the present invention combines the heterogeneous identification and traffic characteristics in the multi-modal network environment, makes full use of the computing power and data distribution characteristics of local controllers, realizes the efficient training of local policies through distributed reinforcement learning, and realizes the collaborative optimization of global policies through the federated learning mechanism. On the one hand, this method reduces the communication volume between local controllers and global controllers and reduces the risk of privacy leakage; on the other hand, through the distributed aggregation of model parameters and the unified optimization of the global model, it realizes the efficient global routing decision-making of multi-modal networks, overcomes the deficiencies of traditional centralized architectures in terms of efficiency and privacy protection, and provides a brand-new solution for intelligent routing optimization in multi-modal network environments.

[0093] The optimization method for global routing of a multimodal network according to an embodiment of the present invention can construct a data plane and a control plane according to the hierarchical structure of a software-defined network. Then, a multimodal analysis is performed on the actual network topology in the data plane according to the modal types supported by the switches, and the actual network topology is divided into logical subnets of different modalities based on the results of the multimodal analysis. Based on the network topology relationships between the logical subnets of different modalities, a set of N paths that meet the preset conditions between the source switch and the destination switch is generated using a preset path-finding algorithm; based on the set of N paths and the global controller in the control plane, a global routing for the multimodal network is generated using a preset global policy model. Thereby, the problems of insufficient collaboration and low utilization rate of global characteristics in traditional network optimization methods in a multi-domain data environment are solved, and the overall transmission efficiency and cross-domain collaboration ability of the network are significantly improved.

[0094] Next, an optimization device for global routing of a multimodal network according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0095] Figure 6 It is a block diagram of an optimization device for global routing of a multimodal network according to an embodiment of the present invention.

[0096] As Figure 6 shown, the optimization device 10 for global routing of a multimodal network includes: a construction module 100, a division module 200, a first generation module 300, and a second generation module 400.

[0097] Among them, the construction module 100 is used to construct a data plane and a control plane according to the hierarchical structure of a software-defined network, where the data plane includes an actual network topology and a routing environment, and the control plane includes at least one domain controller and a global controller; The division module 200 is used to perform a multimodal analysis on the actual network topology according to the modal types supported by the switches, and divide the actual network topology into logical subnets of different modalities based on the results of the multimodal analysis; The first generation module 300 is used to generate a set of N paths that meet the preset conditions between the source switch and the destination switch using a preset path-finding algorithm based on the network topology relationships between the logical subnets of different modalities, where N is a positive integer; The second generation module 400 is used to generate a global routing for the multimodal network using a preset global policy model based on the set of N paths, the actual network topology, the routing environment, at least one domain controller, and the global controller.

[0098] Furthermore, in some embodiments, the second generation module 400 is specifically used for: Obtain the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information of the multimodal network based on the N path sets, the actual network topology, and the routing environment; Based on the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information, use the global controller and the preset global policy model to generate the global routing of the multimodal network.

[0099] Further, in some embodiments, before generating the global routing of the multimodal network based on the N path sets, the actual network topology, the routing environment, at least one domain controller, and the global controller using the preset global policy model, the second generation module 400 further includes: An obtaining unit, configured to obtain the performance parameters of the intra-domain links based on the N path sets, train a preset local model based on the performance parameters of the intra-domain links, and obtain the local policy parameters of each domain based on the preset local model; A generating unit, configured to fuse the local policy parameters of each domain by using the global controller based on the dynamic weight allocation mechanism, and generate a preset global policy model based on the fused parameters.

[0100] Further, in some embodiments, the obtaining unit is specifically configured to: Based on the N path sets, use in-band network telemetry technology to obtain the performance parameters of the intra-domain links in real time, aggregate the performance parameters according to a preset time window, and construct at least one short-term traffic state matrix; Input each short-term traffic state matrix into the current Actor network to obtain the initial path selection corresponding to each current short-term traffic state; Input each short-term traffic state and the initial path selection corresponding to each current short-term traffic state into the current Critic network to obtain the expected return corresponding to each current short-term traffic state; Input the initial path selection corresponding to each current short-term traffic state and the expected return corresponding to each current short-term traffic state into the target network for training to obtain an optimized Actor network and an optimized Critic network; Use a preset feedback mechanism to obtain a preset local model based on the optimized Actor network and the optimized Critic network.

[0101] Further, in some embodiments, after fusing the local policy parameters of each domain by using the global controller based on the dynamic weight allocation mechanism and generating a preset global policy model based on the fused parameters, the generating unit is further configured to: Update the preset local model based on the model parameters of the preset global policy model.

[0102] Further, in some embodiments, after performing multi-modal analysis on the actual network topology according to the modal types supported by the switch, and dividing the actual network topology into logical subnets of different modalities based on the results of the multi-modal analysis, the dividing module 200 is further configured to: Determine the bandwidth resources of each logical subnet based on the service requirements and resource requirements of the logical subnets of different modalities; Perform bandwidth allocation for each logical subnet according to the bandwidth resources of each logical subnet.

[0103] Further, in some embodiments, the first generating module 300 is specifically configured to: Determine the source node and the destination node; Based on the network topology relationship between the logical subnets of different modalities, use a preset path finding algorithm to calculate the path planning situation from the source node to the destination node; Determine at least one parent node and the length of each path based on each path in the path planning situation; Based on the length of each path, trace back from the destination node in the reverse direction according to at least one parent node of each path until returning to the source node, to obtain an N-path set between the source switch and the destination switch that meets the preset conditions.

[0104] It should be noted that the foregoing explanation of the embodiments of the optimization method for multi-modal network global routing also applies to the optimization device for multi-modal network global routing in this embodiment, and will not be elaborated here.

[0105] According to the optimization device for multi-modal network global routing provided by the embodiments of the present invention, by virtue of the hierarchical structure of the software-defined network, a data plane and a control plane can be constructed. Then, multi-modal analysis is performed on the actual network topology in the data plane according to the modal types supported by the switch, and the actual network topology is divided into logical subnets of different modalities. Based on the network topology relationship between the logical subnets of different modalities, an N-path set between the source switch and the destination switch that meets the preset conditions is generated by using a preset path finding algorithm; based on the N-path set and the global controller of the control plane, a global routing of the multi-modal network is generated by using a preset global policy model. Thereby, the problems of insufficient collaboration and low utilization rate of global characteristics in the traditional network optimization method in a multi-domain data environment are solved, and the overall transmission efficiency and cross-domain collaboration ability of the network are significantly improved.

[0106] Figure 7 The structural schematic diagram of the electronic device provided by the embodiments of the present invention. The electronic device may include: A memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702.

[0107] When the processor 702 executes the program, it implements the optimization method for the multi-modal network global routing provided in the above embodiments.

[0108] Furthermore, the electronic device further includes: A communication interface 703, which is used for communication between the memory 701 and the processor 702.

[0109] A memory 701, which is used to store computer programs that can run on the processor 702.

[0110] The memory 701 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0111] If the memory 701, the processor 702, and the communication interface 703 are implemented independently, the communication interface 703, the memory 701, and the processor 702 can be interconnected through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0112] Optionally, in specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a chip, the memory 701, the processor 702, and the communication interface 703 can communicate with each other through an internal interface.

[0113] The processor 702 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0114] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the optimization method for the multi-modal network global routing as described above.

[0115] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0116] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0117] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. An optimization method for global routing of a multi-modal network, characterized in that It includes the following steps: Construct a data plane and a control plane according to the hierarchical structure of the software-defined network, where the data plane includes an actual network topology and a routing environment, and the control plane includes at least one domain controller and a global controller; Conduct multi-modal analysis on the actual network topology according to the modal types supported by the switches, and divide the actual network topology into logical subnets of different modalities based on the results of the multi-modal analysis; Based on the network topology relationships between the logical subnets of different modalities, use a preset path-finding algorithm to generate a set of N paths that meet the preset conditions between the source switch and the destination switch, where N is a positive integer; Based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller, use a preset global policy model to generate the global routing of the multi-modal network.

2. The method according to claim 1, wherein The step of using a preset global policy model to generate the global routing of the multi-modal network based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller includes: Obtain the boundary nodes, boundary link information, transmission parameters of each domain controller, and routing link information of the multi-modal network based on the set of N paths, the actual network topology, and the routing environment; Based on the boundary nodes, the boundary link information, the transmission parameters of each domain controller, and the routing link information, use the global controller and the preset global policy model to generate the global routing of the multi-modal network.

3. The method according to claim 1, characterized in that Before using a preset global policy model to generate the global routing of the multi-modal network based on the set of N paths, the actual network topology, the routing environment, the at least one domain controller, and the global controller, it further includes: Obtain the performance parameters of the intra-domain links based on the set of N paths, train a preset local model based on the performance parameters of the intra-domain links, and obtain the local policy parameters of each domain based on the preset local model; Based on a dynamic weight allocation mechanism, use the global controller to fuse the local policy parameters of each domain, and generate the preset global policy model based on the fused parameters.

4. The method according to claim 3, wherein The step of obtaining the performance parameters of the intra-domain links based on the set of N paths and training a preset local model based on the performance parameters of the intra-domain links includes: Based on the set of N paths, use in-band network telemetry technology to obtain the performance parameters of the intra-domain links in real time, aggregate the performance parameters according to a preset time window, and construct at least one short-term traffic state matrix; Input each short-term traffic state matrix into the current Actor network to obtain an initial path selection corresponding to each current short-term traffic state; Input each short-term traffic state and the initial path selection corresponding to each current short-term traffic state into the current Critic network to obtain an expected return corresponding to each current short-term traffic state; Input the initial path selection corresponding to each current short-term traffic state and the expected return corresponding to each current short-term traffic state into a target network for training to obtain an optimized Actor network and an optimized Critic network; Using a preset feedback mechanism, obtain the preset local model based on the optimized Actor network and the optimized Critic network.

5. The method according to claim 3, wherein After fusing the local policy parameters of each domain by the global controller based on the dynamic weight allocation mechanism and generating the preset global policy model based on the fused parameters, it further includes: Updating the preset local model based on the model parameters of the preset global policy model.

6. The method according to claim 1, characterized in that, After performing multi-modal analysis on the actual network topology according to the modal types supported by the switch and dividing the actual network topology into logical subnets of different modalities based on the multi-modal analysis results, it further includes: Determine the bandwidth resources of each logical subnet based on the service requirements and resource requirements of the logical subnets of different modalities. Perform bandwidth allocation for each logical subnet according to the bandwidth resources of each logical subnet.

7. The method according to claim 1, characterized in that, The generating, based on the network topology relationship between the logical subnets of different modalities, of an N-path set that meets preset conditions between a source switch and a destination switch by using a preset path finding algorithm includes: Determine the source node and the target node; Based on the network topology relationship between the logical subnets of different modalities, use the preset path finding algorithm to calculate the path planning situation from the source node to the target node; Determine at least one parent node and the length of each path based on each path in the path planning situation; Based on the length of each path, trace back from the target node in the reverse direction according to at least one parent node of each path until returning to the source node to obtain an N-path set that meets preset conditions between the source switch and the destination switch.

8. An optimization device for global routing of a multimodal network, characterized in that, It includes: A construction module, configured to construct a data plane and a control plane according to the hierarchical structure of a software-defined network, where the data plane includes an actual network topology and a routing environment, and the control plane includes at least one domain controller and a global controller; A division module, configured to perform multi-modal analysis on the actual network topology according to the modal types supported by the switch and divide the actual network topology into logical subnets of different modalities based on the multi-modal analysis results; A first generation module, configured to generate an N-path set that meets preset conditions between a source switch and a destination switch by using a preset path finding algorithm based on the network topology relationship between the logical subnets of different modalities, where N is a positive integer; A second generation module, configured to generate a global route of a multi-modal network by using a preset global policy model based on the N-path set, the actual network topology, the routing environment, the at least one domain controller, and the global controller.

9. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the program to implement the optimization method for multimodal network global routing according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the optimization method for multimodal network global routing according to any one of claims 1-7.

Citation Information

Patent Citations

  • Network routing method, system and device, and electronic equipment

    CN113765808A

  • Reinforcement learning agent training method, modal bandwidth resource scheduling method and device

    CN114866494A

  • Resource management method and system suitable for multi-mode network

    CN116319301A

  • Data routing method and device, electronic equipment and storage medium

    CN116566894A

  • Multi-modal network cross-domain networking system and cross-domain interconnected high-speed convergence method thereof

    CN117729069A

Cited By

  • Path optimization and intelligent resource scheduling method and system based on digital twinning

    CN120782229A

  • Forwarding control method and security detection method of multi-mode network

    CN121462496A