Cross-group routing optimization method and device, and storage medium

By adopting the Q-learning algorithm and multi-objective reward function adaptive routing optimization method in high-performance computing networks, the congestion and efficiency of static routing algorithms under dynamic network conditions are solved, and the communication efficiency and resource utilization rate under Dragonfly network architecture are improved.

CN120263716AInactive Publication Date: 2025-07-04ZHEJIANG LAB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510732446.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional static or fixed strategy routing algorithms cannot effectively deal with dynamically changing network states in high-performance computing networks, resulting in network congestion, uneven load and reduced transmission efficiency, especially when cross-group communication under the Dragonfly network architecture.

Method used

The Q-learning algorithm in reinforcement learning is used to combine real-time network status monitoring and multi-objective reward function, and dynamically optimize routing decisions between nodes through local intelligent agents, and balance exploration and utilization of ε-greedy strategy. Multi-objective reward function is designed to reflect path quality, realizing adaptive cross-group routing optimization.

Benefits of technology

It improves the communication efficiency and resource utilization of large-scale GPU clusters under the Dragonfly network architecture, enhances the fault tolerance and robustness of the network, and ensures that the system operates stably under network load fluctuations or local failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263716A_ABST
    Figure CN120263716A_ABST
Patent Text Reader

Abstract

The invention provides a cross-group routing optimization method and device and a storage medium, and the method comprises the steps: obtaining a real-time network state of a current network environment when it is detected that a data packet needs to be transmitted in a cross-group manner through a current node, and determining that the data packet needs to be transmitted in the cross-group manner in the real-time network state; and determining a target node from the next hop node to be selected based on the real-time network state and the long-term return value corresponding to the path from the current node to the next hop node to be selected, thereby realizing real-time sensing of the network state and performing hop-by-hop adaptive routing decision. The problems of congestion, non-uniform load and transmission efficiency reduction when a current static or fixed strategy routing algorithm deals with dynamically changing network conditions are solved. And finally, the data packet is routed from the current node to the target node based on the global near-optimal path. The communication efficiency and the resource utilization rate of the large-scale GPU cluster under the Dragonfly network architecture are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of high-performance computing networks, and particularly to cross-group routing optimization methods, devices, and storage media. Background Art

[0002] In high-performance computing (HPC) and large-scale distributed GPU clusters, the Dragonfly network architecture is widely used due to its high bandwidth, low latency, and strong scalability. This architecture divides nodes into multiple groups, with high interconnectivity within each group, and different groups are connected through a limited number of global links, thus effectively reducing the number of global links and simplifying the network design.

[0003] However, traditional static or fixed-strategy routing algorithms have significant deficiencies in dealing with dynamically changing network states. These algorithms usually rely on pre-configured routing tables or centralized controllers (such as SDN controllers) for path decision-making. When dealing with large-scale parallel tasks, cross-group communication occurs frequently, and some global links or intermediate nodes may become bottlenecks, causing serious network congestion and resource allocation imbalance, thereby affecting the data transmission efficiency of the overall system. Summary of the Invention

[0004] To overcome the problems existing in the related art, this specification provides a cross-group routing optimization method, device, and storage media.

[0005] According to the first aspect of the embodiments of this specification, a method is provided, and the method includes: When detecting that a data packet needs to be transmitted across groups through the current node, obtain the real-time network state of the current network environment, where the real-time network state is determined according to the parameters of the current node and the parameters of the path between the current node and the candidate next-hop node; Determine the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network state, where the long-term reward value is updated by a dynamic learning algorithm to reflect the long-term reward obtained by selecting the corresponding action in the current network state; Based on the real-time network state and the long-term reward value, determine a target node from the candidate next-hop nodes; Route the data packet from the current node to the target node.

[0006] According to a cross-group routing optimization method provided by this specification, the determining a target node from the candidate next-hop nodes based on the real-time network state and the long-term reward value includes: Based on the exploration-exploitation balance strategy, determine a target node from the candidate next-hop nodes based on the real-time network status and the long-term reward value.

[0007] According to a cross-group routing optimization method provided in this specification, the step of determining a target node from the candidate next-hop nodes based on the real-time network status and the long-term reward value according to the exploration-exploitation balance strategy includes: Generate a random number and compare the random number with an exploitation threshold; If the random number does not exceed the exploitation threshold, determine the candidate next-hop node corresponding to the largest long-term reward value as the target node; Wherein, the exploitation threshold represents a critical value at which the strategy tends to select a known optimal action, the exploitation threshold is related to a preset exploration probability, and the value of the exploitation threshold is less than 1.

[0008] According to a cross-group routing optimization method provided in this specification, the method further includes: When the random number is less than the exploitation threshold, randomly determine a node from the candidate next-hop nodes as the target node.

[0009] According to a cross-group routing optimization method provided in this specification, the step of determining the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network status includes: Obtain a state-action value list including states and long-term reward values having a mapping relationship with the states, where the state-action value list records the expected long-term rewards for taking different actions in different network states; From the state-action value list, determine the long-term reward value corresponding to the state that matches the real-time network status as the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network status.

[0010] According to a cross-group routing optimization method provided in this specification, after routing the data packet from the current node to the target node, the method further includes: Obtain an immediate reward and a new network status after completing cross-group transmission of the data packet through the current node; According to the immediate reward and the new network status, update the long-term reward value in the state-action value list through a dynamic learning algorithm, where the state-action value list records the expected long-term rewards for taking different actions in different network states.

[0011] According to a cross-group routing optimization method provided in this specification, obtaining the immediate reward after the current node completes the cross-group transmission of data packets includes: Collecting the parameters of the new network state after the current node completes the cross-group transmission of data packets, where the parameters of the network state include the parameters of the current node and the parameters of the path between the current node and the target node; Based on the parameters of the new network state, obtaining the immediate reward through a preset reward function.

[0012] According to a cross-group routing optimization method provided in this specification, the function terms of the reward function correspond to the parameters of the new network state, and the weight coefficients of the function terms are dynamically adjusted according to the operating state of the network.

[0013] According to a cross-group routing optimization method provided in this specification, the parameters of the real-time network state or the new network state include real-time network state parameters and static topology information.

[0014] According to a cross-group routing optimization method provided in this specification, before obtaining the real-time network state of the current network environment when it is detected that a data packet needs to be cross-group transmitted through the current node, the method further includes: Collecting the real-time network state of the current network environment based on a set time interval.

[0015] According to the second aspect of the embodiments of this specification, a device is provided, including: Including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the cross-group routing optimization method as described in any of the above.

[0016] According to the third aspect of the embodiments of this specification, a computer-readable storage medium is provided, including: When the computer program is executed by the processor, it implements the cross-group routing optimization method as described in any of the above.

[0017] The technical solutions provided by the embodiments of this specification may include the following beneficial effects: In the embodiments of this specification, when it is detected that a data packet needs to be transmitted across groups through the current node, the real-time network status of the current network environment is obtained, and the long-term return value corresponding to the path from the current node to the candidate next-hop node is determined under the real-time network status. Based on the real-time network status and the long-term return value, a target node is determined from the candidate next-hop nodes, so as to realize adaptive routing decision-making by sensing the network status in real time and making hop-by-hop decisions, and solve the problems of congestion, uneven load, and decreased transmission efficiency that occur in current static or fixed policy routing algorithms when dealing with dynamically changing network conditions. Finally, the data packet is routed from the current node to the target node based on the globally near-optimal path, improving the communication efficiency and resource utilization rate of a large-scale GPU cluster under the Dragonfly network architecture.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings

[0019] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0020] Figure 1 is a flowchart of a method shown according to an exemplary embodiment of this specification.

[0021] Figure 2 is a block diagram of a device shown according to an exemplary embodiment of this specification.

[0022] Figure 3 is a schematic diagram of a device shown according to an exemplary embodiment of this specification. Detailed Embodiments

[0023] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.

[0024] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0026] This specification provides a method for cross-group routing optimization, a device, and a computer-readable storage medium. The embodiments of this specification will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners may be combined with each other.

[0027] In high-performance computing (HPC) and large-scale distributed GPU clusters, traditional static or fixed-strategy routing algorithms have significant deficiencies in coping with dynamically changing network states. These algorithms usually rely on pre-configured routing tables or centralized controllers (such as SDN controllers) for path decision-making. When dealing with large-scale parallel tasks, cross-group communication occurs frequently, and some global links or intermediate nodes may become bottlenecks, causing serious network congestion and resource allocation imbalance, thus affecting the data transmission efficiency of the overall system. In addition, traditional methods are also difficult to quickly respond and adjust path strategies when facing local failures or heterogeneous communication patterns (such as point-to-point transmission, broadcast, and AllReduce operations), thus affecting the efficiency and stability of the overall system.

[0028] To solve the above technical problems, this specification provides a method for cross-group routing optimization.

[0029] Aiming to utilize the Q-learning algorithm in reinforcement learning, combined with real-time network state monitoring and multi-objective reward function design, an adaptive and distributed routing optimization mechanism is provided, aiming to solve the problem that traditional static or fixed-strategy routing algorithms cannot effectively cope with dynamically changing network conditions and bursty traffic, thereby improving network resource utilization and data transmission efficiency.

[0030] By such technical means, optimizing the communication path selection between different groups (groups) can maximize bandwidth utilization while maintaining low latency, and enhance the fault tolerance and robustness of the network, ensuring that the system can still operate stably and provide efficient services even when the network load fluctuates or local failures occur.

[0031] It should be noted that the Dragonfly network architecture is widely used due to its characteristics of high bandwidth, low latency, and strong scalability. By dividing nodes into multiple groups, with high interconnectivity within each group and connection between different groups through limited global links, the number of global links is effectively reduced and the network design is simplified. Therefore, the cross-group routing optimization method described in this specification is applied to the Dragonfly network architecture to solve the obvious limitations of the existing cross-group routing mechanism in the Dragonfly architecture when facing a dynamic network environment. However, it should be clear that the cross-group routing optimization method described in this specification is also applicable to other hierarchical high-bandwidth network topologies. Similar topologies can be but are not limited to Fat-Tree (fat tree), Clos network, and the specific implementation is similar to that of the Dragonfly architecture and will not be elaborated here.

[0032] This specification provides an embodiment of a cross-group routing optimization method.

[0033] As Figure 1 shown, Figure 1 is a flowchart of a method shown in this specification according to an exemplary embodiment, including the following steps: In step 102, when it is detected that a data packet needs to be transmitted across groups through the current node, obtain the real-time network state of the current network environment, where the real-time network state is determined according to the parameters of the current node and the parameters of the path between the current node and the candidate next-hop node.

[0034] Taking the cross-group dynamic routing optimization of the Dragonfly architecture as an example, in the Dragonfly architecture, local intelligent agents are deployed on each node in the network. These local intelligent agents collect network state information reflecting the real-time network state and optimize the cross-group routing through a dynamic learning algorithm.

[0035] Among them, the dynamic learning algorithm includes but is not limited to the Q-learning algorithm (Q-learning is a model-free reinforcement learning algorithm, and its core lies in continuously trial and error to update the Q-table and gradually learn which actions to take in different states to obtain the maximum long-term reward).

[0036] As an example, this embodiment constructs a multi-dimensional state space that combines real-time network state parameters and static topology information, enabling each node to dynamically select the optimal path based on the current network environment.

[0037] Specifically, under the Dragonfly architecture, the local intelligent agents deployed on each node in the network collect network status information reflecting the real-time network status. The parameters of the real-time network status or the new network status include real-time network status parameters and static topology information, as shown in Table 1 below: Table 1: Composition of the state space

[0038] For ease of processing in Q-learning, all state variables will be normalized or discretized. For example: The bandwidth utilization rate is divided into three levels: [0 - 30%), [30% - 70%), (70% - 100%]; The queue depth is divided into three levels: low, medium, and high; The latency is divided into three levels: <1ms, 1 - 5ms, >5ms; Other parameters such as physical bandwidth and hop count remain at their original values or are mapped to integers within a finite interval.

[0039] Finally, a discretized high-dimensional state vector is formed from the above parameters to form a state space, which is used as the input for Q-learning to determine the real-time network status through Q-learning. And routing decisions are made through Q-learning, which will be described in detail in the following chapters.

[0040] In some other embodiments, before obtaining the real-time network status of the current network environment when detecting that a data packet needs to be transmitted across groups through the current node, the method further includes: Collect the real-time network status of the current network environment based on a set time interval.

[0041] The local agent periodically collects the above parameters from the network interface and the routing table. For example, according to a specific time interval (such as every microsecond or before each routing decision), it collects network status information to update the state space, ensuring that the Q-learning model always makes decisions based on the latest network status.

[0042] Through this embodiment, by real-time perceiving the changes in key metrics such as network load, queue depth, and link bandwidth utilization rate, real-time perceiving the network status and making adaptive routing decisions, it solves the problems of congestion, uneven load, and decreased transmission efficiency that occur in existing static or fixed-strategy routing algorithms when dealing with dynamically changing network conditions, thereby improving the communication efficiency and resource utilization rate of large-scale GPU clusters under the Dragonfly network architecture.

[0043] In step 104, determine the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network state, where the long-term reward value is updated by a dynamic learning algorithm to reflect the long-term reward obtained by selecting the corresponding action in the current network state.

[0044] The core of Q-learning is to update the state-action value list (Q-table) through continuous trial and error, gradually learning which action to take in different states to obtain the maximum long-term reward, and gradually learning the optimal routing strategy. The Q-table records the expected long-term reward values (Q-values) of taking different actions in different states. In subsequent routing decisions, the Q-values in the Q-table will be used as an important basis for selecting the next-hop node. It can be understood that the Q-value represents the expectation of the long-term cumulative reward that can be obtained by selecting a specific action in a certain state. A higher Q-value indicates that selecting the corresponding action in this state may obtain a better long-term reward, so it is more likely to be selected. As the Q-table is continuously updated, the nodes in the network can gradually find the optimal routing selection in different network states, realizing dynamic routing optimization.

[0045] In some embodiments, in step 104 of determining the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network state, it includes the following steps 1041 to 1042: In step 1041, obtain a state-action value list including the long-term reward values having a mapping relationship with the state, where the state-action value list records the expected long-term rewards of taking different actions in different network states; In step 1042, determine, from the state-action value list, the long-term reward value corresponding to the state that matches the real-time network state, as the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network state.

[0046] The candidate next-hop node described in this specification refers to a certain node in the action space, and the action space represents the set of next-hop nodes to which the current node can forward data packets. In the Dragonfly architecture, the action space includes, but is not limited to, neighbor nodes within the same group; global link exit nodes connecting other groups; if the current node is the destination node, the action is empty (terminal state).

[0047] In this specification, there is a mapping relationship between the state in the Q-table and the long-term reward value.

[0048] When determining the optimal path data transmission path through the Q value, based on the information in the state space, the Q values of each action in the corresponding state are looked up in the Q-table. That is, at a certain moment, the state space describes information such as the current network bandwidth utilization and queue depth. Q-learning looks up the target state that matches the real-time network state of the current node from the Q-table based on this information. Then, the action corresponding to the Q value mapped by the target state is more likely to be selected. The action refers to the action of transmitting data from the current node to the path of the candidate next-hop node.

[0049] It should be noted that in the initialization stage, the Q values of all state-action pairs in the Q-table are set to 0 or a random small value. During the training process, the standard Q-learning update formula is used to update the Q-table. As the Q-table is continuously updated, the nodes in the network can gradually find the optimal routing selection in different network states and achieve dynamic routing optimization. The update of this part will be described in detail in the subsequent chapter.

[0050] In step 106, based on the real-time network state and the long-term reward value, a target node is determined from the candidate next-hop nodes.

[0051] In some embodiments, according to the exploration-exploitation balance strategy, based on the real-time network state and the long-term reward value, a target node is determined from the candidate next-hop nodes.

[0052] The exploration-exploitation balance strategy described in this specification includes but is not limited to the ε-greedy strategy. In the cross-group dynamic routing optimization of the Dragonfly architecture, the network state is complex and changeable. The ε-greedy strategy helps the Q-learning decision engine balance exploration and exploitation and make a trade-off between the known optimal path and the possibly better path.

[0053] As an example, in step 106 of determining the target node from the candidate next-hop nodes according to the exploration-exploitation balance strategy, based on the real-time network state and the long-term reward value, it includes steps 1061 to 1063: In step 1061, a random number is generated and compared with the exploitation threshold. In step 1062, if the random number does not exceed the exploitation threshold, the candidate next-hop node corresponding to the maximum long-term reward value is determined as the target node. In step 1063, when the random number is less than the exploitation threshold, a node is randomly determined from the candidate next-hop nodes as the target node. Among them, the exploitation threshold represents a critical value that tends to select the known optimal action. The exploitation threshold is related to a preset exploration probability and the value of the exploitation threshold is less than 1.

[0054] The ε-greedy policy is a mechanism for balancing exploration and exploitation in the Q-learning decision-making process, but it is not the only basis for decision-making. It selects the action with the largest current Q value with a probability of (1 - ε) of exploitation, that is, based on the current available information (such as real-time parameters like current bandwidth utilization, queue depth, end-to-end delay, etc. and static parameters like hop count, physical link bandwidth, etc.), it selects the currently considered optimal route to transmit data; it randomly selects an action with a probability of ε of exploration to explore new paths, that is, it actively tries some routes that may not seem optimal currently, aiming to obtain more information about the network and discover potential better routes that have not been discovered in the current state. It should be noted that exploration is not a random attempt, but a purposeful attempt of different paths based on a certain strategy (such as the ε probability in the ε-greedy policy).

[0055] That is to say, it selects the action corresponding to the largest Q value in the current Q-table with a probability of (1 - ε) of exploitation, and selects the action with the largest current Q value.

[0056] It randomly selects an action with a probability of ε of exploration. Based on the set of neighbor nodes of the current node or the set of optional next-hop nodes (in the Dragonfly architecture, the action space of a node includes neighbor nodes within the same group, global link exit nodes connecting other groups, etc.), it randomly selects a node as the next-hop node. After selection, the data packet will be sent to this node, and then the Q-table will be updated according to the result of this transmission (calculating the reward value through the reward function), so as to accumulate experience about this path for better subsequent routing decisions and avoid falling into local optima.

[0057] Exemplarily, in a network, when ε is set to 0.1, the action with the largest Q value will be selected 90% of the time, and an action will be randomly selected 10% of the time.

[0058] It should be noted that in the initial stage, due to limited understanding of the network, a larger ε value can make the node explore new paths more frequently, obtain more network information, and discover potential better routes; as learning progresses, gradually reduce the ε value and increase the probability of using existing experience (selecting the action with the largest Q value) to make the routing decision more stable and efficient. For example, in the network initialization stage, ε is set to 0.5, and the node has a 50% probability of randomly exploring new paths; after a period of learning, ε is adjusted to 0.1. At this time, the node mainly uses the existing Q values to select routes, but will still occasionally explore new paths to adapt to network changes.

[0059] In this embodiment, the algorithm balances between leveraging existing knowledge and exploring new possibilities, determines the next-hop node according to the real-time network state and action selection strategy of the current node, and improves the accuracy of cross-group routing optimization.

[0060] In step 108, route the data packet from the current node to the target node.

[0061] The data packet is forwarded to the target node along the selected path. During the routing process, monitor the path performance metrics, such as the transmission result: whether it reaches successfully, end-to-end delay, change in queue depth. And network state changes: bandwidth utilization, link load balancing. Obtain the observation result after the path execution through the information in the monitoring process.

[0062] In some embodiments, after routing the data packet from the current node to the target node, the method further includes steps 110 to 112: In step 110, obtain the immediate reward and the new network state after completing the cross-group transmission of the data packet through the current node; In step 112, according to the immediate reward and the new network state, update the long-term return value in the state-action value list through a dynamic learning algorithm, and the state-action value list records the expected long-term return of taking different actions in different network states.

[0063] After each data packet transmission is completed, calculate the immediate reward, and use this reward value to update the Q value of the corresponding state-action pair in the Q-table according to the Q-learning update formula.

[0064] The obtaining of the immediate reward after completing the cross-group transmission of the data packet through the current node includes: Collect the parameters of the new network state after completing the cross-group transmission of the data packet through the current node, and the parameters of the network state include the parameters of the current node and the parameters of the path between the current node and the target node; Based on the parameters of the new network state, obtain the immediate reward through a preset reward function.

[0065] As an example, the immediate reward is obtained through a multi-objective reward function, and the design of the reward function determines the learning direction of the Q-learning model. This specification proposes a multi-objective weighted reward function, comprehensively considering multiple network performance metrics (such as the end-to-end delay, available bandwidth, queue depth, and transmission success rate of this transmission, etc.), balancing the importance of each metric through weight coefficients, and guiding the Q-learning model to learn the end-to-end performance optimal path, so as to guide the model to learn an end-to-end performance optimal path.

[0066] Exemplarily, the reward function R is defined as follows:

[0067] Wherein, w 1. w 2. w 3. w 4 is the weight coefficient, which can be adjusted according to the specific application scenario and is used to balance the importance of each index; : The end-to-end delay of this transmission; : The available bandwidth of the current link; : The queue length of the next-hop node; : The transmission success rate of the path closest to this path.

[0068] In some embodiments, the function terms of the reward function correspond to the parameters of the new network state, and the weight coefficients of the function terms are preset according to the application scenario or dynamically adjusted according to the network condition.

[0069] In practical applications, according to different scenarios, the emphasis on network performance indicators is different, and the weight coefficients are preset w 1. w 2. w 3. w 4. For example, in the scenario of deep learning model training, the timeliness of data transmission is crucial, and the end-to-end delay will seriously affect the training efficiency. Therefore, the value of w 1 (the weight related to delay) can be appropriately increased, and the other weights can be relatively reduced to highlight the importance of reducing delay and guide the Q-learning model to preferentially select the path with low delay. In the scientific simulation scenario, the data volume is huge, and more attention is paid to the full utilization of the link bandwidth. Then the value of w 2 (the weight related to the available bandwidth) can be increased, so that the model tends to select the path with large bandwidth.

[0070] Based on the parameters of the new network state, the immediate reward is calculated through the above reward function.

[0071] In some other embodiments, the weight coefficients of the function terms are dynamically adjusted according to the network operating state (such as high load or stable communication phase). The adjusted weights are used for subsequent reward calculation to guide the Q-learning model to adapt to different network scenarios and optimize the routing decision.

[0072] As an example, in the scenario where the network is under high load, the congestion and delay problems are prominent. Increase w1 (weight related to latency, preferring paths with low latency) and w 3 (weight related to queue depth, reducing data waiting time at nodes), emphasizing congestion and latency reduction; In the stable communication phase, link utilization and transmission reliability are more critical. Improve w 2 (weight related to available bandwidth) and w 4 (weight related to transmission success rate), improving link utilization and reliability.

[0073] During network operation, by continuously monitoring network performance metrics and routing decision effects, the weights are iteratively optimized according to certain algorithms or rules. For example, record the network performance changes after each routing decision. If it is found that the path selected according to the current weights always results in too high latency, then appropriately increase w the value of 1; if the link utilization is low, then appropriately increase w the value of 2. Through multiple iterations, find the weight combination that best suits the current network state and application requirements, and achieve a balance between different goals.

[0074] After each data packet transmission is completed, according to the standard Q-learning update formula, combined with the current actual network state, selected action, immediate reward, new network state, and learning rate and discount factor to update the Q value in the Q-table, As an example, during the training process, the standard Q-learning update formula is as follows: +

[0075] where, : the current real-time network state: : the currently selected action (next-hop node); : the immediate reward obtained after executing the action; : the new network state; : the learning rate (0 ≤ ≤ 1), controlling the update step size; : the discount factor (0 ≤ ≤ 1), indicating the degree of emphasis on future rewards.

[0076] In subsequent routing decisions, the Q-values in the Q-table will serve as an important basis for selecting the next-hop node. A higher Q-value indicates that choosing the corresponding action in that state may result in a better long-term reward, and thus is more likely to be selected. As the Q-table is continuously updated, the nodes in the network can gradually find the optimal routing choices in different network states, achieving dynamic routing optimization.

[0077] For example, if after choosing node B as the next-hop, the data packet is successfully transmitted with a low transmission delay and high bandwidth utilization, obtaining a high immediate reward, then the Q-value corresponding to the current state and the action of choosing node B will be updated to reflect this successful experience.

[0078] This embodiment is applicable to application scenarios that need to process a large number of parallel computing tasks, such as high-performance computing (HPC) environments like deep learning model training, scientific simulations, and big data analysis. By optimizing the communication path selection between different groups (groups), the present invention can maximize the bandwidth utilization while maintaining a low latency, and enhance the fault tolerance and robustness of the network, ensuring that the system can still operate stably and provide efficient services even in the case of network load fluctuations or local failures.

[0079] This specification provides a cross-group routing optimization method, device, and computer-readable storage medium. When it is detected that a data packet needs to be transmitted across groups through the current node, the real-time network state of the current network environment is obtained. By constructing a state space that includes real-time parameters such as bandwidth utilization, queue depth, end-to-end delay, and static parameters such as hop count and physical link bandwidth, and combining with the local intelligent agents deployed on each node, an adaptive routing decision is made hop by hop. Determine the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network state. Based on the real-time network state and the long-term reward value, a target node is determined from the candidate next-hop nodes; among them, the ε-greedy strategy is adopted to balance exploration and exploitation, and a multi-objective reward function is designed to reflect the path quality, thereby driving the Q-learning model to gradually converge to a globally near-optimal path. Subsequently, the data packet is routed from the current node to the target node based on the globally near-optimal path. Throughout the process, the system supports microsecond-level path switching, avoiding the high latency problem brought by centralized control, and significantly improving the network robustness and resource utilization in burst traffic scenarios. It effectively solves the path selection problem of cross-group communication in a large-scale GPU cluster under the Dragonfly topology structure, improves the overall communication efficiency and network throughput capacity, and is applicable to high-performance computing scenarios such as artificial intelligence training and scientific computing.

[0080] As Figure 2 shown, Figure 2The block diagram of a device shown in this specification according to an exemplary embodiment, the device includes: A status acquisition module, configured to obtain the real-time network status of the current network environment when it is detected that a data packet needs to be transmitted across groups through the current node, where the real-time network status is determined according to the parameters of the current node and the parameters of the path between the current node and the candidate next-hop node; A first routing optimization module, configured to determine the long-term return value corresponding to the path from the current node to the candidate next-hop node in the real-time network status, where the long-term return value is updated by a dynamic learning algorithm to reflect the long-term return obtained by selecting the corresponding action in the current network status; A second routing optimization module, configured to determine a target node from the candidate next-hop nodes based on the real-time network status and the long-term return value; A data transmission module, configured to route the data packet from the current node to the target node.

[0081] For the functions and effects of each module / sub-module / unit in the above device, the specific implementation process can be found in the corresponding steps of the above method, and the same technical effects can be achieved, which will not be elaborated here.

[0082] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0083] Figure 3 Illustrates the schematic physical structure diagram of a cross-group routing optimization device, as Figure 3 shown, the cross-group routing optimization device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the cross-group routing optimization method.

[0084] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0085] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the cross-group routing optimization method provided by the above-mentioned various methods.

[0086] On another aspect, this application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the cross-group routing optimization method provided by the above-mentioned various methods.

[0087] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0088] Those skilled in the art will readily think of other implementations of this specification after considering the specification and practicing the invention herein. This specification is intended to cover any variations, uses, or adaptations of this specification, which follow the general principles of this specification and include the common general knowledge or conventional technical means in the technical field not claimed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this specification are pointed out by the following claims.

[0089] It should be understood that the present specification is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present specification is only limited by the appended claims.

[0090] The above are only the preferred embodiments of the present specification and are not intended to limit the present specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present specification shall be included within the scope of protection of the present specification.

Claims

1. A cross-group routing optimization method, characterized in that The method includes: When it is detected that a data packet needs to be transmitted across groups through the current node, obtaining the real-time network status of the current network environment, where the real-time network status is determined according to the parameters of the current node and the parameters of the path between the current node and the candidate next-hop node; Determining the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network status, where the long-term reward value is updated by a dynamic learning algorithm to reflect the long-term reward obtained by selecting the corresponding action in the current network status; Based on the real-time network status and the long-term reward value, determining a target node from the candidate next-hop nodes; Routing the data packet from the current node to the target node.

2. The cross-group routing optimization method according to claim 1, wherein The determining a target node from the candidate next-hop nodes based on the real-time network status and the long-term reward value includes: Determining a target node from the candidate next-hop nodes according to the exploration-exploitation balance strategy based on the real-time network status and the long-term reward value.

3. The cross-group routing optimization method according to claim 2, wherein The determining a target node from the candidate next-hop nodes according to the exploration-exploitation balance strategy based on the real-time network status and the long-term reward value includes: Generating a random number and comparing the random number with an exploitation threshold; If the random number does not exceed the exploitation threshold, determining the candidate next-hop node corresponding to the largest long-term reward value as the target node; Wherein, the exploitation threshold represents the critical value at which the strategy tends to select the known optimal action, the exploitation threshold is related to the preset exploration probability and the value of the exploitation threshold is less than 1.

4. The cross-group routing optimization method according to claim 3, wherein The method further includes: When the random number is less than the exploitation threshold, randomly determining a node from the candidate next-hop nodes as the target node.

5. The cross-group routing optimization method according to claim 1, wherein, The determining the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network status includes: Obtaining a state-action value list including the long-term reward value having a mapping relationship with the state, where the state-action value list records the expected long-term rewards for taking different actions in different network statuses; Determining, from the state-action value list, the long-term reward value corresponding to the state matching the real-time network status as the long-term reward value corresponding to the path from the current node to the candidate next-hop node in the real-time network status.

6. The cross-group routing optimization method according to claim 1, wherein After routing the data packet from the current node to the target node, the method further includes: Obtaining the immediate reward and the new network status after the data packet is transmitted across groups through the current node; Updating the long-term reward value in the state-action value list according to the immediate reward and the new network status through a dynamic learning algorithm, where the state-action value list records the expected long-term rewards for taking different actions in different network statuses.

7. The cross-group routing optimization method according to claim 6, characterized in that, The obtaining the immediate reward after the data packet is transmitted across groups through the current node includes: Collect the parameters of the new network state after the current node completes the cross-group transmission of data packets. The parameters of the network state include the parameters of the current node and the parameters of the path between the current node and the target node. Based on the parameters of the new network state, obtain the immediate reward through a preset reward function.

8. The cross-group routing optimization method according to claim 7, characterized in that, The function terms of the reward function correspond to the parameters of the new network state, and the weight coefficients of the function terms are dynamically adjusted according to the network operation state.

9. The cross-group routing optimization method according to claim 8, wherein, The parameters of the real-time network state or the new network state include real-time network state parameters and static topology information.

10. The cross-group routing optimization method according to claim 1, characterized in that Before detecting that it is necessary to cross-group transmit data packets through the current node and obtaining the real-time network state of the current network environment, the method further includes: Collect the real-time network state of the current network environment based on a set time interval.

11. A computer device, characterized in that, It includes a memory, a processor, and a cross-group routing optimization program stored on the memory and executable on the processor. When the processor executes the cross-group routing optimization program, it implements the steps of the cross-group routing optimization method according to any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, A cross-group routing optimization program is stored on the computer-readable storage medium. When the cross-group routing optimization program is executed, it implements the steps of the cross-group routing optimization method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Mobile ad hoc network routing optimization method and device based on deep reinforcement learning

    CN117061411A

  • Cloud adaptive congestion control algorithm based on deep reinforcement learning

    CN118802758A

  • Route scheduling method based on deep reinforcement learning

    CN118945100A

  • Multi-path routing method and device, medium and equipment

    CN119544597A

  • SDN-based air-sea cross-domain network reinforcement learning routing algorithm

    CN119996290A