A method for digital twin-assisted data center network load balancing
By building a digital twin framework in the data center network, collecting and computing network data, and using the DDPG algorithm to make routing decisions, the problem of insufficient consideration of algorithm model training dependency and dynamics in load balancing in data center network is solved, and the autonomous and efficient traffic scheduling of the network is realized.
Patent Information
- Application Number
- CN202211740947.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In data center network load balancing based on reinforcement learning, algorithm model training relies on data samples simulated by digital twin networks, lacks authenticity, fails to consider the dynamics of the network, and untrained strategies sent to the real network will cause performance deterioration.
By building a data center network framework based on digital twins, data from the physical data center network is collected, link utilization and delay are calculated, new flow size is judged, and routing decisions are made using ECMP or elephant flow scheduling module, link weight calculation is used to calculate and stream table is issued.
It realizes load balancing of the data center network, enhances the SDN function, realizes the autonomy of the data center network, and helps the data center to perform traffic scheduling in an intelligent and efficient way.
Smart Images

Figure CN116233133B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of digital twins and relates to a method for digital twin-assisted data center network load balancing. Background Art
[0002] As the infrastructure of the next generation of Internet application services, data centers have attracted widespread attention from industry and academia and have become a hot research area. Load balancing is one of the research issues in data center network traffic engineering. Its purpose is to distribute traffic in multiple equal-cost multi-paths as much as possible, make full use of network resources, and avoid high-load traffic causing network congestion. In recent years, the emergence of digital twin technology can build a real-time virtual network for the physical network and enhance the simulation, optimization, and verification capabilities that the physical network lacks, which makes up for the shortcomings of data center network traffic scheduling based on SDN. Therefore, artificial intelligence algorithm-enabled digital twins can globally and accurately grasp the status of the data center network. After simulation, optimization, verification, and control of the virtual network, load balancing of the data center network can be achieved.
[0003] At present, the main purpose of introducing digital twin technology in the network field is to use digital twins as simulators to generate training samples for artificial intelligence algorithms or estimate the rewards of reinforcement learning. However, the introduction of digital twins in data center network load balancing based on reinforcement learning faces the following problems: first, the algorithm model training completely relies on the data samples simulated by the digital twin network, that is, the model data samples lack authenticity; second, the network state has changed before a policy is executed, that is, the dynamic nature of the network is not considered; finally, the untrained policy sent to the real network will cause network performance to deteriorate, that is, the security of the network policy is not verified. Summary of the invention
[0004] In view of this, an object of the present invention is to provide a method for digital twin-assisted data center network load balancing.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for digital twin-assisted data center network load balancing includes the following steps:
[0007] S1: Build a data center network framework based on digital twins, including the physical data center network layer, twin network layer, and network application layer;
[0008] S2: Collect data on the physical data center network, including the configuration of servers, switches, and links, the status of servers, switches, and links, the network topology data structure, and the traffic matrix;
[0009] S3: Calculate link utilization, latency, packet loss rate, switch load, and traffic matrix;
[0010] S4: In the current collection cycle, determine the size of the new flow arriving at the edge switch. If it is a mouse flow, use the ECMP method for routing; if it is an elephant flow, transfer it to the elephant flow scheduling module for scheduling;
[0011] S5: The Elephant Flow Scheduling Module uses DDPG to make scheduling decisions. Its output is a set of link weights. The path calculation submodule calculates the optimal forwarding path, and then the flow table is sent to the physical data center network through the traffic management module.
[0012] S6: Through the southbound interface protocol, the physical data center network layer receives the flow table and selects the optimal path for rerouting the large flow.
[0013] Further, the construction of a data center network framework based on digital twins in step S1 specifically includes the following steps:
[0014] S11: Use Mininet simulation to build a physical data center network layer, where the physical data center network layer is a Fat-tree network topology structure;
[0015] S12: Constructing a twin network layer, including a data storage module, a model mapping module and a twin management module; the data storage module stores various types of collected data in a Redis database through a controller; the model mapping module maps the collected data into a functional model and a basic model, and operates through the twin management module;
[0016] S13: Build a network application layer, obtain valuable information from the twin network layer database through the northbound interface protocol, and realize network visualization and network logging.
[0017] Furthermore, the physical data center network layer is composed of network elements, servers, and links. In the data center network, the network elements are switches, routers, and equipped applications, which are used to support the filtering and forwarding of network data packets.
[0018] Furthermore, the digital twin network layer is used to establish a real-time mapping model between the physical data center network and the virtual twin network system, collect various configuration and operation data of the physical network entity through the southbound interface protocol, and store it in the database for the establishment, simulation and optimization of the basic model and functional model;
[0019] The data storage module collects data from the physical data center network and divides the data into four different types for storage: the configuration of servers, switches, and links; the status of servers, switches, and links; the network topology data structure; and the traffic matrix;
[0020] The model mapping module extracts, defines and describes the key features of each entity in the physical network based on the collected data, builds the network element model and topology model of the data center network, and provides training data and test data for the input of the functional model;
[0021] The network element model is a real-time and accurate mapping of servers, switches, and links. The topology model is obtained by connecting and combining the network element models according to the network topology data structure.
[0022] The functional model is to establish a network simulation, analysis, and optimization data model based on the network data in the database;
[0023] The twin management module is responsible for managing and updating each mapping model in the network twin layer, and has the functions of model update, status synchronization, model interaction, and application association.
[0024] Furthermore, the network application layer supports the following network applications based on the physical data center network and the digital twin network: QoS configuration, network logging, network visualization, and policy verification;
[0025] The network application layer inputs QoS requirements and network policies to the twin network layer through the northbound interface protocol, and simulates the specified network policy in the twin network layer through the model instance; if the designed policy can meet the QoS requirements, the policy is deployed to the real network environment through the southbound interface protocol.
[0026] Furthermore, for the physical data center network DCN<V,E> , use e i,j Indicates that from v i to v j The maximum load of the link is B i,j , the actual load is b i,j At time t, suppose there are k flows in the network, and the set of flows is represented by F = {f1, f2, ..., f k}, one of the flows f s In link e i,j The occupied bandwidth on Define f s The quaternary group is (o s ,d s ,r s ,t s ), s∈{1,2,…,k}, where o s ∈V represents the source node; d s ∈V represents the destination node; r s Represents the flow f s Required bandwidth; t s Indicates the size of the flow, and its value equal to 1 indicates f s For large flow, a value equal to 0 means f s For small streams;
[0027] Initially, the KSP algorithm calculates the path set between all edge node pairs The lth path has Links, is the source node v o To the destination node v d Directed sequence of <v o →v i →v j →…→v d connected; path load relative balance CV l :
[0028]
[0029]
[0030]
[0031] st
[0032]
[0033]
[0034] In the formula, represents the average link load of the lth path, σ represents the standard deviation of the link load of the lth path; Formula (4) is the link load constraint, which ensures that the traffic distributed on the link does not exceed the maximum load of the link; Formula (5) is the traffic conservation constraint, which ensures that the amount of data flowing out of the source destination node is equal to the amount of data flowing into the destination node, where Γ(v i )={v j :e i,j ∈E},Γ′(v i )={v j :e j,i ∈E};
[0035] Further, in step S3, the controller sends a directional message to the specified switch to obtain link delay information and calculate link e i,j The specific steps are as follows:
[0036] S31: The controller sends a Packet_out message carrying the current timestamp to the switch v i , which indicates that the switch v i Send it to the adjacent switch v j ;
[0037] S32: Switch v j Receive switch v iIf the sent message cannot match the flow table entry, it will be forwarded to the controller as a Packet_in message. After receiving the message, the controller unpacks it to obtain the timestamp, and then subtracts the timestamp from the current timestamp to obtain the time difference ΔT1. The time from the controller sending the message to receiving the message is:
[0038]
[0039] S33: The controller sends a Packet_out message carrying a timestamp to the switch v j , and then the controller receives the i Receive the Packet_in message and record the time difference ΔT2:
[0040]
[0041] S34: The controller sends a message to the switch v i and switch v j Send an Echo request message with a timestamp. After receiving the message, the switch replies with an Echo reply message. After receiving the Echo reply message, the controller calculates the timestamp to the switch v i and switch v j The round trip time is thus and Switch v i To switch v j The round trip time is:
[0042]
[0043] S35: Assume that switch v i With switch v j The round-trip time of link e is the same, and the processing time of the switch is ignored. i,j The delay is:
[0044]
[0045] When calculating the link delay, if the link is severely congested and the controller in step S32 cannot receive the message, the currently calculated link delay is set to a high value. i,j The delay is:
[0046]
[0047] In summary, the delay of the lth path is expressed as:
[0048]
[0049] Taking into account the relative balance of path load and path delay, the established objective function is as follows:
[0050]
[0051] Among them, α∈(0,1).
[0052] Furthermore, the mathematical model of the digital twin network is:
[0053] In the physical data center network, there are n network element devices, and V is defined as {v1, v2, …, v n} is a set of network element devices, with m links, and E = {e1, e2, …, e m} is a link set, the digital twins of the network element model and link model are expressed as:
[0054] DT v (t) = Θ(C v ,S v (t),M v (t)) (32)
[0055] DT e (t) = Θ(C e ,S e (t)) (33)
[0056] Where C represents static configuration data. For network element device v, it can be the maximum transmission rate, backplane bandwidth, and MAC address capacity; for link e, it can be the maximum capacity; S(t) represents the operation status that changes over time, which is determined by multi-dimensional features. For network element device v, its x-dimensional state feature is defined as M v (t) represents the operation behavior of network element device v, which is characterized by the y-dimensional characteristic behavior and is defined as
[0057] The topology model is obtained by connecting and combining the network element model and the link model according to the network topology data structure. The digital twin of the topology model is represented as follows:
[0058] DT(t)=Θ(DT V (t),DT E (t)) (34)
[0059] Among them, DT V (t) represents the set of all network element models in the physical data center network, that is, DT E (t) represents the set of all link models in the physical data center network, namely
[0060] The relationship between the physical data center network and the digital twin network is formally defined as:
[0061]
[0062] Among them, DCN<V,E> represents the physical data center network, V represents the set of all network nodes, and E represents the set of all links in the network; SIP represents the southbound interface protocol, through which communication between the physical data center network and the digital twin network is realized.
[0063] Furthermore, digital twin-assisted data center network load balancing is implemented based on the DDPG algorithm. The four networks of the DDPG algorithm are:
[0064] Actor current network: responsible for iterative update of policy network parameters θ, responsible for selecting the current action A according to the current state S, used to interact with the environment to generate S′, R;
[0065] Actor target network: responsible for selecting the optimal next action A′ based on the next state S′ sampled in the experience replay pool, and the network parameters θ′ are regularly copied from θ;
[0066] Critic current network: Iterative update of the replication value network parameter ω, responsible for calculating the current Q value Q(S,A,ω), the target Q value y i =R+γQ′(S′,A′,ω′);
[0067] Critic target network: responsible for calculating the Q′(S′,A′,ω′) part of the target Q value, and the network parameter ω′ is regularly copied from ω;
[0068] The scheduling method of the elephant flow scheduling module in step S4 is as follows:
[0069] State: It is an n×n traffic matrix TM, where b 1,n is switch v1 and switch v n The actual load of the connected links is expressed as follows:
[0070]
[0071] Action: The action space is a set of link weights, which are expressed as follows:
[0072] W=[w1,w2,…,w m ] T (37)
[0073] After obtaining a set of link weights, the routing path of the new flow is obtained through the path calculation module, and then the flow table entry is obtained through the flow table management module. The flow table is sent to the twin network, and then sent to the physical network after verification, finally realizing the path selection and forwarding of the large flow;
[0074] Reward: Based on the current state and action, the agent receives a reward from the environment; in the data center network, the reward function is as follows:
[0075] R(s(t),a(t))=-(min α·CV l +(1-α)·T l ) (38).
[0076] The beneficial effects of the present invention are: digital twin network technology can enhance SDN functions and realize the autonomy of data center networks. The present invention completes the load balancing of data center networks and helps data centers achieve network autonomy in a more intelligent and efficient traffic scheduling manner.
[0077] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0079] Figure 1 This is a network framework diagram for data based on digital twins;
[0080] Figure 2 This is a flowchart of digital twin-assisted data center network load balancing. DETAILED DESCRIPTION
[0081] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and the features in the embodiments can be combined with each other without conflict.
[0082] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0083] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0084] The present invention provides a method for digital twin-assisted data center network load balancing, which is mainly divided into two parts: one part is to design a data center network framework based on digital twins, and the other part is to realize data center network load balancing under the network framework. Figure 1 As shown in Figure 1, it is divided into three layers: network application layer, twin network layer and physical data network layer. The flowchart of digital twin-assisted data center network load balancing is shown in Figure 1. Figure 2 shown.
[0085] A method for digital twin-assisted data center network load balancing mainly includes the following steps:
[0086] Step 1: Build the physical data center network layer. The physical data center network layer adopts the Fat-tree network topology structure and uses Mininet for simulation construction to make it act as a real physical network environment.
[0087] Step 2: Build the twin network layer. The twin network layer is divided into three modules: data storage, model mapping, and twin management. The data storage module stores the collected data in the Redis database through the controller, which provides data support for the model mapping module and the twin management module. The model mapping module is divided into two sub-modules: the functional model and the basic model. The twin management module can make the model operational.
[0088] Step 3: Build the network application layer. The example of the present invention implements functions such as network visualization and network log in the network application layer. By obtaining valuable information in the twin network layer database through the northbound interface protocol, convenient application services can be provided for network administrators.
[0089] Step 4: Through steps 1, 2, and 3, build a data center network framework based on digital twins. The framework diagram is as follows: Figure 1 shown.
[0090] Step 5: Initially, the controller collects data from the physical data center network, including the configuration of servers, switches, and links, the status of servers, switches, and links, the network topology data structure, and the traffic matrix.
[0091] Step 6: Calculate link utilization, latency, packet loss rate, switch load, and traffic matrix.
[0092] Step 7: During the current collection cycle, determine the size of the new flow arriving at the edge switch. If it is a mouse flow, use the ECMP method for routing; if it is an elephant flow, transfer it to the elephant flow scheduling module for scheduling.
[0093] Step 8: The Elephant Flow Scheduling Module uses DDPG to make scheduling decisions. Its output is a set of link weights. The optimal forwarding path is calculated through the path calculation submodule, and then the flow table is sent to the physical data center network through the traffic management module.
[0094] Step 9: Through the southbound interface protocol, the physical data center network layer receives the flow table and selects the optimal path to reroute the large flow.
[0095] Specifically, the method also includes the following contents:
[0096] Designing a data center network framework based on digital twins
[0097] The framework is divided into three layers: the physical data center network layer, the twin network layer, and the network application layer. The physical data center network layer provides the twin network layer with basic network equipment configuration, environmental information, operating state, link topology and other information; the twin network layer is responsible for the construction and operation of the basic model and functional model of the physical network entity; the network application layer can be a platform for integrating services and applications, aiming to achieve intelligent network management. The detailed components of the three-layer network architecture of the data center based on digital twins are as follows:
[0098] Physical data center network layer: The physical data center network layer consists of network elements, servers, and links. In a data center network, network elements can be switches, routers, and equipped applications. These network element devices support the filtering and forwarding of network data packets. Network element devices and applications generate a large amount of data every day. These data may be the backplane throughput of the switch, packet buffer size, MAC address table, interface packet forwarding rate, etc. Then, by extracting valuable information from these data, we can obtain real-time physical network status, such as the health status of network element devices, traffic matrix, available bandwidth changes, etc., so as to detect network anomalies.
[0099] Digital twin network layer: The digital twin network layer is responsible for establishing a real-time mapping model between the physical data center network and the virtual twin network system, including three key submodules: data storage, model mapping, and twin management. The data storage module collects various configuration and operation data of physical network entities through southbound interface protocols (such as NETCONF, OpenFlow, XMPP, I2RS and other protocols) and stores them in the database for the establishment, simulation and optimization of basic models and functional models. Specifically, the data storage module collects data from the physical data center network and divides the data into four different types for storage: configuration of servers, switches, and links, status of servers, switches, and links, network topology data structure, and traffic matrix. Based on the collected data, the model mapping module needs to extract, define, and describe the key features of each entity in the physical network. On the one hand, it is to build the network element model and topology model of the data center network, that is, the basic model, and on the other hand, it can provide training data and test data for the input of the functional model. Among them, the network element model is a real-time and accurate mapping of servers, switches, and links, and the topology model is obtained by connecting and combining the network element models according to the network topology data structure. The functional model is to establish data models such as network simulation, analysis, and optimization based on the network data in the database. A major feature of the digital twin network is to achieve real-time mirroring of the physical network, so the model mapping module also needs to be put into operation. To this end, the twin management module is responsible for managing and updating the various mapping models in the network twin layer, and has functions such as model update, state synchronization, model interaction, and application association. For example, through the state synchronization function, traffic is replayed in the topology model to achieve synchronous mapping of the twin network to the physical network, so that the functional model can make corresponding network optimization strategies.
[0100] Network application layer: Based on the physical data center network and the digital twin network, the network application layer can support various network applications, such as QoS configuration, network logs, network visualization, policy verification and other applications. The network application layer can input QoS (latency, throughput, etc.) requirements and network policies to the twin network layer through northbound interface protocols (such as RESTCONF, NETCONF, SNMP and other protocols), and simulate the specified network policies in the twin network layer through model instances. In this process, the digital twin network can simulate the behavior of the physical network and predict the effects of the designed policies without deploying network policies on the real data center network. If the designed policy can meet the QoS requirements, the policy can be deployed to the real network environment through the southbound interface protocol.
[0101] Defining the mathematical model of load balancing
[0102] For the physical data center network DCN〈V,E〉, use e i,j Indicates that from v i to v j The maximum load of the link is B i,j , the actual load is b i,j At time t, suppose there are k flows in the network, and the set of flows is represented by F = {f1, f2, ..., f k}, one of the flows f s In link e i,j The occupied bandwidth on Define f s The four-tuple of s∈{1,2,…,k}, where: o s ∈V represents the source node; d s ∈V represents the destination node; r s Represents the flow f s Required bandwidth; t s Indicates the size of the flow, and its value equal to 1 indicates f s For large flow, a value equal to 0 means f s For small stream.
[0103] Initially, the KSP algorithm calculates the path set between all edge node pairs The lth path has Links, is the source node v o To the destination node v d Directed sequence of <v o →v i →v j →…→v dIn order to dispatch new flows reaching edge nodes to appropriate paths for forwarding, network traffic needs to be evenly distributed to the entire network links to achieve load balancing in the data center network. Therefore, based on the concept of coefficient of variation in statistics, the present invention proposes a relative balance degree of path load CV. l :
[0104]
[0105]
[0106]
[0107] st
[0108]
[0109]
[0110] In the formula, represents the average link load of the lth path, and σ represents the standard deviation of the link load of the lth path. Formula (4) is the link load constraint, which ensures that the traffic distributed on the link does not exceed the maximum load of the link. Formula (5) is the traffic conservation constraint, which ensures that the amount of data flowing out of the source destination node is equal to the amount of data flowing into the destination node, where Γ(v i )={v j :e i,j ∈E},Γ′(v i )={v j :e j,i ∈E}.
[0111] Path load relative balance CV l Can measure and flow f s The smaller the load balance degree of different paths to the same destination node, the smaller the discrete degree of the link load contained in the path, and the more relatively uniform the load of the path. In data statistical analysis, when the relative balance of the path load is greater than 15%, the load of the path may be abnormal. If the newly arrived flow is assigned to the path, it may lead to uneven distribution of network resources and waste of network resources. Although the relative balance of the path load can measure whether the traffic is evenly distributed on each link to a certain extent, it cannot analyze the transmission delay of the new flow assigned to the current path. If the traffic distribution of the entire network link is relatively uniform, but the link delay is large, it will affect the flow completion time, which is intolerable for delay-sensitive flows. Therefore, when assigning a new flow to the current path, the link delay of the path also needs to be considered.
[0112] Link delay cannot be obtained through switch port status information. The controller needs to send a directional message to the specified switch to obtain link delay information and calculate link e i,j The specific steps are as follows:
[0113] (1) The controller sends a Packet_out message carrying the current timestamp to the switch v i , which indicates that the switch v i Send it to the adjacent switch v j .
[0114] (2) Switch v j Receive switch v i If the sent message cannot match the flow table entry, it will be forwarded to the controller as a Packet_in message. After receiving the message, the controller will unpack it to get the timestamp, and then subtract the timestamp from the current timestamp to get the time difference ΔT1. Then the time from the controller sending the message to receiving the message is:
[0115]
[0116] (3) Similarly, the controller sends a Packet_out message with the same timestamp to the switch v j , and then the controller receives the i Receive the Packet_in message and record the time difference ΔT2:
[0117]
[0118] (4) The controller sends a message to the switch v i and switch v j Send an Echo request message with a timestamp. After receiving the message, the switch replies with an Echo reply message. After receiving the Echo reply message, the controller calculates the timestamp to the switch v i and switch v j The round trip time is thus and Therefore, the switch v i To switch v j The round trip time is:
[0119]
[0120] (5) Assume that switch v i With switch v j The round-trip time of link e is the same, and the processing time of the switch is ignored, because the calculation of each link delay will generate switch overhead. i,j The delay is:
[0121]
[0122] When calculating the link delay, if the link is severely congested and the controller in step (2) cannot receive the message, this paper sets the currently calculated link delay to a high value. i,j The delay is:
[0123]
[0124] In summary, the delay of the lth path can be expressed as:
[0125]
[0126] Therefore, considering the relative balance of path load and path delay, the established objective function is as follows:
[0127] minα·CV l +(1-α)·T l (50)
[0128] Wherein, α∈(0,1), and in the embodiment of the present invention, the value of α is 0.5.
[0129] Defining the mathematical model of the digital twin network
[0130] In the physical data center network, there are n network element devices, and V is defined as {v1, v2, …, v n} is a set of network element devices, with m links, and E = {e1, e2, …, e m} is a link set, the digital twins of the network element model and link model can be expressed as:
[0131] DT v (t) = Θ(C v ,S v (t),M v (t)) (51)
[0132] DT e (t) = Θ(C e ,S e (t)) (52)
[0133] Among them, C represents static configuration data. For network element device v, it can be the maximum transmission rate, backplane bandwidth, MAC address capacity, etc.; for link e, it can be the maximum capacity. S(t) represents the operation status that changes over time, and the operation status is determined by multi-dimensional features. For network element device v, its x-dimensional state feature is defined as Such as the CPU usage of network element equipment, packet buffer size, port status information, etc. Similarly, the link also has multi-dimensional status characteristics, but this article only considers the current load characteristics of the link. v (t) represents the operation behavior of network element device v, which is characterized by the y-dimensional characteristic behavior and is defined as
[0134] The topology model is obtained by connecting and combining the network element model and the link model according to the network topology data structure. The digital twin of the topology model can be expressed as:
[0135] DT(t)=Θ(DT V (t),DT E (t)) (53)
[0136] Among them, DT V (t) represents the set of all network element models in the physical data center network, that is, DT E (t) represents the set of all link models in the physical data center network, namely
[0137] The combination of the network element model set and the link model set can obtain the digital twin of the topology model. Various types of data are collected through the data storage module of the twin network layer, and the network element model and link model are constructed in a digital form, and then the topology model is constructed to achieve a comprehensive and accurate mapping of the physical data center network. Assisted by the adaptive and self-learning capabilities of the functional model, the digital twin network can finally achieve real-time control and optimization of the physical data center network. Based on the above analysis, the relationship between the physical data center network and the digital twin network can be formally defined as:
[0138]
[0139] Among them, DCN<V,E> represents the physical data center network, V represents the set of all network nodes, and E represents the set of all links in the network; SIP represents the southbound interface protocol, through which communication between the physical data center network and the digital twin network is realized.
[0140] Algorithm design of the elephant flow scheduling module
[0141] The embodiment of the present invention implements digital twin-assisted data center network load balancing based on the DDPG algorithm. The four networks of the DDPG algorithm are described as follows:
[0142] Actor current network: responsible for iterative update of policy network parameters θ, responsible for selecting the current action A according to the current state S, and used to interact with the environment to generate S′, R.
[0143] Actor Target Network: Responsible for selecting the optimal next action A′ based on the next state S′ sampled from the experience replay pool. The network parameters θ′ are periodically copied from θ.
[0144] Critic current network: Iterative update of the replication value network parameter ω, responsible for calculating the current Q value Q(S,A,ω), the target Q value y i =R+γQ′(S′,A′,ω′).
[0145] Critic target network: responsible for calculating the Q'(S', A', ω') part of the target Q value. The network parameters ω' are copied from ω regularly.
[0146] First, the data of new flows arriving at the edge switch is collected. The twin network layer determines whether the new flow is a small flow or a large flow. If it is a small flow, the polling mechanism is used for routing. Otherwise, the large flow scheduling module is used to decide the routing. The design of the large flow scheduling module is described as follows:
[0147] State: The state of DDPG learning is a space that reflects the data center network environment. The state of the embodiment of the present invention is an n×n traffic matrix TM, where b 1,n is switch v1 and switch v n The actual load of the connected links is expressed as follows:
[0148]
[0149] Action: In DDPG, the agent maps the state space to the action space to learn the optimal strategy. In the system of the embodiment of the present invention, the action space is a set of link weights, and the link weights are expressed as follows:
[0150] W=[w1,w2,…,w m ] T (56)
[0151] After obtaining a set of link weights, the routing path of the new flow can be obtained through the path calculation module, and then the flow table entry can be obtained through the flow table management module. The flow table is sent to the twin network, and then sent to the physical network after verification, finally realizing the path selection and forwarding of large flows.
[0152] Reward: Based on the current state and action, the agent receives a reward from the environment. In a data center network, since the reward is related to the objective function of network optimization, the reward function is as follows:
[0153] R(s(t),a(t))=-(min α·CV l +(1-α)·T l ) (57)
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A method for digital twin-assisted data center network load balancing, characterized in that: The following steps are involved: S1: Build a data center network framework based on digital twins, including the physical data center network layer, twin network layer, and network application layer; S2: Collect data on the physical data center network, including the configuration of servers, switches, links, the status of servers, switches, links, network topology data structure, and traffic matrix; for the physical data center network DCN<V,E> , use e i,j Indicates that from v i to v j The maximum load of the link is B i,j , the actual load is b i,j At time t, suppose there are k flows in the network, and the set of flows is represented by F = {f1, f2, ..., f k }, one of the flows f s In link e i,j The occupied bandwidth on Define f s The quaternary group is (o s ,d s ,r s ,t s ), s∈{1,2,…,k}, where o s ∈V represents the source node; d s ∈V represents the destination node; r s Represents the flow f s Required bandwidth; t s Indicates the size of the flow, and its value equal to 1 indicates f s For large flow, a value equal to 0 means f s For small streams; Initially, the KSP algorithm calculates the path set between all edge node pairs The lth path has Links, is the source node v o To the destination node v d Directed sequence of <v o →v i →v j →…→v d >Connected; relative balance of path load CV l : In the formula, represents the average link load of the lth path, σ represents the standard deviation of the link load of the lth path; Formula (4) is the link load constraint, which ensures that the traffic distributed on the link does not exceed the maximum load of the link; Formula (5) is the traffic conservation constraint, which ensures that the amount of data flowing out of the source destination node is equal to the amount of data flowing into the destination node, where Γ(v i )={v j :e i,j ∈E},Γ′(v i )={v j :e j,i ∈E}; S3: Calculate link utilization, latency, packet loss rate, switch load, and traffic matrix; S4: In the current collection cycle, determine the size of the new flow arriving at the edge switch. If it is a mouse flow, use the ECMP method for routing; if it is an elephant flow, transfer it to the elephant flow scheduling module for scheduling; S5: The Elephant Flow Scheduling Module uses DDPG to make scheduling decisions. Its output is a set of link weights. The path calculation submodule calculates the optimal forwarding path, and then the flow table is sent to the physical data center network through the traffic management module. S6: Through the southbound interface protocol, the physical data center network layer receives the flow table and selects the optimal path for rerouting the large flow.
2. The method for digital twin-assisted data center network load balancing according to claim 1 is characterized in that: The construction of a data center network framework based on digital twins described in step S1 specifically includes the following steps: S11: Use Mininet simulation to build a physical data center network layer, where the physical data center network layer is a Fat-tree network topology structure; S12: Constructing a twin network layer, including a data storage module, a model mapping module and a twin management module; the data storage module stores various types of collected data in a Redis database through a controller; the model mapping module maps the collected data into a functional model and a basic model, and operates through the twin management module; S13: Build a network application layer, obtain valuable information from the twin network layer database through the northbound interface protocol, and realize network visualization and network logging.
3. The method for digital twin-assisted data center network load balancing according to claim 2 is characterized in that: The physical data center network layer is composed of network elements, servers, and links. In the data center network, the network elements are switches, routers, and equipped applications, which are used to support the filtering and forwarding of network data packets.
4. The method for digital twin-assisted data center network load balancing according to claim 2 is characterized in that: The digital twin network layer is used to establish a real-time mapping model between the physical data center network and the virtual twin network system, collect various configuration and operation data of the physical network entity through the southbound interface protocol, and store it in the database for the establishment, simulation and optimization of the basic model and functional model; The data storage module collects data from the physical data center network and divides the data into four different types for storage: the configuration of servers, switches, and links; the status of servers, switches, and links; the network topology data structure; and the traffic matrix; The model mapping module extracts, defines and describes the key features of each entity in the physical network based on the collected data, builds the network element model and topology model of the data center network, and provides training data and test data for the input of the functional model; The network element model is a real-time and accurate mapping of servers, switches, and links. The topology model is obtained by connecting and combining the network element models according to the network topology data structure. The functional model is to establish a network simulation, analysis, and optimization data model based on the network data in the database; The twin management module is responsible for managing and updating each mapping model in the network twin layer, and has the functions of model update, status synchronization, model interaction, and application association.
5. The method for digital twin-assisted data center network load balancing according to claim 2 is characterized in that: The network application layer is based on the physical data center network and the digital twin network, and supports the following network applications: QoS configuration, network log, network visualization, and policy verification; The network application layer inputs QoS requirements and network policies to the twin network layer through the northbound interface protocol, and simulates the specified network policies in the twin network layer through model instances; If the designed policy can meet the QoS requirements, the policy is deployed to the real network environment through the southbound interface protocol.
6. The method for digital twin-assisted data center network load balancing according to claim 1, characterized in that: In step S3, the controller sends a directional message to the specified switch to obtain link delay information and calculate link e i,j The specific steps are as follows: S31: The controller sends a Packet_out message carrying the current timestamp to the switch v i , which indicates that the switch v i Send it to the adjacent switch v j ; S32: Switch v j Receive switch v i If the sent message cannot match the flow table entry, it will be forwarded to the controller as a Packet_in message. After receiving the message, the controller unpacks it to obtain the timestamp, and then subtracts the timestamp from the current timestamp to obtain the time difference ΔT1. The time from the controller sending the message to receiving the message is: S33: The controller sends a Packet_out message carrying a timestamp to the switch v j , then the controller receives the i Receive the Packet_in message and record the time difference ΔT2: S34: The controller sends a message to the switch v i and switch v j Send an Echo request message with a timestamp. After receiving the message, the switch replies with an Echo reply message. After receiving the Echo reply message, the controller calculates the timestamp to the switch v i and switch v j The round trip time is thus and Switch v i To switch v j The round trip time is: S35: Assume that switch v i With switch v j The round-trip time of link e is the same, and the processing time of the switch is ignored. i,j The delay is: When calculating the link delay, if the link is severely congested and the controller in step S32 cannot receive the message, the currently calculated link delay is set to a high value. i,j The delay is: In summary, the delay of the lth path is expressed as: Taking into account the relative balance of path load and path delay, the established objective function is as follows: minα·CV l +(1-a)·T l (12) Among them, α∈(0,1).
7. The method for digital twin-assisted data center network load balancing according to claim 1, characterized in that: The mathematical model of the digital twin network is: In the physical data center network, there are n network element devices, and V is defined as {v1, v2, …, v n } is a set of network element devices, with m links, and E = {e1, e2, …, e m } is a link set, the digital twins of the network element model and link model are expressed as: DT v (t)=Θ(C v ,S v (t),M v (t)) (13) DT e (t)=Θ(C e ,S e (t))(14) Among them, C represents static configuration data. For network element device v, it is the maximum transmission rate, backplane bandwidth, and MAC address capacity; for link e, it is the maximum capacity; S(t) represents the operating state that changes over time, which is determined by multi-dimensional features. For network element device v, its x-dimensional state feature is defined as M v (t) represents the operation behavior of network element device v, which is characterized by the y-dimensional characteristic behavior and is defined as The topology model is obtained by connecting and combining the network element model and the link model according to the network topology data structure. The digital twin of the topology model is represented as follows: DT(t)=Θ(DT V (t),DT E (t)) (15) Among them, DT V (t) represents the set of all network element models in the physical data center network, that is, DT E (t) represents the set of all link models in the physical data center network, namely The relationship between the physical data center network and the digital twin network is formally defined as: Among them, DCN<V,E> represents the physical data center network, V represents the set of all network nodes, and E represents the set of all links in the network; SIP represents the southbound interface protocol, through which communication between the physical data center network and the digital twin network is realized.
8. The method for digital twin-assisted data center network load balancing according to claim 1, characterized in that: Digital twin-assisted data center network load balancing is implemented based on the DDPG algorithm. The four networks of the DDPG algorithm are: Actor current network: responsible for iterative update of policy network parameters θ, responsible for selecting the current action A according to the current state S, used to interact with the environment to generate S′, R; Actor target network: responsible for selecting the optimal next action A′ based on the next state S′ sampled in the experience replay pool, and the network parameters θ′ are regularly copied from θ; Critic current network: Iterative update of the replication value network parameter ω, responsible for calculating the current Q value Q(S,A,ω), the target Q value y i =R+γQ′(S′,A′,ω′); Critic target network: responsible for calculating the Q′(S′,A′,ω′) part of the target Q value, and the network parameter ω′ is regularly copied from ω; The scheduling method of the elephant flow scheduling module in step S4 is as follows: State: It is an n×n traffic matrix TM, where b 1,n is switch v1 and switch v n The actual load of the connected links is expressed as follows: Action: The action space is a set of link weights, which are expressed as follows: In=[in1,in2,…,in m ] T (18) After obtaining a set of link weights, the routing path of the new flow is obtained through the path calculation module, and then the flow table entry is obtained through the flow table management module. The flow table is sent to the twin network, and then sent to the physical network after verification, finally realizing the path selection and forwarding of the large flow; Reward: Based on the current state and action, the agent receives a reward from the environment; in the data center network, the reward function is as follows: R(s(t),a(t))=-(min α·CV l +(1-a)·T l ) (19).
Citation Information
Patent Citations
Perception task processing method and system in digital twin Internet of Vehicles
CN115086917A
Systems and methods for fleet management
US11417154B1