Flow control method for distributed in-network neural network based on programmable data plane, and device
By building a distributed on-network neural network based on programmable data planes in the network, the problem that network traffic control in the prior art is difficult to achieve global optimal load balancing, efficient traffic control and network status monitoring are achieved, and network performance and reliability are improved.
Patent Information
- Application Number
- PCT/CN2024/136071
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-12
AI Technical Summary
The prior art is difficult to achieve global optimal load balancing in network traffic control, and the load balancing algorithm is difficult to make decisions in complex and dynamic network environments, resulting in problems such as traffic hotspots and congestion that cannot be discovered and solved in a timely manner.
The traffic control method of distributed on-network neural network based on programmable data plane is adopted, and the communication link is determined through programmable switches and the source routing method of actively sending data packets is used to determine the communication link, which is unified between the network packet forwarding process and the neural network model training process, and adapts to different network topology through model pruning.
Approximate optimal flow control with in-band decentralized collaboration is realized, which can more accurately observe and process network status, reduce traffic hotspots and congestion, and improve network performance, reliability and scalability.
Smart Images

Figure CN2024136071_12062025_PF_FP_ABST
Abstract
Description
Flow control method and device for distributed on-line neural network based on programmable data plane Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and in particular relates to a flow control method and device for a distributed on-line neural network based on a programmable data plane. Background Art
[0002] Currently, most AI model training accelerators require direct mathematical isomorphism between the hardware physics and the mathematical operations in neural networks (for example, based on GPUs or FPGAs). However, their excessive energy consumption increasingly limits their scalability. Therefore, this embodiment proposes a distributed on-network neural network. Based on multiple interconnected programmable switches, the distributed on-network neural network can perform both training and inference.
[0003] In a network, the essence of flow control is how to distribute traffic appropriately across links, achieving a reasonable temporal and spatial distribution of traffic to improve network performance, reliability, and scalability. However, as a key flow control technology, load balancing currently lacks a globally optimal load balancing model, and current approaches are limited. This is primarily due to the following reasons: First, the complexity and dynamic nature of the network environment leads to significant variations in network load over time and location. Load balancing algorithms must make decisions within this complex environment, making their design and optimization extremely difficult. Furthermore, the limitations of load balancing algorithms also restrict the optimization of load balancing solutions.
[0004] Currently, load balancing algorithms that make decisions in the programmable data plane typically only perform traffic scheduling and routing at the packet level and do not monitor or collect information about network status. Consequently, these algorithms struggle to obtain global network status information or observe changes in the network's overall state. This can lead to potential problems, such as traffic hotspots and congestion, remaining undetected and unresolved. Load balancing algorithms that make decisions in the control plane typically require a centralized controller to collect, analyze, and process network-wide state information. This can result in lengthy decision cycles and additional bandwidth overhead, reducing load balancing efficiency and real-time performance. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and provide a flow control method and device for a distributed on-line neural network based on a programmable data plane. By using a programmable switch to form neurons through matrix-vector linear operations, and then the programmable switch takes the source routing method of actively sending data packets to determine the neuron communication link, the network's packet forwarding process can be unified with the training process of the neural network model, and is widely used in various network tasks such as flow control, flow type prediction, congestion prediction, and anomaly detection. In addition, through model pruning, the distributed on-line neural network can adapt to different network topologies. The present application also provides corresponding or further systems, electronic devices, and computer-readable storage media for the above method.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for flow control of a distributed on-line neural network based on a programmable data plane, comprising the following steps:
[0008] Constructing a distributed on-net neural network, wherein the distributed on-net neural network includes a plurality of neurons, wherein the neurons are formed by a programmable switch, and the programmable switch adopts a source routing method of actively sending data packets to determine a neuron communication link;
[0009] Pruning the distributed neural network model according to the needs of building distributed neural networks with different structures under different network topologies; the model pruning is at least one of pruning neurons and pruning weight connections between neurons;
[0010] The distributed on-line neural network is iteratively trained using a labeled distributed spatiotemporal network behavior feature dataset. The loss value of the loss function is calculated until the model parameters converge to the label value. After the model training is completed, the optimal neuron weights are obtained.
[0011] A mapping relationship is established between the optimal neuron weight and the network behavior characteristics of flow control; the optimal neuron weight is converted into the network behavior characteristics of flow control through the mapping relationship, and the weighted probability forwarding strategy is used to complete the approximately optimal flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping, routing scheduling or flow scheduling.
[0012] As a preferred technical solution, the distributed on-line neural network uses programmable switches to form neurons through matrix-vector linear operations, specifically:
[0013] The matrix-vector linear operation is specifically as follows:
[0014] The output signal x of the i-th neuron in the previous layeri The weight w between the i-th neuron in the previous layer and this neuron i Multiply them together to get w i x i , then sum the outputs of multiple neurons in the upper layer, subtract the activation bias θ of the neurons, and finally obtain the output y after the activation function f as the input of the neurons in the lower layer.
[0015] As a preferred technical solution, the matrix-vector linear operation is implemented by the programmable switch executing the Map-Reduce mechanism based on the matching-action table MAT to perform equivalent calculations, specifically:
[0016] The vector-to-vector operation in matrix-vector linear operations is equivalent to the Map operation in Map-Reduce, and the vector-to-scalar operation in matrix-vector linear operations is equivalent to the Reduce operation in Map-Reduce.
[0017] The Map-Reduce mechanism is located between the two MATs of the programmable switch, realizing pipeline operation with MAT, parser, and scheduler.
[0018] As a preferred technical solution, the programmable switch adopts a source routing method of actively sending data packets to determine the neuron communication link, specifically:
[0019] Determine a source routing path based on the need to transfer model parameters between neurons of the distributed neural network model;
[0020] When neurons communicate, the programmable switch actively sends a data packet with a customized header using the P4 programming language, and then embeds the model parameters into the data packet; the customized header includes the addresses of the switches along the way required for source routing;
[0021] The data packets are forwarded according to the source routing path determined above, completing the transmission of the distributed neural network model parameters between neurons, ensuring that the delay of communication transmission between neurons is deterministic.
[0022] As a preferred technical solution, the labeled distributed spatiotemporal network behavior feature dataset is specifically:
[0023] In the neurons of all input layers of the distributed neural network, the same network behavior feature is extracted in each time slot, and different network behavior features are extracted in different time slots; the neurons of all input layers are the detection nodes;
[0024] The network behavior features of multiple detection nodes in multiple time slots are used to build a labeled distributed spatiotemporal joint network behavior feature dataset.
[0025] As a preferred technical solution, the distributed on-line neural network uses programmable switches to form neurons through matrix-vector linear operations, specifically:
[0026] The matrix-vector linear operation is specifically as follows:
[0027] The output signal x of the i-th neuron in the previous layer i The weight w between the i-th neuron in the previous layer and this neuron i Multiply them together to get w i x i , then sum the outputs of multiple neurons in the upper layer, subtract the activation bias θ of the neurons, and finally obtain the output y after the activation function f as the input of the neurons in the lower layer.
[0028] As a preferred technical solution, the matrix-vector linear operation is implemented by the programmable switch executing the Map-Reduce mechanism based on the matching-action table MAT to perform equivalent calculations, specifically:
[0029] The vector-to-vector operation in matrix-vector linear operations is equivalent to the Map operation in Map-Reduce, and the vector-to-scalar operation in matrix-vector linear operations is equivalent to the Reduce operation in Map-Reduce.
[0030] The Map-Reduce mechanism is located between the two MATs of the programmable switch, realizing pipeline operation with MAT, parser, and scheduler.
[0031] As a preferred technical solution, the programmable switch adopts a source routing method of actively sending data packets to determine the neuron communication link, specifically:
[0032] Determine a source routing path based on the need to transfer model parameters between neurons of the distributed neural network model;
[0033] When neurons communicate, the programmable switch actively sends a data packet with a customized header using the P4 programming language, and then embeds the model parameters into the data packet; the customized header includes the addresses of the switches along the way required for source routing;
[0034] The data packets are forwarded according to the source routing path determined above, completing the transmission of the distributed neural network model parameters between neurons, ensuring that the delay of communication transmission between neurons is deterministic.
[0035] As a preferred technical solution, the labeled distributed spatiotemporal network behavior feature dataset is specifically:
[0036] In the neurons of all input layers of the distributed neural network, the same network behavior feature is extracted in each time slot, and different network behavior features are extracted in different time slots; the neurons of all input layers are the detection nodes;
[0037] The network behavior features of multiple detection nodes in multiple time slots are used to build a labeled distributed spatiotemporal joint network behavior feature dataset.
[0038] As a preferred technical solution, the construction of a distributed spatiotemporal network behavior feature dataset with labels is specifically as follows:
[0039] All input layer neurons of the distributed neural network extract link flow density from the forwarded packets in each time slot, where the link flow density is the number of packets per unit time;
[0040] A labeled distributed spatiotemporal network behavior feature dataset is constructed using the link traffic density of multiple key time slots of multiple key switches on the path that the packet passes through.
[0041] As a preferred technical solution, the neuron waits for all its inputs to arrive within a time slot, and the time slot is used for synchronization between neurons, so that each neuron can process its input before new input arrives.
[0042] As a preferred technical solution, a distributed on-line neural network is iteratively trained using a labeled distributed spatiotemporal network behavior feature dataset. The iterative training requires multiple iterations of forward propagation and backward propagation so that the predicted value of the on-line neural network model converges to the label value. One iteration of forward propagation and backward propagation specifically includes the following steps:
[0043] The forward propagation converts the network behavior feature data set into a standard numerical value, and then inputs the numerical value in parallel into the input layer of the distributed neural network, propagates forward layer by layer in the intermediate layer of the distributed neural network, outputs the predicted value at the output layer, and finally calculates the error between the predicted value and the label value by calculating the loss value of the loss function;
[0044] The backward propagation propagates the error from the output layer back to the input layer layer by layer based on the chain rule, and in this process updates the weights and biases of the distributed neural network.
[0045] As a preferred technical solution, the method of achieving approximately optimal flow control in a data center network specifically includes the following steps:
[0046] The data center network uses the Equal Cost Multi-Path (ECMP) routing method to forward packets. The programmable switch collects network behavior characteristics once every time slot and stores them in the programmable switch's register. The network controller collects data from the register every time interval T to construct a labeled distributed spatiotemporal network behavior characteristic dataset. Every time interval T, the labeled distributed spatiotemporal network behavior characteristic dataset adds a set of data.
[0047] Using a distributed spatiotemporal network behavior feature dataset with labels, a distributed on-network neural network is iteratively trained. When the sliding average of the loss value decreases to a predetermined threshold, the training is stopped.
[0048] In the network controller, the weights of each neuron are obtained from the trained neural network model and sent to the programmable switch;
[0049] The programmable switch receives the neuron weights issued and stores them in its own registers;
[0050] When a packet arrives at a programmable switch, the programmable switch uses a weighted probabilistic forwarding strategy to select a forwarding port for the packet.
[0051] As a preferred technical solution, the optimization objectives of the label flow control are specifically: end-to-end delay, flow completion time and deadline satisfaction rate when the network is completely free of congestion.
[0052] As a preferred technical solution, the mapping relationship between the neuron weights and the network behavior characteristics of the flow control is established as follows:
[0053] The mapping relationship between neuron weights and network behavior characteristics of flow control is an equal relationship;
[0054] The link flow density of the links between programmable switches is a network behavior feature of the flow control, and the value of the link flow density of the links between programmable switches is equal to the value of the neuron weight.
[0055] As a preferred technical solution, the loss function loss in the distributed neural network training is specifically:
[0056] Among them, MSE represents the mean square error function, represents the predicted end-to-end delay, flow completion time, and deadline satisfaction rate of the i-th packet in the data center network, D i represents the corresponding i-th label, and n represents the number of features.
[0057] As a preferred technical solution, the weighted probability forwarding strategy is specifically as follows: based on the mapping between link traffic density and neuron weight, the probability of a packet being forwarded by the i-th switch port is:
[0058] Among them, W i is the weight of the i-th switch port, P i is the probability that port i is selected, and the weight of the switch port is the neuron weight of the corresponding neuron.
[0059] As a preferred technical solution, under the K=4 fat-tree network topology, the labeled distributed spatiotemporal network behavior feature dataset includes the following steps:
[0060] Set the link traffic density of ports 3 and 4 of the programmable switch;
[0061] At the tth time slot, a sample without spatiotemporal joint correlation characteristics of the dataset used to train the neural network is obtained as follows:
[0062] n=4
[0063] in, represents the link traffic density of the i-th port of the n-th switch at the t-th time slot.
[0064] As a preferred technical solution, under the K=4 fat-tree network topology, the acquisition of the labeled distributed spatiotemporal joint network behavior feature dataset further includes the following steps:
[0065] Set the link traffic density of ports 3 and 4 of the programmable switch;
[0066] Set the input of the first neuron in the input layer of the neural network
[0067] Set the input of the second neuron in the input layer
[0068] Set the input I of the nth neuron in the input layer n (t+n-1) is:
[0069] At the tth time slot, a sample with time-related characteristics of the data set used to train the neural network on the network is obtained, as follows:
[0070] n=4
[0071] Among them, I n (t+n-1) represents the input of the nth neuron in the input layer of the neural network at the tth time slot, represents the link traffic density of the i-th port of the n-th switch at the t-th time slot.
[0072] As a preferred technical solution, under the K=4 fat-tree network topology, the acquisition of the labeled distributed spatiotemporal joint network behavior feature dataset further includes the following steps:
[0073] Set the link traffic density of ports 3 and 4 of the programmable switch;
[0074] The programmable switches s1, s9, and s17 are used as the input of the first neuron in the input layer of the neural network, as shown in the following formula:
[0075] The programmable switches s2, s10, and s19 are used as the input of the second neuron in the input layer of the neural network, as shown in the following formula:
[0076] The programmable switches s3, s11, and s18 are used as the input of the third neuron in the input layer of the neural network.
[0077] The programmable switches s4, s12, and s20 are used as the input of the fourth neuron in the input layer of the neural network.
[0078] At the tth time slot, a sample with spatiotemporal joint correlation characteristics of the data set used to train the on-line neural network is obtained as follows:
[0079] Among them, I n (t) represents the input of the nth neuron in the input layer of the neural network at the tth time slot, represents the link traffic density of the i-th port of the n-th switch at the t-th time slot.
[0080] As a preferred technical solution, under the K=4 fat-tree network topology, a specific method of setting the end-to-end delay as a label of the distributed spatiotemporal joint network behavior feature dataset is as follows:
[0081] When the network is completely uncongested, at time slot t, the average end-to-end delays of packets between switches s1 and s5, s2 and s6, s3 and s7, and s4 and s8 are D s1-s5 (t), D s2-s6 (t), D s3-s7 (t), D s4-s8 (t);
[0082] At time slot t, the label D(t) of the sample is set to: D(t) = [D s1-s5 (t),D s2-s6 (t),D s3-s7 (t),D s4-s8 (t)]
[0083] Then, at the tth time slot, the labeled samples in the aforementioned distributed spatiotemporal joint network behavior feature dataset are C1(t)-D(t), C2(t)-D(t), and C3(t)-D(t); C1(t) refers to the sample without spatiotemporal joint correlation characteristics, C2(t) refers to the sample with time correlation characteristics, and C3(t) refers to the sample with spatiotemporal joint correlation characteristics;
[0084] A distributed spatiotemporal network behavior feature dataset with labels is constructed using the multiple labeled samples.
[0085] In a second aspect, the present invention further provides an in-band decentralized collaborative load balancing system for a distributed on-line neural network, which is applied to the flow control method of the distributed on-line neural network based on a programmable data plane, and includes an on-line neural network construction module, a model pruning module, an on-line neural network training module, and a flow control module;
[0086] The on-line neural network construction module is used to construct a distributed on-line neural network, wherein the distributed on-line neural network includes a plurality of neurons, and the neurons are formed by a programmable switch, and the programmable switch adopts a source routing method of actively sending data packets to determine the neuron communication link;
[0087] The model pruning module is used to prune the model of the distributed on-line neural network according to the requirements of constructing distributed on-line neural networks with different structures under different network topologies; the model pruning is at least one of pruning neurons and pruning weight connections between neurons;
[0088] The network training module is used to iteratively train the distributed on-line neural network using the labeled distributed spatiotemporal network behavior feature dataset, calculate the loss value of the loss function, until the model parameters converge to the label value, and obtain the optimal neuron weights after completing the model training;
[0089] The flow control module is used to establish a mapping relationship between the optimal neuron weight and the network behavior characteristics of the flow control; through the mapping relationship, the optimal neuron weight is converted into the network behavior characteristics of the flow control, and the weighted probability forwarding strategy is used to complete the approximately optimal flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping, routing scheduling or flow scheduling.
[0090] In a third aspect, the present invention provides an electronic device, comprising:
[0091] at least one processor; and,
[0092] a memory communicatively connected to the at least one processor; wherein,
[0093] The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the traffic control method of the distributed in-network neural network based on the programmable data plane.
[0094] In a fourth aspect, the present invention provides a computer-readable storage medium comprising computer-readable instructions, which, when executed on a computing device or a computing device cluster, enable the computing device or computing device cluster to execute a flow control method for a distributed in-network neural network based on a programmable data plane.
[0095] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0096] (1) The present invention forms neurons by using a programmable switch through matrix-vector linear operations, and then the programmable switch uses a source routing method of actively sending data packets to determine the neuron communication link, which can unify the network's packet forwarding process and the training process of the neural network model, thereby achieving approximately optimal flow control with in-band decentralized collaboration.
[0097] (2) The distributed deployment paradigm adopted by the present invention allows the input of neurons to come from many spatiotemporally distributed network devices, which actually adds an ability to associate multiple spatiotemporally joint observations of network states and infer their behavior in real time. With this ability, the neural network model can make more accurate decisions for network tasks based on the distributed spatiotemporal joint correlation of network behavior.
[0098] (3) The present invention can make the distributed on-line neural network adapt to different network topologies by pruning the model of the distributed on-line neural network according to needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0100] FIG1 is a flow chart of a method for controlling traffic in a distributed neural network based on a programmable data plane according to an embodiment of the present invention;
[0101] FIG2 is a schematic diagram of an implementation of a neuron in a programmable switch according to an embodiment of the present invention;
[0102] FIG3 is a distributed on-line neural network under a K=4 fat-tree network topology according to an embodiment of the present invention;
[0103] FIG4 is a schematic diagram of a probabilistic port selection method according to an embodiment of the present invention;
[0104] FIG5 shows the performance of a flow control method based on a distributed neural network according to an embodiment of the present invention;
[0105] FIG6 is a schematic diagram of the structure of a traffic control system of a distributed on-line neural network based on a programmable data plane according to an embodiment of the present invention;
[0106] FIG7 is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0107] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0108] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0109] Referring to FIG. 1 , this embodiment proposes a method for controlling traffic in a distributed neural network based on a programmable data plane, including the following steps:
[0110] S1. Construct a distributed on-line neural network, wherein the distributed on-line neural network includes multiple neurons, and the neurons are formed by programmable switches. The programmable switches use a source routing method of actively sending data packets to determine neuron communication links.
[0111] In this embodiment, the setting of neurons adopts the Map-Reduce mechanism equivalent to matrix-vector linear operations. As shown in part (a) of Figure 2, the calculation of a neuron in the neural network can be described by the MP model. Specifically, the calculation process is as follows: the output signal x from the i-th neuron in the previous layer is i The weight w between the i-th neuron in the previous layer and this neuron i Multiply them together to get w i x i , then sum the outputs of multiple neurons in the upper layer, subtract the activation bias θ of the neurons, and finally obtain the output y after the activation function f as the input of the neurons in the lower layer. It can be expressed as follows:
[0112] Simply put, neuron computing includes matrix-vector linear operations and element-by-element nonlinear operations (activation functions), among which matrix-vector linear operations are the main part that consumes computing resources. Programmable switches can support the element-by-element nonlinear operations required by activation functions, but do not support matrix-vector linear operations. However, as can be seen from the above process, this embodiment shows that matrix-vector linear operations can be decomposed into many multiplications and additions, and the calculation process of all neurons is the same. This feature is very similar to the Map and Reduce processes of the Map-Reduce mechanism in cloud computing, and the switch can support these multiplications and additions. Therefore, this embodiment can decompose the matrix-vector linear operation process of neurons into the Map and Reduce processes of the Map-Reduce mechanism, thereby realizing neuron computing in the switch.
[0113] Currently, programmable switches are designed based on a match-action table (MAT), which is then coupled with pipeline technology to achieve high-speed forwarding. Therefore, as shown in part (b) of Figure 2, this embodiment utilizes the computing resources of the programmable switch to design a module that executes the Map-Reduce mechanism to perform the computation equivalent to a neuron. This module is located between the two MATs of the switch to achieve pipeline operation with the MATs, parser, and scheduler. Vector-to-vector operations are equivalent to the map in Map-Reduce, and vector-to-scalar operations are equivalent to the reduce in Map-Reduce. Simply put, this implements the design of a neuron. Furthermore, this embodiment can also extend this mechanism to be executed by multiple switches. That is, each switch performs one multiplication or one addition in the Map-Reduce operation in part (b) of Figure 2. In other words, multiple switches collaborate to complete the computation of a neuron.
[0114] Furthermore, both forward and backward propagation during neural network training or inference require the transfer of model parameter data (e.g., gradients) between neurons in each layer. In the distributed on-net neural network, programmable switches are responsible for computing tasks for worker nodes, meaning that model parameter data must be transmitted between programmable switches. Communication between switches in packet networks is best-effort, while communication between neurons in the distributed on-net neural network requires reliable and continuous communication. This creates a mismatch between the distributed neural network at the upper layer and the underlying packet network.
[0115] To overcome the above contradictions, this embodiment adopts the following preferred solutions:
[0116] Preferably, the programmable switch adopts a source routing method of actively sending data packets to determine the neuron communication link, specifically:
[0117] S11. Determine a source routing path by calculating an optimization algorithm based on the need for model parameters transmitted between neurons of the distributed neural network model;
[0118] S12. During inter-neuron communication, the programmable switch actively sends a data packet with a customized header using the P4 (Programming Protocol-Independent Packet Processors) programming language, and then embeds the model parameters into the data packet; the customized header includes the addresses of the switches along the way required for source routing;
[0119] S13. The data packet is forwarded according to the source routing path determined above, completing the transmission of the distributed neural network model parameters between neurons, ensuring that the delay of communication transmission between neurons is deterministic.
[0120] It should be noted that the communication link between neurons in this embodiment is mainly determined by the communication delay between neurons. Therefore, when the source routing method of actively sending data packets is adopted, in order to ensure the certainty of the communication delay between neurons, the establishment of the communication link is based on the following ideas:
[0121] (1) According to which programmable switches the distributed neural network needs to transfer model parameters, a source routing path is determined through an optimization algorithm.
[0122] (2) When a neuron needs to transmit model parameters, the switch uses P4 to actively send a data packet with a customized header. The header of the data packet contains the addresses of the switches along the way required for source routing and embeds the parameter data to be transmitted into the data packet.
[0123] (3) Since programmable switches support source routing of custom paths, these data packets will be transmitted along the predetermined path. The forwarding process of these data packets also means that the model parameters are transferred between the required programmable switches, thus completing the communication between neurons.
[0124] Next, regarding the implementation of programmable switches, given the resource constraints of programmable switches, this embodiment uses a binary neural network (BNN) as the implementation method for the on-network neural network. Weights and activations in a BNN are represented and calculated using only two binary values: +1 and -1. This makes the BNN highly memory- and computationally efficient, and allows it to run efficiently on low-power and embedded devices.
[0125] It should also be noted that, for the distributed on-line neural network itself in the present invention, as shown in FIG1 , the computing resources and links of the programmable switches in the network are used to construct a distributed on-line neural network, wherein each programmable switch contributes a small portion of resources to realize the computing function of N (N≥1) neurons, and the neurons are connected to each other through the links of the programmable switches.
[0126] The distributed on-line neural network can construct any neural network for any network topology through model pruning technology. Figure 1 shows a general neural network that can be selected according to task requirements, for example, commonly used fully connected neural networks, convolutional neural networks CNN (Convolutional Neural Networks), recurrent neural networks RNN (Recurrent Neural Network), long / short term memory LSTM (Long / Short Term Memory), graph neural networks GNN (Graph Neural Networks), etc. Common network topologies include VL2, Fat-tree, BCube, DCell, etc.
[0127] The core idea of the distributed neural network is:
[0128] (1) This embodiment uses the continuous iteration of the neural network training process to adjust the weights between each neuron. When the training is completed (that is, when the neural network model converges), the globally optimal set of weights is found;
[0129] (2) Establish a mapping relationship (function) between neuron weights and network task characteristics (for example, for the network task of load balancing, its characteristic is link traffic density). In other words, finding the globally optimal set of weights is equivalent to finding the globally optimal distribution of network task characteristics.
[0130] (3) The optimal weights found are converted into the characteristics of the network task through the mapping function to guide the completion of the network task, so that the network task can be completed in the most optimized way, for example, to achieve optimal load balancing.
[0131] During network operation, the characteristics of the network task change over time. Since this embodiment establishes a mapping relationship between the network task characteristics and the neuron weights, the change in the characteristics of the network task is actually the change in the weights in the neural network model.
[0132] S2. Pruning the distributed neural network model according to the requirements of constructing distributed neural networks with different structures under different network topologies; the model pruning is at least one of pruning neurons and pruning weight connections between neurons;
[0133] The network topology is changeable. Common network topologies include VL2, Fat-tree, BCube, DCell, etc. Model pruning can delete neurons and the weight connections between them as needed to construct neural networks with different structures. Therefore, this embodiment uses model pruning technology to make the distributed neural network adapt to different network topologies.
[0134] As shown in Figure 3, for the Fat-tree topology with K=4, the distributed neural network, for neuron 1 in the first layer, when not pruned, it is connected to neurons 1 to 4 in the second layer. However, in the Fat-tree topology with K=4, the s1 programmable switch represented by neuron 1 in the first layer has no connection relationship with the s3 programmable switch and s4 switch represented by neurons 3 and 4 in the second layer, respectively. Therefore, this embodiment trims the connection relationship between them to achieve model pruning.
[0135] S3. Use the distributed spatiotemporal network behavior feature dataset with labels to iteratively train the distributed on-line neural network and calculate the loss value of the loss function until the model parameters converge to the label value. After completing the model training, the optimal neuron weights are obtained.
[0136] The feature engineering in this embodiment has the characteristics of distributed spatiotemporal joint correlation of network behavior. Specifically, as shown in FIG1 , a distributed neural network based on back propagation is shown. The input layer neurons of the distributed neural network, i.e., the detection nodes of the system, extract features from the forwarded data packets. In(t) represents the network behavior (feature) extracted by the nth detection node at the tth time slot, and the coefficient represents the weight from neuron j in layer n to neuron i in layer m, and n represents the number of selected features. These features are used to complete network tasks such as flow control, flow type prediction, congestion prediction, and anomaly detection, as shown in Table 1. This shows the input distributed across the neural network. Different network tasks require different features. Within each time slot, all detection nodes extract the same feature, while different features are extracted in different time slots. As shown in Figure 1, for the input layer, the input of the nth detection node is In(t-n+1), reflecting the temporal correlation of network behavior.
[0137] The detection node converts the extracted features into standardized numerical values and then inputs these values into the distributed neural network in parallel. Each neuron in the distributed neural network waits for all its inputs to arrive within a time slot. This time slot is used for synchronization between neurons, allowing each neuron to process its input before new inputs arrive. After the parallel inputs pass through the distributed neural network, the results are output in parallel at the output layer. For anomaly detection tasks, the output results can be used to determine whether a coordinated attack from multiple sources has occurred, and timely alerts and responses can be issued.
[0138] The process of training the network includes forward propagation and backward propagation. The forward propagation converts the network behavior feature data set into a standard numerical value, and then inputs the numerical value in parallel into the input layer of the distributed on-net neural network, propagates forward layer by layer in the middle layer of the distributed on-net neural network, outputs the predicted value in the output layer, and finally calculates the error between the predicted value and the label value by calculating the loss value of the loss function; the backward propagation propagates the error from the output layer back to the input layer layer by layer based on the chain rule, and updates the weights and biases of the distributed on-net neural network in this process. Furthermore, in order to complete more complex tasks, this embodiment can also use multiple distributed on-net neural networks to cascade into an on-net neural network chain, wherein the output of the first distributed on-net neural network is used as the input of the second distributed on-net neural network. In other words, the analysis results of the first on-net neural network (the spatiotemporal behavior of the feature sequence) are used as the features of the second on-net neural network for further analysis.
[0139] It should also be explained that the "time slot" mentioned in the present invention is actually the time interval s on the data plane.
[0140] Table 1
[0141] Next, for the neural network training process, the model is first forward propagated to obtain the predicted value of the first iteration. Then, the difference between the predicted value of the forward propagation and the label value is used to guide the correction of the weights between neurons by minimizing the loss function. After multiple iterations of correction, the predicted value of the model can be infinitely close to the label value, thereby achieving the predetermined goal (for example, classification). The backward propagation of the neural network is to correct the weights in reverse layer by layer. That is, starting from the output layer, the weights of the previous layer are adjusted layer by layer through the loss function until the input layer is completed. The essence of training the model is actually to find the optimal weights between each neuron.
[0142] S4. Establish a mapping relationship between the optimal neuron weight and the network behavior characteristics of flow control; convert the optimal neuron weight into the network behavior characteristics of flow control through the mapping relationship to complete the flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping or routing flow scheduling.
[0143] In the packet forwarding process of the network, the essence of flow control (congestion control, traffic shaping, load balancing, routing / flow scheduling, etc.) is how to reasonably distribute packets to each link to achieve the best distribution of traffic in time and space, so as to improve the performance, reliability and scalability of the network. To be more specific, the control process of the flow control method is essentially to adjust the link traffic density between each switch to achieve optimization goals, such as low end-to-end delay, small flow completion time, large deadline satisfaction rate, etc. In order to better understand the present invention, the present application provides a second embodiment - an in-band decentralized collaborative load balancing method of a distributed on-line neural network based on a programmable data plane.
[0144] This embodiment unifies the network's packet forwarding process with the training process of the neural network model, establishes a mapping relationship between the neuron weights of the neural network and the link traffic density in the network, and uses continuous iteration of the neural network training process to adjust the weights until a globally optimal set of weights is found, that is, a model of the globally optimal link density distribution. The link traffic density is then considered the weights between neurons. During network operation, because traffic density changes over time, this embodiment establishes a mapping relationship between link traffic density and neuron weights. Therefore, changes in link traffic density are actually changes in weights in the neural network model. Therefore, this embodiment can consider the network operation process as a training process for a neural network model.
[0145] Specifically, the in-band decentralized collaborative load balancing method of a distributed neural network based on a programmable data plane includes the following steps:
[0146] (1) The network uses the ECMP load balancing strategy to guide the flow of traffic. The data plane collects network behavior characteristics every time interval s (i.e., a time slot) and stores them in the registers of the programmable switch. The control plane collects data from the registers every time interval T and constructs a data set. Every time interval T, the data set is incremented. The length of a time slot is set to 100ms. Step (1) is shown in lines 1-10 of Algorithm 1.
[0147] (2) Using the above dataset, the neural network is continuously trained iteratively. When the sliding average of the loss decreases to the threshold q = 0.0001, the training is stopped, as shown in lines 11-13 of Algorithm 1.
[0148] (3) In the control plane, the weights between neurons can be obtained from the trained model, and the weights between neurons are sent to the programmable switch.
[0149] (4) The programmable switch receives the weights issued and stores them in its own register. The switch selects a forwarding port for the packet based on the weighted probability forwarding strategy, as shown in lines 14-18 of Algorithm 1.
[0150] This embodiment requires establishing a mapping relationship between link traffic density and neuron weights. Specifically, a DCN with a K=4 fat-tree network topology is constructed as an on-network neural network, with each switch being a neuron. As shown in Figure 3, neurons 1-4 in the input layer of the neural network correspond to switches s1, s2, s3, and s4 in the network topology. Similarly, each neuron in each layer of the neural network corresponds one-to-one to a switch in the network topology. The links between switches are the connections between neurons, and the traffic density of the links is considered the weight between neurons.
[0151] Taking the second layer as an example, Represents the weight between the i-th neuron in the second layer and the j-th neuron in the first layer. After mapping the link traffic density to the neuron weight, the traffic density between switches is described as shown in formula (1), where: It represents the link traffic density between the i-th switch in layer 2 and the j-th switch in layer 1, and 0 means there is no direct connection between the switches.
[0152] Since the traffic density between each switch in the entire network has a fixed value, the traffic density set shown in formula (1) can actually be regarded as a traffic distribution pattern for the entire network. In other words, such a neural network model can reflect the changes in the traffic distribution pattern.
[0153] In this embodiment, the method for constructing a data set for training a distributed online neural network is as follows:
[0154] As shown in Figure 3, the input layer neurons of the distributed neural network, i.e., the detection nodes of the system, extract features from the forwarded packets. n (t) represents the network behavior (such as number of packets, packet size, etc.) collected by the nth node at the tth time slot, and the coefficient represents the weight from neuron j in layer n to neuron i in layer m.
[0155] The key to load balancing strategy is how to measure the current network congestion. The number of packets passing through the port can well reflect the current busyness of the programmable switch port. Therefore, this embodiment sets the number of packets passing through the switch port per unit time as the input feature of the neural network, where: It represents the number of packets per unit time at the ith port of the nth switch in the tth time slot.
[0156] Specifically, for the K=4 fat-tree network topology, since the input layer neurons correspond to the four switches s1, s2, s3, and s4, and the traffic flow direction decisions are all made at ports 3 and 4 of the switches, in the tth time slot, the number of packets per unit time at ports 3 and 4 of switches s1, s2, s3, and s4 in Figure 3 are respectively Therefore, at the tth time slot, a sample without spatiotemporal joint correlation in the dataset used to train the neural network is shown in formula (2).
[0157] n=4 (2)
[0158] Furthermore, the present invention considers the temporal correlation of features in the neural network input layer. Then, the input of the first neuron in the input layer The input to the second neuron in the input layer The input of the nth neuron in the input layer n = 4. Therefore, at the tth time slot, a sample with time correlation in the dataset used to train the neural network is as shown in formula (3).
[0159] n=4 (3)
[0160] Furthermore, the present invention considers the spatial correlation of features at the neural network input layer, that is, reflects the spatiotemporal joint correlation characteristics of features. This embodiment uses the network behavior of key switches along the packet's path as input to the neural network. More specifically, referring to Figure 3, for a K=4 fat-tree network topology, this embodiment uses s1, s9, and s17 as inputs to the first neuron in the input layer. Take s2, s10, and s19 as the input of the second neuron in the input layer. Take s3, s11, and s18 as the input of the third neuron in the input layer. Take s4, s12, and s20 as the input of the fourth neuron in the input layer. Therefore, at the tth time slot, a sample of the data set used to train the on-net neural network is shown in formula (4).
[0161] When constructing a data set, whether to consider the temporal correlation or spatial correlation of features can be determined according to the actual situation. That is, the samples in the data set can be determined by formula (2), (3), or (4) according to the actual situation.
[0162] In this embodiment, the optimization objectives for load balancing are low end-to-end latency, small FCT, and large DMR. This embodiment sets the optimization objectives as the labels of the aforementioned datasets. For example, when this embodiment uses end-to-end latency as the optimization objective, this embodiment can set the end-to-end latency as the label of the aforementioned dataset. In load balancing, end-to-end latency refers to the time required for a data packet to be transmitted from the source host to the destination host. This latency primarily includes the sending latency at the sender, the queuing latency of intermediate routers, the processing latency, and the propagation latency. End-to-end latency can often reflect network congestion. In a highly loaded network, router buffers may fill up, causing data packets to wait in queue for transmission longer, thereby increasing queuing and processing latency. Furthermore, when network bandwidth is insufficient, the transmission latency of data packets can also increase. The accumulation of these factors can lead to increased end-to-end latency, reflecting network congestion.
[0163] The specific method of setting the end-to-end delay as the label of the above dataset is as follows: when the network is completely uncongested, at the tth time slot, the average end-to-end delay of the packet between switches s1 and s5, s2 and s6, s3 and s7, and s4 and s8 is D s1-s5 (t), D s2-s6 (t), D s3-s7 (t), D s4-s8(t), which are recorded as a matrix as shown in formula (5). This embodiment uses the matrix as the label of the samples C1(t), C2(t), C3(t) as described in formula (2) or (3) or (4).
[0164] At time slot t, the label D(t) of the sample is set to: D(t) = [D s1-s5 (t),D s2-s6 (t),D s3-s7 (t),D s4-s8 (t)] (5)
[0165] Then, in the tth time slot, one labeled sample in the aforementioned distributed spatiotemporal joint network behavior feature dataset is C1(t)-D(t), C2(t)-D(t), and C3(t)-D(t); C1(t) refers to a sample without spatiotemporal joint correlation characteristics, C2(t) refers to a sample with time correlation characteristics, and C3(t) refers to a sample with spatiotemporal joint correlation characteristics.
[0166] This embodiment uses a plurality of labeled samples to construct a dataset for training an on-line neural network.
[0167] In this embodiment, the mean square error (MSE) is used as the loss function for training the distributed neural network, as shown in the following formula:
[0168] in, Denotes the predicted packet delay from the i-th port to the next end in the network, D i represents the average delay from the ith port to the next port, and n represents the number of features;
[0169] The mean square error (MSE) can measure the difference between the predicted value output by the neural network and the actual true value (that is, the label of the data set). Therefore, as shown in formula (6), the loss function of the neural network is designed as MSE, where the output of the neural network is is the predicted end-to-end delay of the packet in the network. As the neural network trains, the model converges and the loss decreases. This process can be thought of as a packet searching for the optimal path in the network to achieve the lowest end-to-end delay.
[0170] In this embodiment, in order to achieve the optimization effect of in-band decentralized collaborative load balancing, this embodiment adopts a weighted probability forwarding strategy, which is
[0171] According to the mapping between link traffic density and neuron weight, the probability of a packet being forwarded by port i is: P i =W i / (∑W i) (7)
[0172] Among them, W i is the weight of switch port i, P i is the probability that port i is selected.
[0173] After training is complete, this embodiment obtains an on-network neural network model, i.e., optimal neuron weights. When a packet arrives at a switch, the probability of the packet going to each port is proportional to the traffic density of the corresponding link. In other words, this embodiment adjusts the traffic density of each link through the forwarding decisions of each port.
[0174] Therefore, this embodiment defines a forwarding probability based on the mapping between link traffic density and neuron weight. See formula (7), the probability P i It may be a floating point number. However, the programmable data plane does not support division + and floating point data very well. Therefore, the present invention uses the random function that can achieve uniform distribution in the programmable data plane to implement the probabilistic forwarding described in formula (7). That is, the forwarding port is selected in a grid-selecting manner. As shown in Figure 4, the weight of port i is determined by W i grids represent, there are ∑W i grids, numbered 1, 2, 3... ∑W i , the random function randomly generates a i ] range, the data packet will be forwarded from the port to which the grid belongs.
[0175] To overcome variations in network topology, DPN utilizes model pruning techniques to align the neural network structure with the actual network topology. Using the number of packets per port per unit time as the dataset feature and the optimal end-to-end packet delay as the sample label, it trains a load balancing model tailored to the network topology using the Mean Square Error (MSE) loss function. Each switch probabilistically selects a port to forward packets based on this load balancing model. Experimental results, as shown in Figure 5, show that compared to typical flow control methods such as ECMP, LetFlow, and DRILL, this flow control method based on a distributed neural network reduces flow completion time by 42.5%, 35.5%, and 59.8%, respectively.
[0176] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0177] Based on the same concept as the flow control method of the distributed on-line neural network based on the programmable data plane in the above-mentioned embodiment, the present invention also provides a flow control system of the distributed on-line neural network based on the programmable data plane, which can be used to execute the flow control method of the distributed on-line neural network based on the programmable data plane. For ease of explanation, the structural diagram of the flow control system embodiment of the distributed on-line neural network based on the programmable data plane only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and it may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0178] Referring to FIG6 , in another embodiment of the present application, a flow control system 10 of a distributed on-line neural network based on a programmable data plane is provided. The system includes a network construction module 11 , a model pruning module 12 , a network training module 13 , and a flow control module 14 ;
[0179] The network construction module 11 is used to construct a distributed on-line neural network, which includes multiple neurons. The neurons are formed by programmable switches, and the programmable switches use a source routing method that actively sends data packets to determine neuron communication links;
[0180] The model pruning module 12 is used to prune the model of the distributed on-line neural network according to the requirements of constructing distributed on-line neural networks with different structures under different network topologies; the model pruning is at least one of pruning neurons and pruning weight connections between neurons;
[0181] The network training module 13 is used to iteratively train the distributed on-line neural network using the labeled distributed spatiotemporal network behavior feature dataset, calculate the loss value of the loss function, until the model parameters converge to the label value, and obtain the optimal neuron weight after completing the model training;
[0182] The flow control module 14 is used to establish a mapping relationship between the optimal neuron weight and the network behavior characteristics of the flow control; the optimal neuron weight is converted into the network behavior characteristics of the flow control through the mapping relationship to complete the flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping or routing flow scheduling.
[0183] Preferably, the on-line neural network flow control system 10 of the distributed on-line neural network based on a programmable data plane adopts the on-line neural network flow control method of the distributed on-line neural network based on a programmable data plane of the second embodiment.
[0184] It should be noted that the flow control system of the distributed on-line neural network based on the programmable data plane of the present invention corresponds one-to-one to the flow control method of the distributed on-line neural network based on the programmable data plane of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the flow control method of the distributed on-line neural network based on the programmable data plane are applicable to the embodiment of the flow control method of the distributed on-line neural network based on the programmable data plane. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.
[0185] In addition, in the implementation of the flow control system of the distributed in-network neural network based on the programmable data plane in the above-mentioned embodiment, the logical division of each program module is only an example. In actual application, the above-mentioned functions can be distributed to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the flow control system of the distributed in-network neural network based on the programmable data plane is divided into different program modules to complete all or part of the functions described above.
[0186] Please refer to Figure 7. In one embodiment, an electronic device that implements a flow control method for a distributed on-line neural network based on a programmable data plane is provided. The electronic device 20 may include a first processor 21, a first memory 22 and a bus. It may also include a computer program stored in the first memory 22 and executable on the first processor 21, such as a flow control program 23 for a distributed on-line neural network based on a programmable data plane.
[0187] The first memory 22 includes at least one type of readable storage medium, including flash memory, mobile hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, disk, optical disk, etc. In some embodiments, the first memory 22 can be an internal storage unit of the electronic device 20, such as a mobile hard disk of the electronic device 20. In other embodiments, the first memory 22 can also be an external storage device of the electronic device 20, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 20. Furthermore, the first memory 22 can also include both an internal storage unit of the electronic device 20 and an external storage device. The first memory 22 can not only be used to store application software and various types of data installed in the electronic device 20, such as the code of the flow control program 23 of the distributed neural network based on the programmable data plane, but can also be used to temporarily store data that has been output or is about to be output.
[0188] In some embodiments, the first processor 21 may be composed of an integrated circuit, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The first processor 21 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing or executing programs or modules stored in the first memory 22, as well as calling data stored in the first memory 22, to perform various functions of the electronic device 20 and process data.
[0189] FIG7 only shows an electronic device with components. Those skilled in the art will understand that the structure shown in FIG7 does not constitute a limitation on the electronic device 20 , and may include fewer or more components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0190] The first memory 22 in the electronic device 20 stores a distributed on-line neural network flow control program 23 based on a programmable data plane, which is a combination of multiple instructions. When running in the first processor 21, the program 23 can achieve the following:
[0191] Constructing a distributed on-net neural network, wherein the distributed on-net neural network includes a plurality of neurons, wherein the neurons are formed by a programmable switch, and the programmable switch adopts a source routing method of actively sending data packets to determine a neuron communication link;
[0192] Pruning the distributed neural network model according to the needs of building distributed neural networks with different structures under different network topologies; the model pruning is at least one of pruning neurons and pruning weight connections between neurons;
[0193] The distributed on-line neural network is iteratively trained using a labeled distributed spatiotemporal network behavior feature dataset. The loss value of the loss function is calculated until the model parameters converge to the label value. After the model training is completed, the optimal neuron weights are obtained.
[0194] A mapping relationship is established between the optimal neuron weight and the network behavior characteristics of flow control; the optimal neuron weight is converted into the network behavior characteristics of flow control through the mapping relationship, and the weighted probability forwarding strategy is used to complete the approximately optimal flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping, routing scheduling or flow scheduling.
[0195] Furthermore, if the modules / units integrated in the electronic device 20 are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0196] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0197] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0198] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A flow control method for a distributed neural network based on a programmable data plane, characterized in that: The steps include: Constructing a distributed on-net neural network, the distributed on-net neural network comprising a plurality of neurons, the neurons being formed by a programmable switch, the programmable switch adopting a source routing method of actively sending data packets to determine a neuron communication link; According to the needs of building distributed on-line neural networks with different structures under different network topologies, the distributed on-line neural network is model pruned; the model pruning is at least one of pruning neurons and pruning weight connections between neurons; The distributed on-line neural network is iteratively trained using a labeled distributed spatiotemporal joint network behavior feature dataset, and the loss value of the loss function is calculated until the model parameters converge to the label value. After the model training is completed, the optimal neuron weight is obtained; A mapping relationship is established between the optimal neuron weight and the network behavior characteristics of the flow control; the optimal neuron weight is converted into the network behavior characteristics of the flow control through the mapping relationship, and the weighted probability forwarding strategy is used to complete the approximately optimal flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping, routing scheduling or flow scheduling.
2. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The distributed on-line neural network uses programmable switches to form neurons through matrix vector linear operations, specifically: The matrix vector linear operation is specifically as follows: The output signal x of the i-th neuron in the previous layer i The weight w between the i-th neuron in the previous layer and this neuron i Multiply them together to get w i x i , then sum the outputs of multiple neurons in the upper layer, subtract the activation bias θ of the neurons, and finally obtain the output y after the activation function f as the input of the neurons in the lower layer.
3. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 2, characterized in that: The matrix-vector linear operation is implemented by the programmable switch executing the Map-Reduce mechanism based on the matching-action table MAT to perform equivalent calculations, specifically: The vector-to-vector operation in matrix-vector linear operations is equivalent to the Map operation in Map-Reduce, and the vector-to-scalar operation in matrix-vector linear operations is equivalent to the Reduce operation in Map-Reduce. The Map-Reduce mechanism is located between two MATs of the programmable switch, realizing pipeline operation with MAT, parser, and scheduling.
4. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The programmable switch uses a source routing method of actively sending data packets to determine the neuron communication link, specifically: Determine a source routing path according to the need of model parameters transmitted between neurons of the distributed neural network model; When communicating between neurons, the programmable switch actively sends a data packet with a customized header using the P4 programming language, and then embeds the model parameters into the data packet; the customized header includes the addresses of the switches along the way required for source routing; The data packet is forwarded according to the source routing path determined above, completing the transmission of the distributed neural network model parameters between neurons, ensuring that the delay of communication transmission between neurons is deterministic.
5. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The distributed spatiotemporal network behavior feature dataset with labels is specifically: In the neurons of all input layers of the distributed neural network, the same network behavior feature is extracted in each time slot, and different network behavior features are extracted in different time slots; All neurons in the input layer are detection nodes; The network behavior features of multiple detection nodes in multiple time slots are used to build a labeled distributed spatiotemporal joint network behavior feature dataset.
6. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 5, characterized in that: The construction of a distributed spatiotemporal network behavior feature dataset with labels is specifically as follows: All input layer neurons of the distributed neural network extract link flow density from the forwarded packets in each time slot, where the link flow density is the number of packets per unit time; A labeled distributed spatiotemporal network behavior feature dataset is constructed using the link traffic density of multiple key time slots of multiple key switches on the path passed by the packet.
7. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The neuron waits for all its inputs to arrive within a time slot, which is used for synchronization between neurons so that each neuron can process its input before new inputs arrive.
8. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The distributed on-line neural network is iteratively trained using a labeled distributed spatiotemporal joint network behavior feature data set. The iterative training requires multiple iterations of forward propagation and backward propagation so that the predicted value of the on-line neural network model converges to the label value. One iteration of forward propagation and backward propagation specifically includes the following steps: The forward propagation converts the network behavior feature data set into a standard numerical value, and then inputs the numerical value in parallel into the input layer of the distributed network neural network, propagates forward layer by layer in the intermediate layer of the distributed network neural network, outputs the predicted value in the output layer, and finally calculates the error between the predicted value and the label value by calculating the loss value of the loss function; The backward propagation propagates the error from the output layer to the input layer layer by layer based on the chain rule, and in this process updates the weights and biases of the distributed neural network.
9. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The method of completing the approximately optimal flow control of the data center network specifically includes the following steps: The data center network adopts the ECMP routing method to forward packets. The programmable switch collects network behavior characteristics once in each time slot and stores them in the register of the programmable switch. The network controller collects data from the register every time T to build a distributed space-time joint network behavior characteristic data set with labels. Every time T, the distributed space-time joint network behavior characteristic data set with labels adds a set of data. Using the distributed spatiotemporal network behavior feature dataset with labels, the distributed on-line neural network is iteratively trained. When the sliding average of the loss value decreases to a predetermined threshold, the training is stopped. In the network controller, each neuron weight is obtained from the trained neural network model and sent to the programmable switch; The programmable switch receives the neuron weights sent down and stores them in its own registers; When a packet arrives at a programmable switch, the programmable switch uses a weighted probabilistic forwarding strategy to select a forwarding port for the packet.
10. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The optimization objectives of the label flow control are specifically: end-to-end delay, flow completion time and deadline satisfaction rate when the network is completely free of congestion.
11. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The mapping relationship between the neuron weight and the network behavior characteristics of the flow control is established as follows: The mapping relationship between neuron weights and network behavior characteristics of flow control is an equal relationship; The link flow density of the links between programmable switches is a network behavior characteristic of the flow control, and the value of the link flow density of the links between programmable switches is equal to the value of the neuron weight.
12. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The loss function in the distributed neural network training is: Among them, MSE represents the mean square error function, represents the predicted end-to-end delay, flow completion time, and deadline satisfaction rate of the ith packet in the data center network, D i represents the corresponding i-th label, and n represents the number of features.
13. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 1, characterized in that: The weighted probability forwarding strategy is specifically as follows: according to the mapping between link traffic density and neuron weight, the probability of a packet being forwarded by the i-th switch port is: Among them, W i is the weight of the i-th switch port, P i is the probability that port i is selected, and the weight of the switch port is the neuron weight of the corresponding neuron.
14. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 9, characterized in that: Under the K=4 fat-tree network topology, the distributed spatiotemporal network behavior feature dataset with labels includes the following steps: Set the link traffic density of ports 3 and 4 of the programmable switch; At the tth time slot, a sample without spatiotemporal joint correlation characteristics of the data set used to train the in-network neural network is obtained as follows: in, represents the link traffic density of the i-th port of the n-th switch at the t-th time slot.
15. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 9, characterized in that: Under the K=4 fat-tree network topology, the acquisition of the labeled distributed spatiotemporal joint network behavior feature data set also includes the following steps: Set the link traffic density of ports 3 and 4 of the programmable switch; Set the input of the first neuron in the input layer of the neural network Set the input of the second neuron in the input layer Set the input I of the nth neuron in the input layer n (t+n-1) is: At the tth time slot, a sample with time-related characteristics of the data set used to train the neural network on the network is obtained, as follows: Among them, I n (t+n-1) represents the input of the nth neuron in the input layer of the neural network at the tth time slot, represents the link traffic density of the i-th port of the n-th switch at the t-th time slot.
16. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 9, characterized in that: Under the K=4 fat-tree network topology, the acquisition of the labeled distributed spatiotemporal joint network behavior feature data set also includes the following steps: Set the link traffic density of ports 3 and 4 of the programmable switch; The programmable switches s1, s9, and s17 are used as the input of the first neuron in the input layer of the neural network, as shown below: The programmable switches s2, s10, and s19 are used as the input of the second neuron in the input layer of the neural network, as shown below: The programmable switches s3, s11, and s18 are used as the input of the third neuron in the input layer of the neural network. The programmable switches s4, s12, and s20 are used as the input of the fourth neuron in the input layer of the neural network. At the tth time slot, a sample with spatiotemporal joint correlation characteristics of the data set used to train the on-line neural network is obtained, as follows: Among them, I n (t) represents the input of the nth neuron in the input layer of the neural network at the tth time slot, represents the link traffic density of the i-th port of the n-th switch at the t-th time slot.
17. The method for controlling the flow of a distributed neural network based on a programmable data plane according to claim 9, characterized in that: Under the K=4 fat-tree network topology, the specific method of setting the end-to-end delay as a label of the distributed spatiotemporal joint network behavior feature data set is: When the network is completely free of congestion, at the tth time slot, the average end-to-end delays of packets between switches s1 and s5, s2 and s6, s3 and s7, and s4 and s8 are D s1-s5 (t), D s2-s6 (t), D s3-s7 (t), D s4-s8 (t); At the tth time slot, the label D(t) of the sample is set to: D(t)=[D s1-s5 (t),D s2-s6 (t),D s3-s7 (t),D s4-s8 (t)] Then, in the tth time slot, a labeled sample in the aforementioned distributed spatiotemporal joint network behavior feature dataset is C1(t)-D(t), C2(t)-D(t), C3(t)-D(t); C1(t) refers to a sample without spatiotemporal joint correlation characteristics, C2(t) refers to a sample with time correlation characteristics, and C3(t) refers to a sample with spatiotemporal joint correlation characteristics; A distributed spatiotemporal network behavior feature dataset with labels is constructed using the multiple labeled samples.
18. A flow control system based on a distributed neural network with a programmable data plane, characterized in that: A flow control method for a distributed on-line neural network based on a programmable data plane applied to any one of claims 1-17, comprising an on-line neural network construction module, a model pruning module, an on-line neural network training module and a flow control module; The on-line neural network construction module is used to construct a distributed on-line neural network, wherein the distributed on-line neural network includes a plurality of neurons, wherein the neurons are formed by a programmable switch, and the programmable switch adopts a source routing method of actively sending data packets to determine a neuron communication link; The model pruning module is used to prune the model of the distributed on-line neural network according to the needs of building distributed on-line neural networks with different structures under different network topologies; the model pruning is at least one of pruning neurons and pruning weight connections between neurons; The network training module is used to iteratively train the distributed on-line neural network using the distributed spatiotemporal network behavior feature data set with labels, calculate the loss value of the loss function, until the model parameters converge to the label value, and obtain the optimal neuron weight after completing the model training; The flow control module is used to establish a mapping relationship between the optimal neuron weight and the network behavior characteristics of the flow control; the optimal neuron weight is converted into the network behavior characteristics of the flow control through the mapping relationship, and the weighted probability forwarding strategy is used to complete the approximately optimal flow control of the data center network; the flow control includes load balancing, congestion control, traffic shaping, routing scheduling or flow scheduling.
19. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the traffic control method of a distributed in-network neural network based on a programmable data plane as described in any one of claims 1-17.
20. A computer-readable storage medium, characterized in that: It includes computer-readable instructions, which, when executed on a computing device or a computing device cluster, enable the computing device or the computing device cluster to execute the traffic control method of a distributed on-line neural network based on a programmable data plane as described in any one of claims 1 to 17.
Citation Information
Patent Citations
Network traffic feature extraction method and device, equipment and storage medium
CN114374655A
Flow control method and device for distributed in-network neural network based on programmable data plane
CN117834534A
Fusion neuron model, neural network structure and training and inference methods therefor, storage medium, and device
WO2022134391A1
Cited By
Maximum flow scheduling method and device based on time-varying synchronous flow neural network
CN120725076A