Intelligent calculation fusion network co-flow scheduling method oriented to federated learning large model synchronization

By constructing a resource adapter and a deep reinforcement learning training framework in the intelligent computing fusion network, bandwidth and priority policies are generated, which solves the problems of insufficient computation completion time and insufficient network link status awareness, and improves parameter synchronization efficiency and training efficiency.

CN120880993APending Publication Date: 2025-10-31BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511114840.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies cannot effectively perceive computation completion time and dynamic network link status in intelligent computing convergence networks, resulting in low parameter synchronization efficiency, and existing scheduling strategies cannot balance the computation network latency of each training node.

Method used

A smart computing converged network resource adapter is constructed. By establishing a network communication tunnel to transmit parameter data packets, dynamic network load status information is collected and summarized. A co-flow scheduling decision neural network is used to generate bandwidth and priority strategies, and a deep reinforcement learning training framework is combined to optimize the scheduling strategy.

Benefits of technology

This technology effectively reduces the latency of high-performance nodes waiting for low-performance nodes in intelligent computing fusion networks, improves parameter synchronization efficiency, and reduces overall training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880993A_ABST
    Figure CN120880993A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent calculation fusion network co-flow scheduling method oriented to federated learning large model synchronization. The method comprises the following steps of: adding dynamic network load state information of equipment and network flow transmission time of data packet access equipment into a parameter data packet by an intelligent calculation fusion network component; the parameter synchronization platform summarizes the collected dynamic network load state information and network flow transmission time of each intelligent calculation fusion network component; the intelligent computing fusion network resource adapter inputs the dynamic network load state information and the network flow transmission time of each intelligent computing fusion network component into a trained co-flow scheduling decision neural network, generates a co-flow scheduling strategy of bandwidth and priority, and issues the co-flow scheduling strategy to the intelligent computing fusion network component; and the intelligent computing fusion network component reserves resources locally according to the co-flow scheduling strategy. According to the method, the mechanism can be called to effectively reduce the time delay of waiting for low-performance nodes by high-performance nodes, the parameter synchronization efficiency is improved, and the overall training time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network co-flow scheduling technology, and in particular to a method for intelligent computing fusion network co-flow scheduling for large-scale model synchronization in federated learning. Background Technology

[0002] The release of the DeepSeek-R1 large language model has attracted widespread attention due to its "low cost and high efficiency." This research achievement brings tremendous opportunities for the practical application and development of lightweight large model technology, and provides solid and reliable technical support for the local deployment of large language models in high-security fields such as industry, healthcare, law, finance, and government offices. The use of a federated learning training framework can integrate heterogeneous computing resources while protecting data security, significantly improving the training efficiency of federated learning large language models, which is currently a hot topic in the industry.

[0003] However, heterogeneous computing resources can lead to significant differences in computation completion times among multiple training nodes. If computation completion times could be sensed at the network side and flexible co-flow scheduling implemented, waiting time could be reduced and synchronization efficiency improved. However, training nodes often transmit data via wide area networks (WANs), and the high dynamism of network links makes it difficult for the controller to achieve optimal scheduling. Therefore, a flexible co-flow scheduling mechanism that can sense computation completion times and adapt to dynamic network link states is needed to balance the computation-network latency of each training node.

[0004] However, existing network technologies lack computational awareness and dynamic adaptation capabilities. The proposed theory of intelligent computing convergence networks effectively addresses these shortcomings. Intelligent computing convergence networks propose a "three-layer, three-domain" architecture. The controller in the resource adaptation layer can support the perception of various network information and, based on the network status stored in the knowledge domain and the intelligent scheduling algorithm for operation, achieves flexible scheduling of network resources, effectively meeting diverse service needs.

[0005] One existing approach proposes a co-flow scheduling algorithm based on deep reinforcement learning for data center networks. This method assumes that the data center network is an abstract switching structure, allocating bandwidth resources and priorities at the inlet and outlet respectively. Although the method considers the interrelationship of parameters synchronization flows of different training nodes for the same training task during co-flow transmission in its reward design, its assumption of an abstract switching structure does not take into account the dynamic load state between network nodes, and therefore it is not suitable for the co-flow scheduling problem in intelligent computing converged network scenarios.

[0006] One drawback of the aforementioned prior art is that while it achieves awareness of the local computation completion time and the synchronization completion time of other parameters within the co-flow group, its method is only applicable to data center networks where network link states are relatively static. In intelligent computing converged networks, highly dynamic network links are prone to congestion and packet loss due to excessive instantaneous traffic. Using this method, a parameter synchronization flow is highly likely to be assigned to a sudden high-load link, resulting in an excessively long flow completion time and affecting parameter synchronization efficiency.

[0007] Another existing approach addresses the dynamic network link states in wide area networks (WANs) by proposing a deep reinforcement learning method based on a generative diffusion model. This method, during the training of the deep reinforcement learning decision model, progressively transforms the scheduling policy into Gaussian white noise through a forward diffusion process, and then performs a reverse policy generation process based on a neural network. By integrating the dynamic network link states at multiple time points, the Gaussian white noise is gradually restored to an optimal scheduling policy. While this method can effectively extract dynamic network link load information during the decision-making process, it fails to achieve optimal decision-making due to the lack of awareness of the differentiated calculation completion times between synchronous traffic flows with the same set of parameters.

[0008] Another drawback of the above-mentioned prior art is that while the scheme achieves dynamic generative decision-making for a single flow, its scheduling strategy is only for a single flow and cannot perceive the local computation completion time and the completion time of other parameter synchronization flows within the co-flow group. The scheduling strategy generated according to this method is very likely to give high priority and high bandwidth to parameter synchronization flows with low computation completion time, thus crowding out precious intelligent computing fusion network resources. Summary of the Invention

[0009] Embodiments of the present invention provide a co-flow scheduling method for intelligent computing fusion networks oriented towards the synchronization of large models in federated learning, so as to effectively improve the transmission efficiency of parameter streams in intelligent computing fusion networks.

[0010] To achieve the above objectives, the present invention adopts the following technical solution.

[0011] A method for co-current scheduling of intelligent computing fusion networks for synchronizing large federated learning models includes constructing an intelligent computing fusion network comprising multiple edge nodes, a parameter synchronization platform, an intelligent computing fusion network resource adapter, and intelligent computing fusion network components, wherein the intelligent computing fusion network components include network communication tunnels and routers. The method includes:

[0012] The intelligent computing converged network resource adapter establishes network communication tunnels for each edge node, parameter synchronization platform and intelligent computing converged network component in the intelligent computing converged network with the lowest bandwidth and priority allocation;

[0013] After local training is completed at a certain edge node, the local computation completion time and network stream transmission time are added to the parameter data packet. The parameter data packet is then uploaded to the parameter synchronization platform through the network communication tunnel. Each intelligent computing fusion network component that the parameter data packet passes through during transmission also adds the device's dynamic network load status information and the network stream transmission time of the data packet entering and leaving the device to the parameter data packet.

[0014] After the parameter synchronization platform receives the parameter data packet, it will summarize the dynamic network load status information and network flow transmission time of each intelligent computing converged network component and send the summarized information to the intelligent computing converged network resource adapter.

[0015] The intelligent computing converged network resource adapter inputs the dynamic network load status information and network flow transmission time of each intelligent computing converged network component into the trained co-flow scheduling decision neural network. The co-flow scheduling decision neural network adopts a multi-layer perceptron structure to generate a co-flow scheduling strategy with bandwidth and priority, and distributes the co-flow scheduling strategy to the intelligent computing converged network components participating in parameter data packet transmission. The intelligent computing converged network components reserve resources locally according to the co-flow scheduling strategy and update the relevant forwarding flow table.

[0016] Preferably, the intelligent computing converged network resource adapter establishes network communication tunnels for each edge node, parameter synchronization platform, and intelligent computing converged network component in the intelligent computing converged network with minimum bandwidth and priority allocation, including:

[0017] The intelligent computing converged network resource adapter establishes SRv6 network communication tunnels for each edge node, parameter synchronization platform, and intelligent computing converged network component in the intelligent computing converged network with the lowest bandwidth and priority allocation. These network communication tunnels include those between edge nodes and parameter synchronization platforms, between edge nodes and intelligent computing converged network components, between different intelligent computing converged network components, and between parameter synchronization platforms and intelligent computing converged network resource adapters. The parameter synchronization platform distributes the initialization training model to the edge nodes participating in training through the network communication tunnels.

[0018] Preferably, the method further includes:

[0019] A deep reinforcement learning training framework for co-flow group perception is constructed. The co-flow scheduling decision neural network is trained using the deep reinforcement learning training framework through a group perception trajectory collection mechanism and a training mechanism based on a dual evaluation network. The group perception trajectory collection mechanism distinguishes the collected sample data according to the co-flow group. The training mechanism based on the dual evaluation network completes the training using the grouped sample data and outputs the co-flow scheduling strategy.

[0020] During the training process of the co-flow scheduling decision neural network, the group-aware trajectory collection mechanism uses different co-flow group identifiers as a basis. Each training instance serving the same training task has a parameter synchronization stream. These parameter synchronization streams will be marked with the same identifier during transmission and together form a co-flow group. Based on this, the collected environment-action interaction sample trajectories are classified. Each sample trajectory τ includes {state s, action a, policy probability density π, reward}. State s includes network latency state and network load state. Action a is the scheduling action taken by the current parameter synchronization stream. Policy probability density π is the selection probability of each predetermined bandwidth and priority scheduling policy. Reward is the ratio of the sum of the current stream's computation and network completion time to the maximum completion time of the co-flow group. G is the number of parameter synchronization streams in the current co-flow group.

[0021] After distinguishing the trajectories collected by each co-flow group, the maximum completion time within the co-flow group is updated, and the co-flow group reward is calculated according to equation (1).

[0022]

[0023] The training mechanism based on the dual-evaluation network includes a dual-evaluation network, a policy network, a target network, and an Adam optimizer. The dual-evaluation network consists of two evaluation networks with identical structures but independent parameters, each performing Q-value estimation on the same state-action pair. The Q-value represents the expected cumulative reward obtained by taking an action under a given co-current scheduling policy that allocates bandwidth and priority given the current network load, current network completion time, and maximum network completion time. The policy network is a neural network model that outputs action decisions; it directly represents the agent's policy function, guiding the action taken in a given state. The target network is a copy of the neural network used to calculate the target value during deep reinforcement learning model training. The Adam optimizer is a gradient optimization algorithm used to update the parameters of each neural network, incorporating momentum and adaptive learning rate mechanisms to update the parameters of the dual-evaluation network, policy network, and target network.

[0024] Preferably, the intelligent computing converged network resource adapter inputs the dynamic network load status information and network flow transmission time of each intelligent computing converged network component into a trained co-flow scheduling decision neural network. This co-flow scheduling decision neural network uses a multilayer perceptron structure to generate a bandwidth and priority co-flow scheduling strategy, which is then distributed to the intelligent computing converged network components participating in parameter data packet transmission. The intelligent computing converged network components reserve resources locally according to the co-flow scheduling strategy and update relevant forwarding flow tables, including:

[0025] Step 1: Select a co-flow group transmitted on the intelligent computing fusion network component, and generate an initial noise action x based on the normal distribution probability density with a mean of 0 and a variance of 1 for the parameter synchronization traffic in any co-flow group. G ;

[0026] Step 2: Set the initial noise action x G The network state s and the number of diffusion generation steps g are input to a policy neural network based on a multilayer perceptron. After the activation calculation of each neuron function in each layer, the output is the action x for the next diffusion generation step. g―1 ;

[0027] Step 3: Repeat step 2 G times, output the probability density factor generated by diffusion, and calculate the final output action probability density π based on equation (2). θ The output action a is sampled. In equation (2), pri represents a position in the eight priority action dimensions, and x0 is the probability density of each action dimension output by the last diffusion generation step. The probability density is normalized using the softmax function to obtain the normalized probability density value of each action dimension, and the final output action probability density vector π is generated. θ Each action dimension has a 0 / 1 variable, indicating whether the parameter synchronization stream is assigned the corresponding priority;

[0028]

[0029] The intelligent computing converged network component participating in forwarding the parameter synchronization flow receives a co-flow scheduling policy that includes bandwidth and priority allocation information on how to forward the parameter synchronization flow. The intelligent computing converged network component reserves resources locally according to the co-flow scheduling policy and updates the relevant forwarding flow table.

[0030] As can be seen from the technical solutions provided by the embodiments of the present invention above, the method of the present invention designs a co-flow scheduling mechanism in the resource adaptation layer of the intelligent computing fusion network that can sense the computation completion time and network link status, and proposes a co-flow scheduling algorithm based on generative reinforcement learning in the knowledge domain. When the intelligent computing fusion network processes the synchronization traffic of federated learning large models of different training nodes in the same training task, it can call this mechanism to effectively reduce the latency of high-performance nodes waiting for low-performance nodes, improve parameter synchronization efficiency, and reduce the overall training time.

[0031] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a schematic diagram of a co-flow scheduling architecture based on an intelligent computing fusion network proposed in an embodiment of the present invention;

[0034] Figure 2 A flowchart illustrating a collaborative scheduling method for intelligent computing fusion networks aimed at synchronizing large models in federated learning, provided as an embodiment of the present invention.

[0035] Figure 3 This is a schematic diagram of a co-flow scheduling architecture based on an intelligent computing fusion network, provided as an embodiment of the present invention. Detailed Implementation

[0036] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0037] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.

[0038] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0039] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0040] This invention first establishes a co-flow scheduling architecture based on intelligent computing fusion networks, aiming to achieve co-flow balanced scheduling with computing network latency awareness; secondly, it proposes a deep reinforcement learning training framework with co-flow group awareness, thereby efficiently extracting the computing network latency relationship of a set of co-flows when training the neural network for co-flow scheduling policies; finally, it designs a policy generation mechanism based on a generative diffusion model, which dynamically adjusts bandwidth and priority allocation by integrating multiple co-flow states when outputting co-flow scheduling policies.

[0041] This invention proposes a co-flow scheduling architecture based on an intelligent computing fusion network, as shown in the embodiments below. Figure 1 As shown, it includes multiple edge nodes, a parameter synchronization platform, and an intelligent computing converged network resource adapter, as well as intelligent computing converged network components such as network communication tunnels and routers. Based on Figure 1 The co-flow scheduling architecture shown in this invention achieves network latency awareness and co-flow scheduling decision-making through an intelligent computing fusion network resource adapter. The processing flow of a co-flow scheduling method for intelligent computing fusion networks oriented towards large-scale model synchronization in federated learning, proposed in this embodiment of the invention, is as follows: Figure 2 As shown, the processing steps include the following:

[0042] Step S10: The intelligent computing converged network resource adapter establishes SRv6 (Segment Routing over IPv6) + EVPN (Ethernet Virtual Private Network) network communication tunnels for each edge node and parameter synchronization platform in the intelligent computing converged network, allocating the lowest bandwidth and priority. Network communication tunnels are also established between each edge node and the parameter synchronization platform, and between the parameter synchronization platform and the intelligent computing converged network resource adapter. These network communication tunnels and routers are components of the intelligent computing converged network. Specifically, such as... Figure 1 In the scenario shown, the corresponding network communication tunnels for edge nodes {1,2,3} are {R1,R4,R5}, {R2,R5}, and {R3,R6,R5}. The parameter synchronization platform distributes the initialization training model to the edge nodes participating in the training through these network communication tunnels.

[0043] Step S20: After completing local training using the initialization training model at a certain edge node, the local computation completion time and network stream transmission time are added to the parameter data packet. The parameter data packet is then uploaded to the parameter synchronization platform via the network communication tunnel. As the parameter data packet is transmitted in the network communication tunnel, each intelligent computing fusion network component that the parameter data packet passes through also adds its dynamic network load status information and the network stream transmission time of the data packet entering and leaving the device to the parameter data packet.

[0044] Step S30: After the parameter synchronization platform receives the parameter data packets uploaded by each edge node, it will statistically analyze the computation completion time and network stream transmission time of each edge node. Simultaneously, it will also summarize the dynamic network load status information and network stream transmission time collected from each intelligent computing converged network component. Then, all the acquired information will be sent to the intelligent computing converged network resource adapter.

[0045] Step S40: After receiving the parameter data packet, the intelligent computing converged network resource adapter initiates a round of full network load status collection based on SNMP (Simple Network Management Protocol), counts the dynamic network load status information and network flow transmission time of all intelligent computing converged network components that participated in and did not participate in the parameter data packet transmission, and timestamps each intelligent computing converged network component.

[0046] Step S50: The intelligent computing converged network resource adapter inputs the dynamic network load status information and network flow transmission time of each intelligent computing converged network component into the trained co-flow scheduling decision neural network. This co-flow scheduling decision neural network adopts a multilayer perceptron structure and intelligently generates co-flow scheduling strategies based on bandwidth and priority. These co-flow scheduling strategies are distributed to all intelligent computing converged network components participating in parameter data packet transmission. The intelligent computing converged network components reserve resources locally according to the co-flow scheduling strategies and update the relevant forwarding flow tables. Before the next round of global parameter synchronization, the allocated co-flow scheduling strategies will take effect, flexibly allocating bandwidth and scheduling priority queues for each parameter synchronization traffic.

[0047] To train the aforementioned co-flow scheduling decision neural network and effectively extract the relationships between co-flow groups, this invention proposes a method such as... Figure 3The diagram illustrates a deep reinforcement learning training framework for co-flow group awareness. During training, the network to be trained outputs a co-flow scheduling strategy for the current network flow and records the reward value based on the input network state, current network completion time, and maximum network completion time, thereby acquiring sample data for training. The process of training the aforementioned co-flow scheduling decision neural network using the deep reinforcement learning training framework includes a group-aware trajectory collection mechanism and a training mechanism based on a dual-evaluation network. The group-aware trajectory collection mechanism distinguishes the collected sample data according to co-flow groups, while the training mechanism based on the dual-evaluation network uses the grouped sample data to complete the training and ultimately outputs a flexible co-flow scheduling strategy.

[0048] The group-aware trajectory collection mechanism aims to classify the collected environment-action interaction sample trajectories based on different coflow group identifiers. Each training instance serving the same training task has a parameter synchronization stream, and these parameter synchronization streams are tagged with the same identifier during transmission, collectively forming a coflow group. For example,... Figure 3 As shown, the probability density distribution of the co-flow scheduling action strategy for edge nodes {1, ec, EC} is {a1, a2, a3}, and the trajectory samples of different iteration rounds are distinguished into {G} according to the co-flow identifier. a G b Each trajectory G includes sample trajectories τ of edge nodes {1, ec, EC}, and each sample trajectory τ includes {state s, action a, policy probability density π, reward}. State s refers to the network latency state and network load state mentioned in the previous section, including the input network state, current network completion time, and maximum network completion time information. Action a is the scheduling action taken by the current parameter synchronization flow, which is a multi-dimensional vector with only one dimension of 1 and other dimensions of 0. The policy probability density π is the selection probability of each predetermined bandwidth and priority scheduling policy, and the reward is the ratio of the sum of the current flow calculation and network completion time to the maximum completion time of the co-flow group. G is the number of parameter synchronization flows in the current co-flow group. After distinguishing the trajectories collected by each co-flow group, the group-aware trajectory collection module will update the maximum completion time in the co-flow group and calculate the co-flow group reward according to equation (1). In equation (1), the difference between the network completion time of the current parameter synchronization flow after executing the assigned scheduling decision and the network completion time of other parameter synchronization flows in the co-flow group is measured by subtracting the average reward within the co-flow group from the current reward and dividing by the reward variance within the co-flow group.

[0049]

[0050] The training mechanism based on the dual-evaluation network comprises four parts: the dual-evaluation network, the policy network, the target network, and the Adam optimizer. The dual-evaluation network consists of two evaluation networks with identical structures but independent parameters. Each network estimates the Q-value of the same state-action pair. The Q-value represents the expected cumulative reward for taking the next action under a given co-current scheduling policy, considering the bandwidth and priority allocation given the current network load, current network completion time, and maximum network completion time. During training, the dual-evaluation network combines the outputs of the two networks by taking the minimum value, thus providing a more stable and conservative value estimate. The policy network is a neural network model that outputs action decisions; it directly represents the agent's policy function, guiding the action taken in a given state. The target network is a copy of the neural network used to calculate the target value during deep reinforcement learning model training. Its structure is the same as the evaluation and policy networks, but its parameter updates are slower or synchronized with a delay, stabilizing the target function calculation during training. The Adam optimizer is a gradient optimization algorithm for updating the neural networks, incorporating momentum and adaptive learning rate mechanisms to achieve stable and efficient parameter updates for the dual-evaluation network, policy network, and target network.

[0051] The steps of the above-mentioned co-flow scheduling decision neural network to intelligently generate the co-flow scheduling strategy of bandwidth and priority are as follows:

[0052] Step 1: Select a co-flow group transmitted on the intelligent computing fusion network component, and generate an initial noise action x based on the normal distribution probability density with a mean of 0 and a variance of 1 for the parameter synchronization traffic in any co-flow group. G .

[0053] Step 2: Set the initial noise action x G The network state s (network load status information, current network completion time status information, maximum network completion time status information) and the number of diffusion generation steps g are input into a policy neural network based on a multilayer perceptron. After the activation calculation of each neuron function in each layer, the output is the action x for the next diffusion generation step. g―1 .

[0054] Step 3: Repeat step 2 G times, output the probability density factor generated by diffusion, and calculate the final output action probability density π based on equation (2). θ And sample the output action a. In equation (2), pri represents a position in the eight priority action dimensions, x0 is the probability density of each action dimension output by the last diffusion generation step, we use the softmax function to normalize the probability density to obtain the normalized probability density value of each action dimension, and generate the final output action probability density vector π. θEach action dimension has a 0 / 1 variable, indicating whether a corresponding priority is assigned to the parameter synchronization stream. Network administrators can preset corresponding priority grading methods according to their own circumstances, with each level representing the allocation of a certain amount of bandwidth and the priority processing order in the forwarding queue.

[0055]

[0056] In this way, for parameter synchronization flows within the same co-flow group forwarded in the intelligent computing converged network component, the generation of the co-flow scheduling strategy considers both the computing network latency information of other flows and the dynamic network state information at each moment. Therefore, the allocation of bandwidth and priority can be flexibly adjusted during the strategy generation process to achieve balanced resource adaptation.

[0057] The intelligent computing converged network component that participates in forwarding the parameter synchronization flow will receive the bandwidth and priority allocation strategy for how to forward the parameter synchronization flow. However, a strategy only affects one parameter synchronization flow. The intelligent computing converged network component will not adopt this strategy for other parameter synchronization flows and non-parameter synchronization traffic.

[0058] As network terminals, edge nodes and parameter synchronization platforms do not require such fine-grained control; the relevant bandwidth allocation and priority allocation strategies only apply to the corresponding network components within the intelligent computing converged network.

[0059] This invention presents a cooperative flow scheduling mechanism for intelligent computing fusion networks designed for large-scale federated learning model synchronization. This mechanism comprehensively considers the complexity of relationships between cooperative flow groups and the dynamic nature of wide area network links. Its unique cooperative flow scheduling architecture based on intelligent computing fusion networks and a deep reinforcement learning training framework for cooperative flow group awareness enable joint bandwidth and priority allocation with computing-network latency as the optimization objective. Furthermore, based on this, a policy generation mechanism using a generative diffusion model can constrain the consistency within cooperative flow groups during multi-step diffusion policy generation sampling. Its advantage lies in incorporating cooperative flow group awareness and diffusion generation mechanisms into the cooperative flow scheduling process, comprehensively considering cooperative flow group relationships and dynamic network links during bandwidth and priority allocation to make optimal decisions.

[0060] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0061] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0062] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0063] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for co-flow scheduling in intelligent computing fusion networks for synchronizing large-scale federated learning models, characterized in that, Constructing an intelligent computing converged network comprising multiple edge nodes, a parameter synchronization platform, an intelligent computing converged network resource adapter, and intelligent computing converged network components, wherein the intelligent computing converged network components include network communication tunnels and routers, the method comprising: The intelligent computing converged network resource adapter establishes network communication tunnels for each edge node, parameter synchronization platform and intelligent computing converged network component in the intelligent computing converged network with the lowest bandwidth and priority allocation; After local training is completed at a certain edge node, the local computation completion time and network stream transmission time are added to the parameter data packet. The parameter data packet is then uploaded to the parameter synchronization platform through the network communication tunnel. Each intelligent computing fusion network component that the parameter data packet passes through during transmission also adds the device's dynamic network load status information and the network stream transmission time of the data packet entering and leaving the device to the parameter data packet. After the parameter synchronization platform receives the parameter data packet, it will summarize the dynamic network load status information and network flow transmission time of each intelligent computing converged network component and send the summarized information to the intelligent computing converged network resource adapter. The intelligent computing converged network resource adapter inputs the dynamic network load status information and network flow transmission time of each intelligent computing converged network component into the trained co-flow scheduling decision neural network. The co-flow scheduling decision neural network adopts a multi-layer perceptron structure to generate a co-flow scheduling strategy with bandwidth and priority, and distributes the co-flow scheduling strategy to the intelligent computing converged network components participating in parameter data packet transmission. The intelligent computing converged network components reserve resources locally according to the co-flow scheduling strategy and update the relevant forwarding flow table.

2. The method according to claim 1, characterized in that, The intelligent computing converged network resource adapter establishes network communication tunnels for each edge node, parameter synchronization platform, and intelligent computing converged network component in the intelligent computing converged network with minimum bandwidth and priority allocation, including: The intelligent computing converged network resource adapter establishes SRv6 network communication tunnels for each edge node, parameter synchronization platform, and intelligent computing converged network component in the intelligent computing converged network with the lowest bandwidth and priority allocation. These network communication tunnels include those between edge nodes and parameter synchronization platforms, between edge nodes and intelligent computing converged network components, between different intelligent computing converged network components, and between parameter synchronization platforms and intelligent computing converged network resource adapters. The parameter synchronization platform distributes the initialization training model to the edge nodes participating in training through the network communication tunnels.

3. The method according to claim 2, characterized in that, The method further includes: A deep reinforcement learning training framework for co-flow group perception is constructed. The co-flow scheduling decision neural network is trained using the deep reinforcement learning training framework through a group perception trajectory collection mechanism and a training mechanism based on a dual evaluation network. The group perception trajectory collection mechanism distinguishes the collected sample data according to the co-flow group. The training mechanism based on the dual evaluation network completes the training using the grouped sample data and outputs the co-flow scheduling strategy. During the training process of the co-flow scheduling decision neural network, the group-aware trajectory collection mechanism uses different co-flow group identifiers as a basis. Each training instance serving the same training task has a parameter synchronization stream. These parameter synchronization streams will be marked with the same identifier during transmission and together form a co-flow group. Based on this, the collected environment-action interaction sample trajectories are classified. Each sample trajectory τ includes {state s, action a, policy probability density π, reward}. State s includes network latency state and network load state. Action a is the scheduling action taken by the current parameter synchronization stream. Policy probability density π is the selection probability of each predetermined bandwidth and priority scheduling policy. Reward is the ratio of the sum of the current stream's computation and network completion time to the maximum completion time of the co-flow group. G is the number of parameter synchronization streams in the current co-flow group. After distinguishing the trajectories collected by each co-flow group, the maximum completion time within the co-flow group is updated, and the co-flow group reward is calculated according to equation (1). The training mechanism based on the dual-evaluation network includes a dual-evaluation network, a policy network, a target network, and an Adam optimizer. The dual-evaluation network consists of two evaluation networks with identical structures but independent parameters, each performing Q-value estimation on the same state-action pair. The Q-value represents the expected cumulative reward obtained by taking an action under a given co-current scheduling policy that allocates bandwidth and priority given the current network load, current network completion time, and maximum network completion time. The policy network is a neural network model that outputs action decisions; it directly represents the agent's policy function, guiding the action taken in a given state. The target network is a copy of the neural network used to calculate the target value during deep reinforcement learning model training. The Adam optimizer is a gradient optimization algorithm used to update the parameters of each neural network, incorporating momentum and adaptive learning rate mechanisms to update the parameters of the dual-evaluation network, policy network, and target network.

4. The method according to claim 3, characterized in that, The intelligent computing converged network resource adapter inputs the dynamic network load status information and network flow transmission time of each intelligent computing converged network component into a trained co-flow scheduling decision neural network. This co-flow scheduling decision neural network uses a multilayer perceptron structure to generate a bandwidth and priority co-flow scheduling strategy, which is then distributed to the intelligent computing converged network components participating in parameter data packet transmission. The intelligent computing converged network components reserve resources locally according to the co-flow scheduling strategy and update relevant forwarding flow tables, including: Step 1: Select a co-flow group transmitted on the intelligent computing fusion network component, and generate an initial noise action x based on the normal distribution probability density with a mean of 0 and a variance of 1 for the parameter synchronization traffic in any co-flow group. G ; Step 2: Set the initial noise action x G The network state s and the number of diffusion generation steps g are input to a policy neural network based on a multilayer perceptron. After the activation calculation of each neuron function in each layer, the output is the action x for the next diffusion generation step. g―1 ; Step 3: Repeat step 2 G times, output the probability density factor generated by diffusion, and calculate the final output action probability density π based on equation (2). θ The output action a is sampled. In equation (2), pri represents a position in the eight priority action dimensions, and x0 is the probability density of each action dimension output by the last diffusion generation step. The probability density is normalized using the softmax function to obtain the normalized probability density value of each action dimension, and the final output action probability density vector π is generated. θ Each action dimension has a 0 / 1 variable, indicating whether the parameter synchronization stream is assigned the corresponding priority; The intelligent computing converged network component participating in forwarding the parameter synchronization flow receives a co-flow scheduling policy that includes bandwidth and priority allocation information on how to forward the parameter synchronization flow. The intelligent computing converged network component reserves resources locally according to the co-flow scheduling policy and updates the relevant forwarding flow table.