Communication method and apparatus

By determining the connection relationships between computing nodes and optimizing parameter transmission time, the problem of low training efficiency of neural network models is solved, and more efficient model training is achieved.

WO2026091978A1PCT designated stage Publication Date: 2026-05-07HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-09-23
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The training efficiency of existing neural network models is low, especially when transmitting parameters between computing nodes, which takes too long and affects the overall training speed.

Method used

By receiving the duration requirements for parameter transmission between computing nodes, the connection relationship between computing nodes is determined based on this duration. This controls the parameter transmission duration within a reasonable range, optimizes the connection method of computing nodes, including available bandwidth and data volume, and dynamically adjusts the connection relationship to improve training efficiency.

Benefits of technology

It effectively controls the total training time of the neural network model, improves the efficiency of model training, and ensures efficient training results under fixed number of rounds and fixed training time per round.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025123203_07052026_PF_FP_ABST
    Figure CN2025123203_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a communication method and apparatus. A first node acquires information of a first duration, the first duration being a duration requirement for transmitting parameters of a first model between computing nodes, or a single-round duration requirement for training the first model, or a total duration requirement for training the first model; and on the basis of the first duration, the first node triggers establishment of a connection between the computing nodes. By determining a connection relationship between the computing nodes by means of the first duration, the duration of transmitting the parameters of the first model between the computing nodes can be controlled within a reasonable range, or a single-round duration of model training can be controlled within a reasonable range, or a total duration of model training can be controlled within a reasonable range, so that the overall training duration can be reduced to a certain extent, and the model training efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

A communication method and apparatus

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411514653.1, filed on October 28, 2024, entitled "A Communication Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of communication technology, and in particular to a communication method and apparatus. Background Technology

[0004] To improve the intelligence and automation of networks, artificial intelligence (AI) and machine learning (ML) technologies are being widely applied. AI and ML technologies involve the training of neural network models, and with the widespread use of neural network models, improving their training efficiency has become particularly important. Therefore, how to improve the training efficiency of neural network models is a question that needs to be considered. Summary of the Invention

[0005] This application provides a communication method and apparatus for improving the training efficiency of neural network models.

[0006] In a first aspect, this application provides a communication method, which can be executed by a first node, a module applied to the first node (e.g., a chip, chip system, or processor), or a logical node, logical module, or software capable of implementing all or part of the functions of the first node. The method includes: receiving information about a first duration from a second node; wherein the first duration is the time requirement for transmitting parameters of a first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model; based on the first duration, triggering the establishment of a connection between the first computing node and the second computing node; wherein the computing nodes training the first model include at least the first computing node and the second computing node, and the connection between the first computing node and the second computing node is used for transmitting parameters of the first model between the first computing node and the second computing node.

[0007] The total training time of the first model is determined based on the number of training epochs and the training time per epoch of the first model. With a fixed number of training epochs, the shorter the training time per epoch of the first model, the shorter the total training time of the first model, and thus the higher the model training efficiency. The training time per epoch of the first model is determined based on the time required for local training on the computing nodes and the time required for transmitting the parameters of the first model between computing nodes. Given a known model structure and a known time required for local training on the computing nodes, the shorter the time required for transmitting the parameters of the first model between computing nodes, the shorter the total training time of the first model.

[0008] By determining the connection relationship between computing nodes based on the time required for transmitting the parameters of the first model between computing nodes, the transmission time of the first model's parameters between computing nodes can be controlled within a reasonable range, and the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0009] If the connection relationship between computing nodes is determined by the training time requirement of the first model in a single round, the training time of the first model in one round can be controlled within a reasonable range. With a fixed number of training rounds, the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0010] If the connection relationship between computing nodes is determined by the total training time requirement of the first model, the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0011] In one possible implementation, the first node is either the first computing node or the second computing node; or, the first node manages the computing node.

[0012] If the first node is a compute node, the second node can manage compute nodes (including the first node). For example, the compute node is a network element (NE), the second node is an element management system (EMS), and the third node is a network management system (NMS). The second node or compute node can also be called a model training provider (MLT producer).

[0013] If the first node manages the compute nodes, and the second node manages the first node, for example, the second node is NMS, the first node is EMS, and the compute nodes are NE, then the second node can be called a model training consumer (machine learning training, MLT consumer), and the first node can be called a model training producer (MLT producer).

[0014] In one possible implementation, triggering the establishment of a connection between the first computing node and the second computing node based on the first duration includes: triggering the establishment of a connection between the first computing node and the second computing node based on the first duration, the available bandwidth between the first computing node and the second computing node, and the amount of data of the parameters of the first model.

[0015] In one possible implementation, the first duration is the duration requirement for transmitting parameters of the first model between computing nodes; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the first estimated duration for bidirectional transmission of the parameters of the first model between the first computing node and the second computing node satisfies the duration requirement.

[0016] In one possible implementation, the first duration is the single-round duration requirement for training the first model; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the second estimated duration for the first computing node and the second computing node to jointly train the first model for one round satisfies the single-round duration requirement.

[0017] In one possible implementation, the first duration is the total duration required to train the first model; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the third estimated duration for the first computing node and the second computing node to jointly train the first model satisfies the total duration requirement.

[0018] In one possible implementation, triggering the establishment of a connection between the first computing node and the second node based on the first duration includes: triggering the establishment of a connection between the first computing node and the second computing node based on the first duration and a first quantity; wherein the first quantity is used to indicate the upper limit of the number of computing nodes that a computing node is allowed to connect to, and the connection is used to transmit parameters of the first model.

[0019] In one possible implementation, triggering the establishment of a connection between the first computing node and the second computing node based on the first duration includes: determining a first connection relationship between the computing nodes based on the first duration; wherein the first connection relationship is a connection relationship for transmitting parameters of the first model between the computing nodes, and the first connection relationship includes a connection relationship between the first computing node and the second computing node; and triggering the establishment of a connection between the first computing node and the second computing node based on the first connection relationship.

[0020] In one possible implementation, determining the first connection relationship between the computing nodes based on the first duration includes: determining the first connection relationship between the computing nodes based on the first duration and a first number; wherein the first number indicates an upper limit on the number of computing nodes that a computing node is allowed to connect to. In other words, a computing node can establish connections with at most a first number of computing nodes for transmitting parameters of the first model.

[0021] In this implementation, the available bandwidth of the computing nodes can be saved by limiting the number of computing nodes that establish connections.

[0022] In one possible implementation, the first node is the first computing node, and triggering the establishment of a connection between the first computing node and the second computing node includes: storing the identifier of the second computing node in a first list, wherein the first list is used to store the identifiers of computing nodes that transmit parameters of the first model to the first computing node.

[0023] In one possible implementation, the first node is the second computing node, and triggering the establishment of a connection between the first computing node and the second computing node includes: storing the identifier of the first computing node in a second list, wherein the second list is used to store the identifiers of computing nodes that transmit parameters of the first model to the second computing node.

[0024] In one possible implementation, the first node manages the computing nodes; triggering the establishment of a connection between the first computing node and the second computing node includes: sending a first indication message to the first computing node, the first indication message being used to instruct the first computing node and the second computing node to establish a connection for transmitting parameters of the first model; and / or, sending a second indication message to the second computing node, the second indication message being used to instruct the first computing node and the second computing node to establish a connection for transmitting parameters of the first model.

[0025] In one possible implementation, the first node is either the first computing node or the second computing node; the method further includes: receiving a first strategy from the second node, the first strategy being used to instruct the computing node to determine the connection relationship for transmitting parameters of the first model.

[0026] In this implementation, the strategy can be used to flexibly indicate the nodes that determine the connection relationship.

[0027] In one possible implementation, the first node manages the computing nodes; the method further includes: receiving a second strategy from a second node, the second strategy being used to instruct the node managing the computing nodes to determine connection relationships between the computing nodes for transmitting parameters of the first model.

[0028] In this implementation, the strategy can be used to flexibly indicate the nodes that determine the connection relationship.

[0029] In one possible implementation, after triggering the establishment of a connection between the first computing node and the second computing node based on the first duration, the method further includes: determining a second connection relationship between the computing nodes based on the first duration, the second connection relationship being a connection relationship for transmitting parameters of the first model between the computing nodes; and updating the connection between the computing nodes based on the second connection relationship.

[0030] In this implementation, the available bandwidth between computing nodes changes dynamically. At a certain point in time, the connections between some computing nodes may meet the corresponding duration requirements, but may not be able to meet them for a period of time. Dynamically determining the connection relationships between computing nodes in this implementation can reduce the occurrence of situations where the initial duration requirement is not met, thereby improving model training efficiency.

[0031] In one possible implementation, determining the second connection relationship between the computing nodes based on the first duration includes: determining the second connection relationship between the computing nodes based on the first period and the first duration.

[0032] In this implementation, the connection relationships between computing nodes are determined periodically to improve the training efficiency of the model.

[0033] In one possible implementation, the method further includes: receiving one or more of the following information from the second node: information on a first quantity, the first quantity indicating an upper limit on the number of computing nodes that a computing node is allowed to connect to, the connections being used to transmit parameters of the first model; information on a first period, the first period indicating a period for determining connection relationships between the computing nodes, the connection relationships being connection relationships used for transmitting parameters of the first model between the computing nodes; available bandwidth between the computing nodes; and a third connection relationship between the computing nodes, the third connection relationship being a connection relationship used for transmitting parameters of the first model between the computing nodes.

[0034] In one possible implementation, the method further includes sending one or more of the following information to the second node: the first connection relationship, the amount of data on the parameters of the first model transmitted bidirectionally between the computing nodes, or the amount of data on the parameters of the first model transmitted unidirectionally between the computing nodes, the convergence performance of the first model, the relevant duration of the transmission of the parameters of the first model between the computing nodes during the entire model training process, the relevant duration of a single round of model training during the entire model training process, and the total training time of the first model.

[0035] In this implementation, after the second node (e.g., NMS) learns one or more of the above information, it can combine historical network traffic information to optimize the parameters subsequently sent to the first node. The parameters sent to the first node include, but are not limited to, one or more of the following: the time requirement for transmitting parameters of the first model between computing nodes, the time requirement for training the first model in a single round, the total time requirement for training the first model, the upper limit of the number of computing nodes that a computing node is allowed to connect to (the upper limit of the number of neighbors), the strategy for determining the connection relationship of computing nodes (i.e., the connection relationship between computing nodes is determined by the computing node or the node that manages the computing node), and the period for determining the connection relationship of computing nodes.

[0036] Secondly, this application provides a communication method that can be executed by a second node, a module applied to the second node (e.g., a chip, chip system, or processor), or a logical node, logical module, or software capable of implementing all or part of the functions of the second node. The method includes: sending information of a first duration to a first node; wherein the first duration is the time requirement for transmitting parameters of a first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model; the first duration is used by the first node to trigger the establishment of a connection between the first computing node and the second computing node, the computing nodes training the first model include at least the first computing node and the second computing node, and the connection between the first computing node and the second computing node is used for transmitting parameters of the first model between the first computing node and the second computing node.

[0037] The total training time of the first model is determined based on the number of training epochs and the training time per epoch of the first model. With a fixed number of training epochs, the shorter the training time per epoch of the first model, the shorter the total training time of the first model, and thus the higher the model training efficiency. The training time per epoch of the first model is determined based on the time required for local training on the computing nodes and the time required for transmitting the parameters of the first model between computing nodes. Given a known model structure and a known time required for local training on the computing nodes, the shorter the time required for transmitting the parameters of the first model between computing nodes, the shorter the total training time of the first model.

[0038] By determining the connection relationship between computing nodes based on the time required for transmitting the parameters of the first model between computing nodes, the transmission time of the first model's parameters between computing nodes can be controlled within a reasonable range, and the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0039] If the connection relationship between computing nodes is determined by the training time requirement of the first model in a single round, the training time of the first model in one round can be controlled within a reasonable range. With a fixed number of training rounds, the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0040] If the connection relationship between computing nodes is determined by the total training time requirement of the first model, the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0041] In one possible implementation, the first node is either the first computing node or the second computing node; or, the first node manages the computing node.

[0042] If the first node is a compute node, the second node can manage compute nodes (including the first node). For example, the compute node is NE, the second node is EMS, and the third node is NMS. The second node or compute node can also be called a model training provider (MLT producer).

[0043] If the first node manages the compute nodes, and the second node manages the first node, for example, the second node is NMS, the first node is EMS, and the compute nodes are NE, then the second node can be called a Model Training Consumer (MLT consumer), and the first node can be called an MLT producer.

[0044] In one possible implementation, the first node is either the first computing node or the second computing node, with the second node managing the computing node; the method further includes receiving information about the first duration from a third node. For example, the second node is an EMS (Model Training Service) and the third node is an NMS (Model Training Service). The second node can also be called a Model Training Producer (MLT producer), and the third node can be called a Model Training Consumer (MLT consumer).

[0045] In one possible implementation, the first node manages the computing node, and the second node manages the first node; the method further includes determining the first duration. For example, the first node is an EMS, and the second node is an NMS. The first node can also be called a Model Training Provider (MLT producer), and the second node can be called a Model Training Consumer (MLT consumer).

[0046] In one possible implementation, the method further includes sending one or more of the following information to the first node: information on a first quantity, the first quantity being an upper limit on the number of computing nodes that a computing node is allowed to connect to, the connections being used to transmit parameters of the first model; information on a first period, the first period being an indication of the period for determining connection relationships between the computing nodes, the connection relationships being connection relationships used for transmitting parameters of the first model between the computing nodes; available bandwidth between the computing nodes; a first strategy, the first strategy being used to instruct computing nodes to determine connection relationships used for transmitting parameters of the first model; a second strategy, the second strategy being used to instruct nodes managing the computing nodes to determine connection relationships between computing nodes used for transmitting parameters of the first model; and a third connection relationship between the computing nodes, the third connection relationship being a connection relationship used for transmitting parameters of the first model between the computing nodes.

[0047] In this implementation, limiting the number of computing nodes establishing connections conserves available bandwidth. A policy can be issued to flexibly instruct nodes on which connection relationships are determined. Periodically determining the connection relationships between computing nodes improves the training efficiency of the model.

[0048] In one possible implementation, the first node is either the first computing node or the second computing node, with the second node managing the computing node; the method further includes receiving one or more pieces of information from a third node. For example, the second node is an EMS (Model Training Service) and the third node is an NMS (Model Training Service). The second node can also be called a Model Training Producer (MLT producer), and the third node can be called a Model Training Consumer (MLT consumer).

[0049] In one possible implementation, the first node manages the computing node, and the second node manages the first node; the method further includes determining one or more pieces of information. For example, the first node is an EMS, and the second node is an NMS. The first node can also be called a Model Training Provider (MLT producer), and the second node can be called a Model Training Consumer (MLT consumer).

[0050] In one possible implementation, the method further includes receiving one or more of the following information from the first node: a first connection relationship between the computing nodes, the amount of data for bidirectional transmission of the parameters of the first model between the computing nodes, or the amount of data for unidirectional transmission of the parameters of the first model between the computing nodes, the convergence performance of the first model, the relevant duration for transmission of the parameters of the first model between the computing nodes during the entire model training process, the relevant duration for a single round of model training during the entire model training process, and the total training time of the first model.

[0051] In this implementation, after the second node (e.g., NMS) learns one or more of the above information, it can combine historical network traffic information to optimize the parameters subsequently sent to the first node. The parameters sent to the first node include, but are not limited to, one or more of the following: the time requirement for transmitting parameters of the first model between computing nodes, the time requirement for training the first model in a single round, the total time requirement for training the first model, the upper limit of the number of computing nodes that a computing node is allowed to connect to (the upper limit of the number of neighbors), the strategy for determining the connection relationship of computing nodes (i.e., the connection relationship between computing nodes is determined by the computing node or the node that manages the computing node), and the period for determining the connection relationship of computing nodes.

[0052] In one possible implementation, the first node manages the computing node, and the second node manages the first node; the method further includes: determining a third connection relationship based on a first parameter; wherein the first parameter includes one or more of the following: the historical network traffic of the computing node, the first connection relationship between the computing nodes, and the amount of data of the parameters of the first model transmitted unidirectionally or bidirectionally between the computing nodes.

[0053] In one possible implementation, the first node is either the first computing node or the second computing node, and the second node manages the computing nodes; the method further includes: performing aggregation processing on the first connection relationship between the computing nodes, and sending the aggregated connection relationship between the computing nodes to a third node.

[0054] Thirdly, a communication device is provided, which can be the first node described in the first aspect above. The communication device possesses the functions of the first node. The communication device is, for example, a functional module of the first node, such as a baseband device or a chip system. Alternatively, the communication device can be the second node described in the second aspect above. The communication device possesses the functions of the second node. The communication device is, for example, a functional module of the second node, such as a baseband device or a chip system.

[0055] In one optional implementation, the communication device includes a baseband device and a radio frequency device. In another optional implementation, the communication device includes a processing unit (sometimes also called a processing module) and a transceiver unit (sometimes also called a transceiver module). The transceiver unit is capable of transmitting and receiving functions. When the transceiver unit performs the transmitting function, it can be called a transmitting unit (sometimes also called a transmitting module), and when the transceiver unit performs the receiving function, it can be called a receiving unit (sometimes also called a receiving module). The transmitting unit and the receiving unit can be the same functional module, which is called the transceiver unit and can perform both transmitting and receiving functions; or, the transmitting unit and the receiving unit can be different functional modules, and the transceiver unit is a collective term for these functional modules.

[0056] In one possible implementation, the communication device further includes a storage unit (sometimes also called a storage module), and the processing unit is configured to couple with the storage unit and execute programs or instructions in the storage unit to enable the communication device to perform the functions of the first node described in the first aspect above, or to enable the communication device to perform the functions of the second node described in the second aspect above.

[0057] When the communication device is applied to the first node, there are several possible implementations:

[0058] In one possible implementation, the transceiver unit is configured to receive information about a first duration from the second node; wherein the first duration is the time requirement for transmitting parameters of the first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model; the processing unit is configured to trigger the establishment of a connection between the first computing node and the second computing node based on the first duration; wherein the computing nodes for training the first model include at least the first computing node and the second computing node, and the connection between the first computing node and the second computing node is used for transmitting parameters of the first model between the first computing node and the second computing node.

[0059] In one possible implementation, the communication device is either the first computing node or the second computing node; or, the communication device manages the computing node.

[0060] In one possible implementation, when the processing unit triggers the establishment of a connection between the first computing node and the second computing node based on the first duration, it is specifically used to: trigger the establishment of a connection between the first computing node and the second computing node based on the first duration, the available bandwidth between the first computing node and the second computing node, and the amount of data of the parameters of the first model.

[0061] In one possible implementation, the first duration is the time requirement for transmitting parameters of the first model between computing nodes; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the first estimated duration for bidirectional transmission of the parameters of the first model between the first computing node and the second computing node satisfies the duration requirement; or, the first duration is the time requirement for a single round of training the first model; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the second estimated duration for the first computing node and the second computing node to jointly train the first model for one round satisfies the single round duration requirement; or, the first duration is the total time requirement for training the first model; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the third estimated duration for the first computing node and the second computing node to jointly train the first model satisfies the total time requirement.

[0062] In one possible implementation, when the processing unit triggers the establishment of a connection between the first computing node and the second node based on the first duration, it is specifically used to: trigger the establishment of a connection between the first computing node and the second computing node based on the first duration and a first quantity; wherein, the first quantity is used to indicate the upper limit of the number of computing nodes that a computing node is allowed to connect to, and the connection is used to transmit parameters of the first model.

[0063] In one possible implementation, when the processing unit triggers the establishment of a connection between the first computing node and the second computing node, it is specifically configured to: determine a first connection relationship between the computing nodes based on the first duration; wherein the first connection relationship is a connection relationship for transmitting parameters of the first model between the computing nodes, and the first connection relationship includes the connection relationship between the first computing node and the second computing node; and trigger the establishment of a connection between the first computing node and the second computing node based on the first connection relationship.

[0064] In one possible implementation, when the processing unit determines the first connection relationship between the computing nodes based on the first duration, it is specifically configured to: determine the first connection relationship between the computing nodes based on the first duration and a first quantity; wherein the first quantity is used to indicate the upper limit of the number of computing nodes that a computing node is allowed to connect to.

[0065] In one possible implementation, the communication device is the first computing node, and the processing unit, when triggering the establishment of a connection between the first computing node and the second computing node, is specifically configured to: store the identifier of the second computing node in a first list, wherein the first list is used to store the identifiers of computing nodes that transmit parameters of the first model to the first computing node.

[0066] In one possible implementation, the communication device is the second computing node, and the processing unit, when triggering the establishment of a connection between the first computing node and the second computing node, is specifically configured to: store the identifier of the first computing node in a second list, wherein the second list is used to store the identifiers of computing nodes that transmit parameters of the first model to the second computing node.

[0067] In one possible implementation, the communication device manages the computing node; the transceiver unit is further configured to send a first indication message to the first computing node, the first indication message being used to instruct the first computing node and the second computing node to establish a connection for transmitting parameters of the first model; and / or, send a second indication message to the second computing node, the second indication message being used to instruct the first computing node and the second computing node to establish a connection for transmitting parameters of the first model.

[0068] In one possible implementation, the communication device is either the first computing node or the second computing node; the transceiver unit is further configured to receive a first strategy from the second node, the first strategy being configured to instruct the computing node to determine the connection relationship for transmitting parameters of the first model; the communication device manages the computing node; the transceiver unit is further configured to receive a second strategy from the second node, the second strategy being configured to instruct the node managing the computing node to determine the connection relationship between computing nodes for transmitting parameters of the first model.

[0069] In one possible implementation, the processing unit is further configured to determine a second connection relationship between the computing nodes based on the first duration, the second connection relationship being a connection relationship for transmitting parameters of the first model between the computing nodes; and to update the connection between the computing nodes based on the second connection relationship.

[0070] In one possible implementation, when the processing unit determines the second connection relationship between the computing nodes based on the first duration, it is specifically used to: determine the second connection relationship between the computing nodes based on the first period and the first duration.

[0071] In one possible implementation, the transceiver unit is further configured to receive one or more of the following information from the second node: information on a first quantity, the first quantity indicating an upper limit on the number of computing nodes that a computing node is allowed to connect to, the connections being used to transmit parameters of the first model; information on a first period, the first period indicating a period for determining the connection relationship between the computing nodes, the connection relationship being a connection relationship for transmitting parameters of the first model between the computing nodes; available bandwidth between the computing nodes; and a third connection relationship between the computing nodes, the third connection relationship being a connection relationship for transmitting parameters of the first model between the computing nodes.

[0072] In one possible implementation, the transceiver unit is further configured to send one or more of the following information to the second node: the first connection relationship, the amount of data on the parameters of the first model transmitted bidirectionally between the computing nodes, or the amount of data on the parameters of the first model transmitted unidirectionally between the computing nodes, the convergence performance of the first model, the relevant duration of the transmission of the parameters of the first model between the computing nodes during the entire model training process, the relevant duration of a single round of model training during the entire model training process, and the total training time of the first model.

[0073] When the communication device is applied to the second node, there are several possible implementations:

[0074] In one possible implementation, the transceiver unit is configured to send information of a first duration to the first node; wherein the first duration is the time requirement for transmitting parameters of the first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model; the first duration is used by the first node to trigger the establishment of a connection between the first computing node and the second computing node, the computing nodes for training the first model include at least the first computing node and the second computing node, and the connection between the first computing node and the second computing node is used for transmitting parameters of the first model between the first computing node and the second computing node.

[0075] In one possible implementation, the first node is either the first computing node or the second computing node; or, the first node manages the computing node.

[0076] In one possible implementation, the first node is either the first computing node or the second computing node, and the communication device manages the computing node; the transceiver unit is further configured to receive information about the first duration from the third node; or, the first node manages the computing node, and the communication device manages the first node; the processing unit is further configured to determine the first duration.

[0077] In one possible implementation, the transceiver unit is further configured to send one or more of the following information to the first node: information on a first quantity, the first quantity being an upper limit on the number of computing nodes that a computing node is allowed to connect to, the connections being used to transmit parameters of the first model; information on a first period, the first period being an indication of the period for determining the connection relationships between the computing nodes, the connection relationships being connection relationships used for transmitting parameters of the first model between the computing nodes; available bandwidth between the computing nodes; a first strategy, the first strategy being used to instruct computing nodes to determine connection relationships used for transmitting parameters of the first model; a second strategy, the second strategy being used to instruct nodes managing the computing nodes to determine connection relationships between computing nodes used for transmitting parameters of the first model; and a third connection relationship between the computing nodes, the third connection relationship being a connection relationship used for transmitting parameters of the first model between the computing nodes.

[0078] In one possible implementation, the first node is either the first computing node or the second computing node, and the communication device manages the computing node; the transceiver unit is further configured to receive one or more pieces of information from the third node; or, the first node manages the computing node, and the communication device manages the first node; the processing unit is further configured to determine one or more pieces of information.

[0079] In one possible implementation, the transceiver unit is further configured to receive one or more of the following information from the first node: the first connection relationship between the computing nodes, the amount of data of the parameters of the first model transmitted bidirectionally between the computing nodes, or the amount of data of the parameters of the first model transmitted unidirectionally between the computing nodes, the convergence performance of the first model, the relevant duration of the transmission of the parameters of the first model between the computing nodes during the entire model training process, the relevant duration of a single round of model training during the entire model training process, and the total training time of the first model.

[0080] In one possible implementation, the first node manages the computing node, and the communication device manages the first node; the processing unit is further configured to determine a third connection relationship based on a first parameter; wherein the first parameter includes one or more of the following: the historical network traffic of the computing node, the first connection relationship between the computing nodes, and the amount of data of the parameters of the first model transmitted unidirectionally or bidirectionally between the computing nodes.

[0081] In one possible implementation, the first node is either the first computing node or the second computing node, and the communication device manages the computing node; the processing unit is further configured to perform aggregation processing on the first connection relationship between the computing nodes; the transceiver unit is further configured to send the aggregated connection relationship between the computing nodes to the third node.

[0082] Fourthly, a communication device is provided, including an interface circuit and a processor, and optionally, a memory. The memory stores a computer program. The processor is coupled to the memory and the interface circuit. When the processor reads the computer program or instructions, it causes the communication device to execute the method executed by the first node in the first aspect, or to execute the method executed by the second node in the second aspect. For example, the interface circuit is used to receive signals from other communication devices besides the communication device and transmit them to the processor, or to send signals from the processor to other communication devices besides the communication device. The processor, through logic circuits or executable code instructions, implements the method executed by the first node in the first aspect, or implements the method executed by the second node in the second aspect.

[0083] In one possible implementation, the communication device is a chip or a chip system. The chip system may consist of chips or may include chips and other discrete components.

[0084] Fifthly, a communication device is provided, including a processor, and optionally, a memory; the processor and the memory are coupled; the memory is used to store computer programs or instructions; the processor is used to execute part or all of the computer programs or instructions in the memory, and when the part or all of the computer programs or instructions are executed, it is used to implement the function of the first node in the first aspect above, and to implement the function of the second node in the second aspect above.

[0085] In one possible implementation, the apparatus may further include a transceiver for transmitting signals processed by the processor or receiving signals input to the processor. The transceiver may perform the transmitting or receiving actions performed by the first node in the first aspect or the second node in the second aspect.

[0086] In one possible implementation, the processing unit in the third aspect can be implemented by the processor, the storage unit in the third aspect can be implemented by the memory, and the transceiver unit in the third aspect can be implemented by the transceiver.

[0087] In one possible implementation, the communication device is a chip or a chip system. The chip system may consist of chips or may include chips and other discrete components.

[0088] In a sixth aspect, a computer-readable storage medium is provided for storing a computer program or instructions that, when executed, enable the implementation of the methods in the foregoing aspects.

[0089] In a seventh aspect, a chip is provided, comprising a processor that, when executing a computer program or instructions, implements the methods of the first or second aspect described above. Optionally, the chip may further include a memory, and the chip may be composed of chips or may include chips and other discrete devices. The memory is used to store computer programs or instructions.

[0090] Eighthly, a circuit is provided for performing the methods described in the first or second aspect above. This circuit may include chip circuitry. Optionally, the circuit may also be coupled to a memory.

[0091] Ninthly, a computer program product containing instructions is provided that, when run on a computer, enables the methods described in the first or second aspect to be implemented.

[0092] In a tenth aspect, a communication system is provided, comprising a first node as described in the first aspect and a second node as described in the second aspect.

[0093] Eleventhly, a communication system is provided, including the first computing node and the second computing node mentioned in the first aspect above. Attached Figure Description

[0094] Figure 1a is a schematic diagram of the architecture of a communication system provided in an embodiment of this application;

[0095] Figure 1b is a schematic diagram of the architecture of a communication system provided in an embodiment of this application;

[0096] Figure 1c is a schematic diagram of the architecture of a communication system provided in an embodiment of this application;

[0097] Figure 2 is a schematic diagram of a neural network model provided in an embodiment of this application;

[0098] Figure 3a is a schematic diagram of a centralized federated learning system provided in an embodiment of this application;

[0099] Figure 3b is a schematic diagram of a distributed federated learning method provided in an embodiment of this application;

[0100] Figure 4 is a schematic flowchart of a communication method provided in an embodiment of this application;

[0101] Figure 5 is a schematic flowchart of a communication method provided in an embodiment of this application;

[0102] Figure 6 is a schematic flowchart of a communication method provided in an embodiment of this application;

[0103] Figure 7 is a schematic flowchart of a communication method provided in an embodiment of this application;

[0104] Figure 8 is a schematic diagram of a device structure provided in an embodiment of this application;

[0105] Figure 9 is a schematic diagram of a device structure provided in an embodiment of this application. Detailed Implementation

[0106] The technical solution of this application can be applied to various wireless communication systems, including but not limited to fourth-generation (4G) mobile communication technology systems (also known as long term evolution (LTE) systems), fifth-generation (5G) mobile communication technology systems (also known as new radio (NR) systems), or future mobile communication systems, etc., without any specific limitations.

[0107] Furthermore, the technical solutions provided in this application can be applied to device-to-device (D2D) scenarios, such as NR-D2D scenarios, or to vehicle-to-everything (V2X) communication scenarios, such as NR-V2X scenarios. For example, they can be used in fields such as intelligent driving, assisted driving, or intelligent connected vehicles. As another example, the technical solutions provided in this application can also be applied to factory manufacturing scenarios.

[0108] Furthermore, the technical solutions provided in this application can be applied to scenarios including but not limited to: terrestrial cellular communication, non-terrestrial network (NTN), satellite communication, high altitude platform station (HAPS) communication, integrated access and backhaul (IAB) communication, and reconfigurable intelligent surface (RIS) communication.

[0109] Figure 1a illustrates a schematic diagram of a communication system architecture, which includes a network management system (NMS), an element management system (EMS), and network elements (NEs). The NMS is responsible for the operation, management, and maintenance of the network and can manage multiple EMSs. Each EMS can manage one or more NEs. The NMS can also be referred to as a cross-domain management function unit. The EMS can also be called a domain management function network element, a radio access network (RAN) domain management function network element, or a single-domain management system. Network elements typically refer to RAN equipment or functional network elements of the core network (CN). The EMS and NEs can be deployed in a single device; for example, the EMS and NE can be two functional modules within the same device. Alternatively, the EMS and NEs can be deployed in different devices; this embodiment does not specifically limit the deployment in this regard.

[0110] (R)AN equipment is a device that provides wireless communication functions for terminal devices; it is also known as access network equipment. RAN equipment includes, but is not limited to: next-generation base stations (g node B, gNB) in 5G, evolved node B (eNB), radio network controller (RNC), node B (NB), base station controller (BSC), base transceiver station (BTS), home base station (e.g., home evolved node B, or home node B, HNB), baseband unit (BBU), transmitting and receiving point (TRP), transmitting point (TP), mobile switching center, etc. In systems employing different wireless access technologies, the names of devices with base station functions may vary. For example, in 5th generation (5G) systems, they are called RAN or gNB (5G NodeB); in LTE systems, they are called evolved NodeB (eNB or eNodeB); and in 3rd generation (3G) systems, they are called Node B, etc.

[0111] Core network elements include one or more of the following network elements:

[0112] The access management network element (also known as the mobility management network element) is a control plane network element provided by the operator's network. It is responsible for access control and mobility management of terminal devices accessing the operator's network, including functions such as mobility state management, allocation of temporary user identities, authentication, and user management. In 5G communication systems, this access management network element can be an access and mobility management function (AMF) network element. In future communication systems, the access management network element can still be an AMF network element, or it can have other names; this application does not limit its scope.

[0113] The session management network element is primarily responsible for session management in mobile networks, such as session establishment, modification, and release. Specific functions include assigning IP addresses to users and selecting user plane network elements that provide packet forwarding capabilities. In 5G communication systems, this session management network element can be a session management function (SMF) network element. In future communication systems, the session management network element may still be an SMF network element, or it may have other names; this application does not impose any limitations on this.

[0114] User plane network elements are responsible for forwarding and receiving user data in terminal devices. They can receive user data from the data network and transmit it to the terminal device through the access network equipment; user plane network elements can also receive user data from the terminal device through the access network equipment and forward it to the data network. The transmission resources and scheduling functions that provide services to the terminal device in the user plane network element are managed and controlled by the SMF network element. In 5G communication systems, this user plane network element can be a user plane function (UPF) network element. In future communication systems, the user plane network element can still be a UPF network element, or it can have other names; this application does not limit this.

[0115] The data management network element is used for generating authentication credentials, processing user identifiers (such as storing and managing permanent user identities), access control, and managing subscription data. In 5G communication systems, this data management network element can be a unified data management (UDM) network element. In future communication systems, unified data management can still be a UDM network element, or it can have other names; this application does not limit this.

[0116] Data storage network elements are responsible for storing structured data information, including subscription information, policy information, and network data or service data with standard format definitions. In 5G communication systems, data storage network elements can be unified data repositories (UDRs). In future communication systems, network open function network elements can still be UDR network elements, or they can have other names; this application does not limit this.

[0117] The policy control network element primarily supports providing a unified policy framework to control network behavior, providing policy rules to the control layer network functions, and is also responsible for acquiring user subscription information related to policy decisions. In 4G communication systems, this policy control network element can be a policy and charging rules function (PCRF) network element. In 5G communication systems, this policy control network element can be a policy control function (PCF) network element. In future communication systems, the policy control network element can still be a PCF network element, or it can have other names; this application does not limit this.

[0118] The network storage element can be used to provide network element discovery functionality, providing network element information corresponding to the network element type based on requests from other network elements. NRF also provides network element management services, such as network element registration, updating, deregistration, and network element status subscription and push. In 5G communication systems, this network storage element can be a network repository function (NRF) element. In future communication systems, the network storage element can still be an NRF element, or it can have other names; this application does not limit its scope.

[0119] Network exposure function (NEF) network elements can be used to securely expose services and capabilities provided by 3GPP network function equipment to the outside world. In 5G communication systems, NEF network elements can be NEF network elements. In future communication systems, NEF network elements can still be NEF network elements, or they can have other names; this application does not limit this.

[0120] The network slice selection function (NSSF) network element can be used to select the appropriate network slice for the terminal's services. In 5G communication systems, the network slice selection network element can be an NSSF network element. In future communication systems, the network open function network element can still be an NSSF network element, or it can have other names; this application does not limit this.

[0121] Network data analytics (NWDAF) network elements can collect, analyze, and predict data from various network functions (NFs), such as policy control, session management, user plane, access management, and application function elements (through network capability open function elements). In 5G communication systems, NWDAF network elements can be network data analytics functions (NWDAF). In future communication systems, NWDAF network elements may still be NWDAF network elements, or they may have other names; this application does not limit this.

[0122] It is understood that the aforementioned network element or function can be a network component in a hardware device, a software function running on dedicated hardware, or a virtualized function instantiated on a platform (e.g., a cloud platform). One possible implementation is that the aforementioned network element or function can be implemented by a single device, multiple devices working together, or a functional module within a single device; this application does not specifically limit this.

[0123] To improve the intelligence and automation of networks, artificial intelligence (AI) and machine learning (ML) technologies are being applied in more and more fields. The 3rd Generation Partnership Project (3GPP) working group is conducting research on several related topics on network intelligence, such as the lifecycle management of models, model training, simulation, deployment and inference processes.

[0124] Optionally, embodiments of this application define two names: model training provider (also known as machine learning training (MLT) provider) and model training consumer (also known as MLT consumer). The model training provider can provide model training services, and the model training consumer can schedule the model training services provided by the model training provider.

[0125] In the 3GPP network domain, the Network Management System (NMS) acts as the MLT consumer, while the EMS or the NE managed by the EMS acts as the MLT producer. In the ORAN network domain, the Service Management and Orchestration (SMO) function acts as the MLT consumer, and the network elements directly managed by the SMO act as the MLT producer. The SMO's role is similar to that of the NMS, responsible for the operation, management, and maintenance of various network services and orchestration functions. The network elements directly managed by the SMO can be heterogeneous; for example, it can directly manage the EMS, gNB, NWDAF, etc.

[0126] Figure 1b shows a schematic diagram of a communication system architecture. The RAN domain management service (MnS) consumer manages the RAN domain. The RAN domain includes RAN domain management functions and multiple gNBs. The RAN management functions include ML training functions and AI / ML inference functions.

[0127] Figure 1c shows a schematic diagram of a communication system architecture. The RAN domain (domain) is managed by the MnS consumer. The RAN domain includes RAN domain management functions and multiple gNBs. The gNBs include ML training functions and AI / ML inference functions.

[0128] A neural network model is a network model that mimics the behavioral characteristics of animal neural networks. This network model processes information by adjusting the interconnections between a large number of neurons. Figure 2 is a schematic block diagram of a neural network model 100. The neural network model 100 includes n layers of neurons. Each layer of n layers includes one or more neurons, and all neurons in each layer are connected to all neurons in the next layer. Taking the neural network model 100 in Figure 2 as an example, the first layer includes two neurons, each of the second to (n-1)th layers includes three neurons, and the nth layer includes one neuron. Here, n is an integer greater than or equal to 2, and i is an integer less than or equal to n and greater than or equal to 1.

[0129] The process of training a neural network is the process of learning the weights of each neuron in each layer of the neural network model. The ultimate goal is to obtain optimal weights for each neuron in each layer. The following section, with reference to Figure 2, provides a detailed description of one round of model training in the neural network model 100 training process.

[0130] Training data is obtained from the training dataset and used as input to the first layer of the neural network model 100. This input passes through multiple neurons in layers 1 through n, and a prediction result is output from layer n. Specifically, each neuron in each layer has its corresponding weights. Training data is input to the neurons in layer 1, and the neurons in layer 1 output their values ​​based on their corresponding weights. The output values ​​of the neurons in layer 1 are then used as input to the neurons in layer 2, and the neurons in layer 2 output their values ​​based on their corresponding weights. This process continues until a prediction result is finally output from layer n.

[0131] During the training of neural network model 100, it is desirable for the prediction result output of the nth layer of neural network model 100 to be as close as possible to the prior knowledge of the training data. Prior knowledge, also known as ground truth, generally includes the true results corresponding to the training data provided by humans. Therefore, the weights corresponding to each neuron in neural network model 100 can be updated based on the difference between the current prediction result and the prior knowledge (before the first update, the weights of each neuron are initial values). Specifically, a loss function is calculated based on the prediction result and prior knowledge; based on the loss function, optimization parameters (e.g., first-order momentum, second-order momentum) corresponding to the gradient are calculated along the direction from the nth layer to the 1st layer; and the gradient of the weights corresponding to each neuron is calculated based on the optimization parameters. After obtaining the gradient of the weights corresponding to each neuron, the weights corresponding to each neuron are updated according to each gradient, completing one round of model training. It is important to note that the gradient calculation of the weights corresponding to each neuron is performed layer by layer in the order from the nth layer to the 1st layer. After calculating the gradient of the weights corresponding to the i-th layer neuron, the gradient of the weights corresponding to the (i-1)-th layer neuron is calculated. The gradient can be obtained by taking the derivative of the loss function with respect to the weights.

[0132] During multiple rounds of model training, the weights corresponding to neurons in each layer are continuously updated, thereby enabling the prediction results output by the neural network model 100 to be as close as possible to the prior knowledge of the training data.

[0133] Figure 3a illustrates a schematic diagram of training a neural network model using a centralized federated learning method. One network element acts as the server, and other network elements act as clients. For example, when applied to the communication system shown in Figure 1a, the network element management system (EMS) is the server, and the network element (NE) is the client; or, the network management system (NMS) is the server, and the network element management system (EMS) is the client.

[0134] The training process is as follows: After training the neural network model for one or more rounds, the Client sends the parameters of the trained neural network model (e.g., the weights or gradients of neurons) to the Server. The Server aggregates the parameters from multiple Clients (e.g., by averaging or weighted averaging) and sends the aggregated parameters to each Client. The Client updates its local parameters using the aggregated parameters and continues to train the neural network model after the parameter update. The above process is repeated until training is complete.

[0135] For models with a large number of neurons, the bandwidth consumed when clients and servers interact with parameters is high. This is especially true for the server, which consumes even more bandwidth when transmitting (receiving / sending) parameters to multiple clients. For example, if the server has 100Mbps of available bandwidth and receiving parameters from a single client requires 10Mbps, then a maximum of 10 clients can send parameters to the server simultaneously. Similarly, the server can send aggregated parameters to a maximum of 10 clients at a time. If more than 10 clients participate in model training, an 11th client will interact with the server in the next second, slowing down the iteration speed and increasing the total training latency.

[0136] Figure 3b illustrates a schematic diagram of training a neural network model using a distributed federated learning method. Among the multiple nodes participating in model training, some nodes are interconnected. After training the neural network model for one or more rounds, any node sends the parameters of the neural network model (e.g., neuron weights or gradients) to other nodes connected to it. This node also receives the parameters of the neural network model from other nodes. The node aggregates the multiple parameters, updates its local parameters with the aggregated parameters, and continues to train the neural network model after the parameter update. This process is repeated until training is complete.

[0137] Suppose there are four nodes participating in model training, each connected to the other three nodes. After one training round, node 1 sends the parameters of its trained neural network model to nodes 2, 3, and 4; similarly, after one training round, node 4 sends the parameters of its trained neural network model to nodes 2, 3, and 1; nodes 2 and 3 are similar and will not be described in detail. Node 1 aggregates the parameters of its own trained neural network model with the parameters from the neural network models of nodes 2, 3, and 4 (e.g., by averaging or weighted averaging), updates its local parameters based on the aggregated parameters, and continues training the neural network model based on the updated parameters. Nodes 2, 3, and 4 follow a similar process.

[0138] In the communication system shown in Figure 1a, the nodes participating in model training are the network elements in Figure 1a. The network elements have local training and model aggregation. NMS is a federated learning (FL) MnS consumer, and EMS is a FL MnS producer.

[0139] In distributed federated learning, the connections established between network elements for transmitting neural network model parameters can be determined by either directly using existing direct communication connections (i.e., two network elements communicate directly without indirect communication through other network elements) or by arbitrarily selecting some connections from existing direct communication connections. For example, if four nodes are four base stations, and any base station can directly communicate with the other three base stations, then any one of these four base stations can establish a connection with the other three base stations for transmitting neural network model parameters. Alternatively, for any base station, one or two of the other three base stations can be arbitrarily selected to establish connections for transmitting neural network model parameters.

[0140] This distributed federated learning approach also suffers from the problem that some nodes have limited available bandwidth, making it impossible to transmit neural network model parameters simultaneously (or within a short period of time) with multiple nodes, thus affecting the model's training efficiency.

[0141] Based on this, embodiments of this application provide a communication method in which a first node receives information of a first duration; the first node triggers the establishment of a connection between computing nodes based on the first duration; wherein, the first duration is the time requirement for transmitting parameters of a first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model. By determining the connection relationship between computing nodes through the first duration, the time for transmitting parameters of the first model between computing nodes can be controlled within a reasonable range, or the time requirement for a single round of model training can be controlled within a reasonable range, or the total time requirement for model training can be controlled within a reasonable range, which can reduce the overall training time to a certain extent and improve the model training efficiency.

[0142] The relevant terms used in the embodiments of this application will be explained below. It should be noted that these explanations are for the purpose of making the embodiments of this application easier to understand, and should not be regarded as a limitation on the scope of protection claimed by this application.

[0143] (1) First node, second node, third node, compute node, first compute node, second compute node, third compute node:

[0144] The computing nodes are nodes that train the first model, or nodes that participate in the training task of the first model. In the embodiments of this application, there are at least three computing nodes, referred to as the first computing node, the second computing node, the third computing node, and optionally, a fourth computing node, etc. These at least three computing nodes train the first model using a distributed federated learning method, or use a method similar to the distributed federated learning method to train the first model. Here, "similar" mainly refers to the transmission of the parameters of the trained first model between the connected nodes.

[0145] In scenario a, the first node is a compute node, the second node manages compute nodes (including the first node), and the third node manages the second node. For example, when applied to the communication architecture shown in Figure 1a, the compute node is the NE, the second node is the EMS, and the third node is the NMS. The third node can also be called a Model Training Consumer (MLT consumer), and the second node or compute node can also be called a Model Training Producer (MLT producer). Optionally, this scenario can be applied to situations where the EMS and NE are deployed in different devices.

[0146] In another scenario b, the first node can manage the compute nodes, and the second node can manage the first node. For example, when applied to the communication architecture shown in Figure 1a, the second node is the NMS, the first node is the EMS, and the compute nodes are the NEs. The second node can be called the Model Training Consumer (MLT consumer), and the first node can be called the Model Training Producer (MLT producer). Optionally, this scenario can be applied to situations where the EMS and NE are deployed on one device or on different devices.

[0147] (2) The first model can be: neural network model, artificial neural network model, AI model, ML model, etc.

[0148] (3) Available bandwidth between computing nodes: refers to the available bandwidth between any two computing nodes.

[0149] In the embodiments of this application, there is available bandwidth between two computing nodes only if a direct communication has been established between them (they can communicate directly without the need for forwarding through other nodes / network elements); in other words, the mention of "available bandwidth between computing nodes" in the embodiments of this application indicates that a direct communication has been established between the two computing nodes.

[0150] Available bandwidth can be, for example, the current remaining bandwidth, the average remaining bandwidth over a past period, or the average remaining bandwidth over a predicted next period. "Available bandwidth" can be replaced with "remaining bandwidth" or "idle bandwidth," etc.

[0151] Optionally, in this embodiment, the available bandwidth between computing nodes does not change significantly during the training of the first model, or does not change significantly within the first period. Here, "not changing significantly" can be understood as the variance of the available bandwidth not exceeding a set threshold within a given time period.

[0152] (4) The “duration requirement” in the embodiments of this application can also be replaced by “duration limit”, which is the maximum allowed duration value, or training efficiency.

[0153] In the embodiments of this application, unless otherwise specified, the connections between computing nodes are used for transmitting parameters of the first model between the computing nodes. For example, the connection between the first computing node and the second node is used for transmitting parameters of the first model between the first computing node and the second computing node. The connection used for transmitting parameters of the first model between computing nodes may also be called a "federated connection" or a "federated interactive connection"; the connection relationship used for transmitting parameters of the first model between computing nodes may be called a "federated connection relationship", a "federated interactive connection relationship", or a "federated interconnection relationship", etc.

[0154] For a given computing node, the computing nodes that establish a connection with that computing node (a connection used to transmit parameters of the first model) can be called the neighboring nodes of that computing node. For example, if a connection is established between the first computing node and the second computing node, then the neighboring nodes of the first computing node include the second computing node, and the neighboring nodes of the second computing node include the first computing node.

[0155] To better illustrate the embodiments of this application, the methods provided by the embodiments of this application are described below with reference to the accompanying drawings. Unless otherwise specified below, the steps indicated by dashed lines in the accompanying drawings corresponding to the various embodiments of this application are optional steps. It should be noted that the technical details of the multiple embodiments provided in this application can be referenced to each other, each embodiment described below can exist independently, and multiple embodiments can also be combined with each other as an embodiment in the absence of logical errors.

[0156] Figure 4 is a flowchart illustrating a communication method provided in an embodiment of this application. Figure 4 uses a first node and a second node as examples of the execution entities in this interactive illustration to illustrate the method. However, this embodiment does not limit the execution entities in this interactive illustration. For example, the method executed by the first node in Figure 4 can also be executed by a module applied to the first node (e.g., a chip, chip system, or processor), or by a logical node, logical module, or software that implements all or part of the functions of the first node. Similarly, the method executed by the second node in Figure 4 can also be executed by a module applied to the second node (e.g., a chip, chip system, or processor), or by a logical node, logical module, or software that implements all or part of the functions of the second node.

[0157] As shown in Figure 4, the method may include the following steps:

[0158] Step 401: The second node sends information for the first duration to the first node, and the first node receives the information for the first duration accordingly.

[0159] The first duration includes, but is not limited to, the following possible scenarios: the first duration is the time required for transmitting parameters of the first model between computing nodes; the first duration is the time required for a single round of training the first model; the first duration is the total time required for training the first model.

[0160] The second node can determine the first duration itself, or it can obtain the first duration information from other nodes. In one scenario, the first node is a computing node, and the second node manages computing nodes (including the first node); in this scenario, the second node determines the first duration. Alternatively, the third node sends the first duration information to the second node, and the second node receives the first duration information from the third node; optionally, the third node determines the first duration. In another scenario, the first node manages computing nodes, and the second node manages the first node; in this scenario, the second node determines the first duration.

[0161] Step 402: The first node triggers the establishment of a connection between the first computing node and the second computing node based on the first duration.

[0162] The connection between the first computing node and the second computing node is used to transmit the parameters of the first model between the first computing node and the second computing node.

[0163] The total training time of the first model is determined based on the number of training epochs and the training time per epoch of the first model. With a fixed number of training epochs, the shorter the training time per epoch of the first model, the shorter the total training time of the first model, and thus the higher the model training efficiency. The training time per epoch of the first model is determined based on the time required for local training on the computing nodes and the time required for transmitting the parameters of the first model between computing nodes. Given a known model structure and a known time required for local training on the computing nodes, the shorter the time required for transmitting the parameters of the first model between computing nodes, the shorter the total training time of the first model.

[0164] By determining the connection relationship between computing nodes based on the time required for transmitting the parameters of the first model between computing nodes, the transmission time of the first model's parameters between computing nodes can be controlled within a reasonable range, and the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0165] If the connection relationship between computing nodes is determined by the training time requirement of the first model in a single round, the training time of the first model in one round can be controlled within a reasonable range. With a fixed number of training rounds, the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0166] If the connection relationship between computing nodes is determined by the total training time requirement of the first model, the total training time of the first model can be controlled within a reasonable range, thereby improving the training efficiency of the model.

[0167] The following is an example of how the first node triggers the establishment of a connection between the first computing node and the second computing node in step 402:

[0168] In one possible implementation, a first computing node creates and maintains a first table, which stores identifiers of computing nodes capable of establishing a connection with the first computing node for transmitting parameters of the first model. A second computing node creates and maintains a second table, which stores identifiers of computing nodes capable of establishing a connection with the second computing node for transmitting parameters of the first model. When the first computing node stores the identifier of the second computing node in the first table, and the second computing node stores its identifier in the second table, a connection for transmitting parameters of the first model is established between the first and second computing nodes.

[0169] In one scenario, the first node is the first computing node.

[0170] The first computing node (i.e., the first node) sends an instruction message to the second computing node requesting the establishment of a connection between the first computing node and the second computing node. This instruction message includes the identifier of the first computing node. After receiving the instruction message, if the second computing node agrees to establish the connection between the first computing node and the second computing node, it can send an instruction message of agreement to the first computing node. Optionally, this instruction message includes the identifier of the second computing node.

[0171] If the first computing node determines that the second computing node agrees to establish a connection, it can store the identifier of the second computing node in the first table. If the second computing node, upon receiving an instruction from the first computing node, agrees to establish a connection between the first and second computing nodes, it can store the identifier of the first computing node in the second table.

[0172] If the first node is the second computing node, the principle is the same as if the first node is the first computing node, so it will not be repeated.

[0173] In one scenario, the first node manages the computing nodes.

[0174] For example, a first node sends a first instruction to a first computing node, requesting the establishment of a connection between the first and second computing nodes. The first instruction includes the identifiers of both the first and second computing nodes. The first node then sends a second instruction to the second computing node, also requesting the establishment of a connection between the two nodes. The second instruction includes the identifiers of both the first and second computing nodes. Upon receiving the first instruction, the first computing node can store the identifier of the second computing node in a first table. Similarly, upon receiving the second instruction, the second computing node can store its identifier in a second table.

[0175] For example, a first node sends a first instruction to a first computing node. This first instruction requests the establishment of a connection between the first and second computing nodes. The first instruction includes the identifiers of both the first and second computing nodes. Upon receiving the first instruction, the first computing node sends (e.g., based on the identifier of the second computing node in the first instruction) an instruction to the second computing node requesting the establishment of the connection, without needing to send a second instruction. If the second computing node agrees to establish the connection, it can send an instruction indicating its agreement to the connection to the first computing node. After determining that the second computing node agrees to the connection, the first computing node can store the identifier of the second computing node in a first table. Similarly, after receiving the instruction from the first computing node, if the second computing node agrees to establish the connection, it can store the identifier of the first computing node in a second table.

[0176] It is understandable that step 402 is described using the first and second computing nodes as examples. Computing nodes can include not only the first and second computing nodes, but also a third computing node, a fourth computing node, and so on. That is to say, in addition to triggering the establishment of a connection between the first and second computing nodes, it may also trigger the establishment of a connection between the first computing node and other computing nodes (such as the third and / or fourth computing nodes), and it may also trigger the establishment of a connection between the third and fourth computing nodes.

[0177] Optionally, before step 402, the first node determines the connection relationship between computing nodes based on the first duration. For ease of description later, this connection relationship is referred to as the first connection relationship. The first connection relationship is a connection relationship used for transmitting parameters of the first model between computing nodes. The first connection relationship includes the connection relationship between the first computing node and the second computing node. Then, based on the first connection relationship, the first node triggers the establishment of a connection between the corresponding computing nodes, including triggering the establishment of a connection between the first computing node and the second computing node (i.e., step 402).

[0178] If the first node manages the computation nodes, the first connection relationship can be understood as the connection relationship between all computation nodes training the first model. If the first node is a computation node, the first connection relationship can be understood as the connection relationship between that computation node and other computation nodes. In other words, each computation node can determine its own connection relationship with other computation nodes.

[0179] One possible implementation of step 402 is as follows: the first node triggers the establishment of a connection between the first computing node and the second computing node based on the first duration, the available bandwidth between the first computing node and the second computing node, and the data volume of the parameters of the first model. Optionally, prior to this, the first node determines the first connection relationship between the computing nodes based on the first duration, the available bandwidth between the computing nodes, and the data volume of the parameters of the first model.

[0180] The following details the process by which the first node determines the first connection relationship between computing nodes based on the first duration, the available bandwidth between computing nodes, and the data volume of the parameters of the first model:

[0181] Example 1: The first node determines the first connection relationship between computing nodes based on the time required for transmitting the parameters of the first model between computing nodes, the available bandwidth between computing nodes, and the data volume of the parameters of the first model.

[0182] Optionally, given the available bandwidth between the first computing node and the second computing node and the amount of data for the parameters of the first model, the first estimated duration for bidirectional transmission of the parameters of the first model between the first computing node and the second computing node satisfies the duration requirement (i.e., the first duration). In other words, given the available bandwidth between two connected computing nodes and the amount of data for the parameters of the first model, the first estimated duration for bidirectional transmission of the parameters of the first model between these two computing nodes satisfies the duration requirement.

[0183] One possible implementation involves the first node determining the amount of data transferred bidirectionally between computing nodes during a single round of model training (here, "between computing nodes" refers to any two computing nodes, and the amount of data transferred bidirectionally between any two computing nodes during a single round of model training is the same or substantially the same). Based on the time requirement for transferring the parameters of the first model between computing nodes, the amount of data transferred bidirectionally between computing nodes during a single round of model training, and the available bandwidth between two computing nodes (e.g., the first computing node and the second computing node), it determines whether the two computing nodes can complete the bidirectional transfer of the parameters of the first model within a first time period. If yes, then a connection is allowed to be established between the two computing nodes to transfer the parameters of the first model; if no, a connection is not allowed to be established between the two computing nodes to transfer the parameters of the first model.

[0184] For example: the first duration is 1 second, and the amount of data transmitted bidirectionally between computing nodes in a single round of model training is 16Mb; if the available bandwidth between the first computing node and the second computing node is greater than or equal to 16Mpbs, then the connection between the first computing node and the second computing node is allowed; otherwise, the connection between the first computing node and the second computing node is not allowed.

[0185] To illustrate further: Consider four compute nodes participating in model training, with any two nodes having established a direct communication connection. The available bandwidth between compute node 1 and compute node 2 is 20 Mbps, between compute node 1 and compute node 3 is 17 Mbps, between compute node 1 and compute node 4 is 10 Mbps, between compute node 2 and compute node 3 is 16 Mbps, between compute node 2 and compute node 4 is 18 Mbps, and between compute node 3 and compute node 4 is 20 Mbps. The initial training duration is 1 second, and the bidirectional data transmission between compute nodes in a single round of model training is 16 Mb. When the available bandwidth between two compute nodes is greater than or equal to 16 Mbps, a connection is allowed between these two compute nodes. Therefore, connections for transmitting the parameters of the first model can be established between compute nodes 1 and 2, compute nodes 1 and 3, compute nodes 2 and 4, compute nodes 2 and 3, and compute nodes 3 and 4. However, a connection for transmitting the parameters of the first model cannot be established between compute node 1 and 4.

[0186] Another possible implementation involves the first node determining the amount of data that can be unidirectionally transmitted between computing nodes during a single round of model training, and half the required time for transmitting the parameters of the first model between computing nodes. Based on this half-time, the amount of data that can be unidirectionally transmitted between computing nodes during a single round of model training, and the available bandwidth between the two computing nodes (e.g., the first and second computing nodes), it is determined whether the two computing nodes can unidirectionally transmit the parameters of the first model within half the time. If yes, a connection is allowed between the two computing nodes to transmit the parameters of the first model; otherwise, a connection is not allowed between the two computing nodes to transmit the parameters of the first model.

[0187] Another possible implementation involves the first node determining the amount of data transmitted unidirectionally between computing nodes during a single round of model training, and half of the available bandwidth between the two computing nodes. Based on the time requirement for transmitting the parameters of the first model between computing nodes, the amount of data transmitted unidirectionally between computing nodes during a single round of model training, and half of the available bandwidth between the two computing nodes, it determines whether the two computing nodes can unidirectionally transmit the parameters of the first model within a first time period. If yes, then a connection is allowed to be established between the two computing nodes to transmit the parameters of the first model; if not, a connection is not allowed to be established between the two computing nodes to transmit the parameters of the first model.

[0188] In a single round of model training, the amount of data transmitted unidirectionally between computing nodes is determined based on the amount of data in the parameters of the first model and the model training method (or model optimization algorithm) used to train the first model. For example, if the amount of data in the parameters of the first model is 2Mb and the model training method is adaptive moment estimation (ADAM), then the amount of data transmitted unidirectionally between computing nodes in a single round of model training is 2Mb*4=8Mb, and the amount of data transmitted bidirectionally between computing nodes in a single round of model training is 16Mb.

[0189] The following describes the process by which the first node obtains / determines information such as the available bandwidth between computing nodes, the amount of data transferred unidirectionally (or bidirectionally) between computing nodes during a single round of model training, and the model training method:

[0190] In one scenario, the first node is the computing node, and the second node manages the computing nodes.

[0191] The first node determines the amount of data for the parameters of the first model based on the number of neurons in the first model. The second node sends a model training method to the first node, and correspondingly, the first node receives the model training method from the second node; then, the first node trains the first model based on the model training method; and based on the model training method and the amount of data for the parameters of the first model, determines the amount of data transmitted unidirectionally and bidirectionally between computing nodes in a single round of model training. Optionally, the third node sends the model training method to the second node, for example, by instructing the second node on the model training method via a MIL training request (MLTraining request); or, the second node determines the model training method itself. The first node can also determine the available bandwidth between itself and other computing nodes, or the third node sends the available bandwidth between itself and other computing nodes to the second node, and the second node sends the available bandwidth between itself and other computing nodes to the first node.

[0192] In one scenario, the first node manages the computing nodes.

[0193] The second node or computing node sends the amount of data for the parameters of the first model, or the number of neural network elements in the first model, to the first node; the first node, knowing the number of neurons in the first model, can then determine the amount of data for the parameters. The first node determines the model training method; or, the second node sends the model training method to the first node, and correspondingly, the first node receives the model training method from the second node. For example, the second node instructs the first node on the model training method via MLTraningrequest. Based on the amount of data for the parameters of the first model and the model training method, the first node determines the amount of data transmitted unidirectionally and bidirectionally between computing nodes in a single round of model training. The first node can also send the model training method to the computing node so that the computing node can train the first model according to the model training method. The computing node sends the available bandwidth between itself and other computing nodes to the first node, or the second node sends the available bandwidth between computing nodes to the first node.

[0194] Example 2: The first node determines the first connection relationship between computing nodes based on the single-round training time requirement of the first model, the available bandwidth between computing nodes, and the amount of data of the parameters of the first model.

[0195] Optionally, given the available bandwidth between the first and second computing nodes and the amount of data for the parameters of the first model, the second estimated duration for the first computing node and the second computing node to jointly train one round of the first model satisfies the single-round duration requirement (i.e., the first duration). In other words, given the available bandwidth between two connected computing nodes and the amount of data for the parameters of the first model, the second estimated duration for these two computing nodes to jointly train one round of the first model satisfies the single-round duration requirement.

[0196] The process of two computing nodes jointly training one round of the first model includes: a single computing node executing the training process of one round of the model as shown in Figure 2 (i.e., the local training process on the computing node) and the process of transferring the parameters of the first model obtained through training between the two computing nodes. The time required for a single computing node to execute the training process of one round of the model as shown in Figure 2 can be estimated in advance. The time required for the two computing nodes to transfer the parameters of the first model obtained through training is related to the available bandwidth between the two computing nodes, the amount of data of the first model parameters, and the model training method.

[0197] In one possible implementation, the first node determines the estimated duration b1 for transmitting the parameters of the first model between the two computing nodes based on the available bandwidth between the two computing nodes, the amount of data for the parameters of the first model, and the model training method. Next, based on the duration b2 required for a single computing node to perform one round of model training as illustrated in Figure 2, and the estimated duration b1, a second estimated duration B is determined for the two computing nodes (e.g., the first computing node and the second computing node) to jointly train one round of the first model; for example, B = b2 + b1. If the second estimated duration B is less than or equal to the single-round duration requirement, a connection is allowed between the two computing nodes to transmit the parameters of the first model; if the second estimated duration B is greater than the single-round duration requirement, a connection is not allowed between the two computing nodes to transmit the parameters of the first model.

[0198] For example: If the data size of the parameters of the first model is 2Mb, and the model training method is ADAM, then the amount of data transmitted unidirectionally between computing nodes in a single round of model training is 2Mb * 4 = 8Mb. If the available bandwidth between the two computing nodes is 18Mbps, then the expected duration b1 is 8 / 9, for example, the expected duration b1 = 8 / (18 / 2), or b1 = (8 * 2) / 18. The first duration (i.e., the duration requirement for a single round) is 1.5s, the duration b2 is 0.6s, and the duration b1 is 8 / 9, then B is approximately equal to 1.4889s, which allows the connection to be established between the two computing nodes.

[0199] In one possible implementation, the first node determines the time requirement 'a' for transmitting the parameters of the first model between computing nodes based on the first duration and the duration 'b2' required for a single computing node to perform one round of model training as illustrated in Figure 2. For example, the time requirement 'a' is obtained by subtracting 'b2' from the first duration. Next, the first node determines the first connection relationship between computing nodes based on the time requirement 'a', the available bandwidth between computing nodes, and the amount of data for the parameters of the first model, referring to Example 1, which will not be repeated here.

[0200] Example 3: The first node determines the first connection relationship between computing nodes based on the total time required to train the first model, the available bandwidth between computing nodes, and the amount of data for the parameters of the first model.

[0201] Optionally, given the available bandwidth between the first and second computing nodes and the amount of data for the parameters of the first model, the third estimated training time for the first and second computing nodes to jointly train the first model meets the total training time requirement (i.e., the first training time). In other words, given the available bandwidth between two connected computing nodes and the amount of data for the parameters of the first model, the third estimated training time for these two computing nodes to jointly train the first model meets the total training time requirement.

[0202] In one possible implementation, the first node determines the estimated duration b1 for transmitting the parameters of the first model between the two computing nodes based on the available bandwidth between the two computing nodes and the amount of data for the parameters of the first model. Based on the duration b2 required for a single computing node to perform one round of model training as illustrated in Figure 2 and the estimated duration b1, a second estimated duration B is determined for the two computing nodes (e.g., the first computing node and the second computing node) to jointly train one round of the first model. For example, B = b2 + b1. Next, the first node determines a third estimated duration C for the two computing nodes to jointly train the first model based on the number of training rounds K and the second estimated duration B, for example, C = B * K. If the third estimated duration C is less than or equal to the total time requirement, a connection is allowed between the two computing nodes to transmit the parameters of the first model; if the third estimated duration C is greater than the total time requirement, a connection is not allowed between the two computing nodes to transmit the parameters of the first model.

[0203] In one possible implementation, the first node determines the single-round training time requirement of the first model based on the number of training rounds K and the total training time requirement of the first model. Next, the first node determines the first connection relationship between computing nodes based on the single-round training time requirement of the first model, the available bandwidth between computing nodes, and the amount of data for the parameters of the first model, as shown in Example 2, and will not be repeated here.

[0204] Examples 1, 2, and 3 above illustrate how, based on a first duration and the bandwidth between computing nodes, it is determined whether a connection for transmitting parameters of the first model can be established between two computing nodes. If allowed, the first node can trigger the establishment of the corresponding connection between the computing nodes. In one possible implementation, the number of computing nodes establishing connections is limited. For example, for any given computing node, the number of other computing nodes that establish connections with it does not exceed a first number, which is the upper limit of the number of computing nodes that a node can (is allowed) to connect to. The first number can also be understood as the upper limit of the number of neighbors. In other words, a computing node can establish connections for transmitting parameters of the first model with at most the first number of computing nodes. By limiting the number of computing nodes establishing connections, the available bandwidth of the computing nodes can be saved.

[0205] One possible implementation of the first node triggering the establishment of a connection between the first computing node and the second computing node based on a first duration (i.e., step 402) is as follows: the first node triggers the establishment of a connection between the first computing node and the second computing node based on a first duration and a first quantity. More specifically, the first node triggers the establishment of a connection between the first computing node and the second computing node based on a first quantity, a first duration, the available bandwidth between the first computing node and the second computing node, and the data volume of the parameters of the first model.

[0206] Optionally, prior to this, the first node determines a first connection relationship between computing nodes based on a first duration and a first quantity. This first connection relationship is used for transmitting parameters of the first model between computing nodes, and includes a connection relationship between the first computing node and the second computing node. Then, based on the first connection relationship, the first node triggers the establishment of connections between computing nodes, including triggering the establishment of a connection between the first computing node and the second computing node. More specifically, the first node determines the first connection relationship between computing nodes based on the first quantity, the first duration, the available bandwidth between the first and second computing nodes, and the data volume of the parameters of the first model.

[0207] The following describes the process by which the first node determines the connection relationships between computing nodes based on the first duration and the first quantity:

[0208] Examples 1, 2, and 3 above illustrate how a first node determines the first connection relationship between computing nodes based on a first duration, the available bandwidth between computing nodes, and the data volume of the parameters of the first model. Alternatively, Examples 1, 2, and 3 describe how to determine whether a connection for transmitting the parameters of the first model is allowed between computing nodes. Building on these examples, the first node can, for each computing node, select no more than a first number of other computing nodes from those allowed to establish connections with it to establish a connection.

[0209] When selecting other computing nodes, the first node can choose any other computing node, or refer to one or more of the following parameters, including but not limited to: available bandwidth between computing nodes, data distribution similarity between computing nodes (data distribution similarity can be represented by KL divergence value), location information of computing nodes, and load information of computing nodes. For example, computing nodes can be selected according to one or more of the following principles: 1) Prioritize computing nodes with large available bandwidth; 2) Prioritize computing nodes with high data distribution similarity; 3) Prioritize computing nodes with closer geographical locations; 4) Prioritize computing nodes with similar loads.

[0210] For example: Continuing with Example 1 above, connections can be established between compute nodes 1 and 2, compute nodes 1 and 3, compute nodes 2 and 4, compute nodes 2 and 3, and compute nodes 3 and 4. The initial quantity is 2, prioritizing compute nodes with larger available bandwidth. Compute node 2 connects to 3 compute nodes (compute nodes 1, 3, and 4), and compute node 3 also connects to 3 compute nodes (compute nodes 1, 2, and 4). Therefore, among the 3 compute nodes connected to compute node 3, priority should be given to selecting the 2 compute nodes with larger available bandwidth. Thus, the available bandwidth between compute nodes 2 and 3 is the smallest, and a connection cannot be established between compute nodes 2 and 3.

[0211] The following is an example of the first node obtaining the first quantity of information:

[0212] In one scenario, the first node is a computing node. The second node sends a first quantity of information to the first node, and correspondingly, the first node receives the first quantity information from the second node. Optionally, the second node determines the first quantity; or a third node sends the first quantity of information to the second node, and correspondingly, the second node receives the first quantity information from the third node.

[0213] In one scenario, the first node manages the computing nodes. The first node determines a first quantity; or, the second node sends information about the first data to the first node, and correspondingly, the first node receives information about the first quantity from the second node.

[0214] Based on the preceding introduction, taking the first node as the first computing node as an example, Figure 5 shows a schematic diagram of a communication method, including the following steps:

[0215] Step 501: The third node sends information about the first duration and the first quantity to the second node, and the second node receives the information about the first duration and the first quantity.

[0216] In step 501, the information about the first duration and the information about the first quantity can be in the same message or in different messages. This message can be a distributed federated learning request message.

[0217] The third node can be an NMS or a model training user; the second node can be an EMS or a model training provider.

[0218] Step 502: The second node sends information about the first duration and the first quantity to the computing nodes participating in model training (including the first computing node (i.e., the first node) and the second computing node).

[0219] In step 502, the information on the first duration and the information on the first quantity can be in the same message or in different messages.

[0220] Step 503a: Based on the first duration, the first quantity, the available bandwidth between the first computing node and other computing nodes, and the data volume of the parameters of the first model, the first computing node determines the first connection relationship associated with the first computing node, that is, determines the computing node M (including the second computing node) that establishes a connection with the first computing node for transmitting the parameters of the first model.

[0221] For specific details, please refer to the previous introduction; they will not be repeated here.

[0222] Step 504a: The first computing node establishes a connection between the first computing node and computing node M for transmitting parameters of the first model.

[0223] How to establish connections between computing nodes has been explained in the previous text and will not be repeated here.

[0224] Step 503b: Based on the first duration, the first quantity, the available bandwidth between the second computing node and other computing nodes, and the data volume of the parameters of the first model, the second computing node determines the first connection relationship associated with the second computing node, that is, determines the computing node N (including the first computing node) that establishes a connection with the second computing node for transmitting the parameters of the first model.

[0225] For specific details, please refer to the previous introduction; they will not be repeated here.

[0226] Step 504b: The second computing node establishes a connection between the second computing node and computing node N for transmitting parameters of the first model.

[0227] How to establish connections between computing nodes has been explained in the previous text and will not be repeated here.

[0228] Based on the preceding introduction, taking the management of computing nodes by the first node as an example, Figure 6 shows a schematic diagram of a communication method, including the following steps:

[0229] Step 601: The second node sends information about the first duration and the first quantity to the first node, and the first node receives the information about the first duration and the first quantity.

[0230] In step 601, the information about the first duration and the information about the first quantity can be in the same message or in different messages. This message can be a distributed federated learning request message.

[0231] The second node can be an NMS or a model training user; the first node can be an EMS or a model training provider.

[0232] Step 602: The compute node sends the available bandwidth between itself and other compute nodes to the first node.

[0233] Optionally, the first node may first send a query request to the compute node, which is used to query the available bandwidth between the compute node and other compute nodes. After receiving the query request, the compute node sends the available bandwidth between the compute node and other compute nodes to the first node.

[0234] Alternatively, the first node can send a subscription request to the compute node, which indicates that it is subscribing to the available bandwidth between the compute node and other compute nodes. When the compute node determines that the subscription conditions are met (e.g., periodic reporting or reporting of specific events), it sends the available bandwidth between the compute node and other compute nodes to the first node.

[0235] Step 603: The first node determines the first connection relationship between computing nodes based on the first duration, the first quantity, the available bandwidth between computing nodes, and the data volume of the parameters of the first model.

[0236] For specific details, please refer to the previous introduction; they will not be repeated here.

[0237] Step 604: The first node sends the connection relationship associated with the computing node to the computing node based on the first connection relationship.

[0238] Step 605: Based on the received connection relationship, the computing node establishes a connection with the corresponding computing node for transmitting the parameters of the first model.

[0239] For specific details, please refer to the previous introduction; they will not be repeated here.

[0240] Based on the preceding description, the first node can manage computing nodes, or the first node itself can be a computing node. That is, the node managing computing nodes can determine (or select) the connection relationships between computing nodes used to transmit parameters of the first model, and the node knowing the computing nodes can determine the connection relationships between those computing nodes and other computing nodes used to transmit parameters of the first model. Which node determines the connection relationships between computing nodes can be specified by the protocol or flexibly indicated through a strategy.

[0241] In one scenario, the first node is a compute node. For example, the third node is an NMS, the second node is an EMS, and the compute node (first node) is a network element (NE).

[0242] For example, the second node sends a first strategy to the first node, and correspondingly, the first node receives the first strategy from the second node. The first strategy is used to instruct the computing nodes to determine the connection relationships for transmitting the parameters of the first model. Specifically, the first strategy is used to instruct the computing nodes to determine the connection relationships between the computing nodes and other computing nodes for transmitting the parameters of the first model. Optionally, the second node determines the first strategy; or, the third node sends the first strategy to the second node, and correspondingly, the second node receives the first strategy from the third node. Optionally, in step 501, the third node sends the first strategy to the second node, and in step 502, the second node sends the first strategy to each computing node. The first strategy and the information on the first quantity or the first duration can be located in the same message or in different messages.

[0243] For example, the second node does not need to send the first strategy to the first node. Instead, the second node sends information about the first duration (optionally also sending information about the first quantity) to the first node. The information about the first duration can implicitly indicate that the computing nodes determine the connection relationships between the computing nodes and other computing nodes used to transmit the parameters of the first model. If the connection relationships between computing nodes are determined by the node managing the computing nodes, the node managing the computing nodes (i.e., the second node) does not need to send the information about the first duration to the computing nodes; it only needs to send the determined connection relationships to the computing nodes.

[0244] In one scenario, the first node manages the compute nodes. For example, the second node is the NMS, the first node is the EMS, and the compute nodes are network elements (NEs).

[0245] For example, the second node sends a second strategy to the first node, and correspondingly, the first node receives the second strategy from the second node. The second strategy is used to instruct the node managing the computing nodes to determine the connection relationship between the computing nodes for transmitting the parameters of the first model. Optionally, in step 601, the second node sends the second strategy to the first node. The second strategy and the information about the first quantity or the first duration can be in the same message or in different messages.

[0246] For example, the second node does not send the second strategy to the first node, and the first node determines the connection relationship between the computing nodes used to transmit the parameters of the first model.

[0247] The first and second strategies mentioned above can also be called neighbor selection strategies.

[0248] The available bandwidth between computing nodes is dynamically changing. In one possible implementation, after step 402, the first node can again determine the connection relationship between computing nodes based on the first duration. To distinguish it from the already determined first connection relationship, the connection relationship determined after the first connection relationship is called the second connection relationship. The second connection relationship is the connection relationship used for transmitting parameters of the first model between the computing nodes. Based on the second connection relationship, the first node triggers an update of the connection between computing nodes.

[0249] If the first node manages the computation nodes, the second connection relationship can be understood as the connection relationship between all computation nodes used to train the first model. If the first node is a computation node, the second connection relationship can be understood as the connection relationship between that computation node and other computation nodes. In other words, each computation node can again determine its own connection relationship with other computation nodes.

[0250] One way to trigger an update of the connections between computing nodes is to: trigger the disconnection (or cancellation, deletion) of existing connections between computing nodes used for transmitting parameters of the first model; and, based on the second connection relationship, trigger the establishment of connections between the computing nodes. This method can be understood as a complete refresh of the connection relationships between computing nodes based on the second connection relationship.

[0251] Another way to trigger the update of the connection between the computing nodes is to: trigger the disconnection of a connection that exists in the first connection relationship but does not exist in the second connection relationship; and trigger the establishment of a connection that does not exist in the first connection relationship but exists in the second connection relationship.

[0252] For example: The first connection relationship includes: compute node 1 is connected to compute node 2, compute node 1 is connected to compute node 4, and compute node 3 is connected to compute node 4. The second connection relationship includes: compute node 1 is connected to compute node 2, compute node 1 is connected to compute node 3, compute node 4 is connected to compute node 2, and compute node 4 is connected to compute node 3. Triggering the deletion of the connection between compute node 1 and compute node 4 triggers the establishment of a connection between compute node 1 and compute node 3, and triggers the establishment of a connection between compute node 3 and compute node 4.

[0253] The process of triggering the establishment of a connection between compute nodes can be referred to the example of triggering the establishment of a connection between the first and second compute nodes introduced earlier, and will not be described in detail here. The following describes an example of triggering the disconnection between compute nodes. We will still use the first and second compute nodes as an example:

[0254] In one scenario, the first node is the first computing node.

[0255] The first compute node deletes the identifier of the second compute node stored in the first table.

[0256] The first computing node (i.e., the first node) sends an instruction to the second computing node to disconnect the connection between the first and second computing nodes. This instruction includes the identifier of the first computing node. After receiving the instruction, the second computing node deletes the identifier of the first computing node from the second table. Optionally, the second computing node can also send a disconnection instruction to the first computing node, which optionally includes the identifier of the second computing node.

[0257] If the first node is the second computing node, the principle is the same as if the first node is the first computing node, so it will not be repeated.

[0258] In one scenario, the first node manages the computing nodes.

[0259] For example, a first node sends a third instruction to a first computing node, indicating that the connection between the first and second computing nodes be disconnected. This third instruction includes the identifiers of both the first and second computing nodes. The first node then sends a fourth instruction to the second computing node, also indicating that the connection between the first and second computing nodes be disconnected. This fourth instruction includes the identifiers of both the first and second computing nodes. Upon receiving the third instruction, the first computing node deletes the identifier of the second computing node from a first table. Upon receiving the fourth instruction, the second computing node deletes the identifier of the second computing node from a second table.

[0260] For example, a first node sends a third instruction to a first computing node, where the first instruction instructs the deletion of the connection between the first and second computing nodes. Upon receiving the third instruction, the first computing node deletes the identifier of the second computing node stored in a first table and sends (e.g., based on the identifier of the second computing node in the first instruction) an instruction to the second computing node to delete the connection between them. This eliminates the need for the first node to send a fourth instruction to the second computing node. Upon receiving the instruction from the first computing node, the second computing node deletes the identifier of the first computing node stored in a second table.

[0261] Understandably, the process of determining the connection relationships between computing nodes does not conflict with the model training process and can be executed in parallel. After determining the new connection relationships, the connection between computing nodes is updated. This process takes very little time and can be ignored, without affecting the model training time.

[0262] The available bandwidth between computing nodes is dynamic. At a certain point in time, the connections between some computing nodes may meet the corresponding duration requirements, but may not be able to meet them for a period of time. In this implementation, dynamically determining the connection relationships between computing nodes can reduce the occurrence of situations where the first duration requirement is not met, thereby improving model training efficiency.

[0263] In one possible implementation, the first node periodically determines the connection relationships between computing nodes according to a first cycle. Alternatively, the first node triggers a re-determination of the connection relationships between computing nodes based on a specific reason, such as the available bandwidth fluctuation value between the connected computing nodes being greater than a set threshold, or the available bandwidth being less than a set value.

[0264] The first cycle can be specified by the protocol or flexibly indicated. The first cycle can be in seconds, such as 10 seconds, 8 seconds, etc.

[0265] In one scenario, the first node is a compute node. For example, the third node is an NMS, the second node is an EMS, and the compute node (first node) is a network element (NE).

[0266] The second node sends the first period's information to the first node, and correspondingly, the first node receives the first period's information from the second node. Optionally, the second node determines the first period's information; or, the third node sends the first period's information to the second node, and correspondingly, the second node receives the first period's information from the third node. Optionally, in step 501, the third node sends the first period's information to the second node, and in step 502, the second node sends the first period's information to each computing node. The first period's information, along with the first strategy, or the first quantity information, or the first duration information, can be in the same message or in different messages.

[0267] For example, the third node sends a dynamic connection determination instruction to the second node, and the second node determines the information for the first period based on this instruction. Alternatively, the third node sends a static connection determination instruction to the second node, and the second node also sends a static connection determination instruction to the first node. Here, static connection determination means that the connection between the computation nodes is determined only once during the training of the first model. Alternatively, the second node may not send the first period information to the first node, meaning that the connection between the computation nodes is determined only once during the training of the first model.

[0268] In one scenario, the first node manages the compute nodes. For example, the second node is the NMS, the first node is the EMS, and the compute nodes are network elements (NEs).

[0269] For example, the second node sends information for the first period to the first node, and correspondingly, the first node receives information for the first period from the second node. Optionally, in step 601, the second node sends information for the first period to the first node. The information for the first period, the first strategy, the first quantity, or the first duration can be in the same message or in different messages.

[0270] For example, the second node does not send the first period's information to the first node; the first node determines the first period itself.

[0271] For example, the second node sends an instruction to the first node to dynamically determine the connection relationship, and the first node determines the information for the first period based on this instruction. Alternatively, the second node sends an instruction to the first node to statically determine the connection relationship, and the first node does not need to determine the information for the first period. Or, the second node does not send an instruction to the first node to dynamically determine the connection relationship, which means that the connection relationship between the computing nodes is determined only once during the training of the first model.

[0272] After determining the connection relationships between the computing nodes, in one possible implementation, the first node sends these connection relationships to the second node. Let's take the example of the first node sending the first connection relationship to the second node:

[0273] In one scenario, the first node is a compute node. For example, the third node is an NMS, the second node is an EMS, and the compute node (first node) is a network element (NE).

[0274] The following is an introduction in conjunction with Figure 5:

[0275] Step 505a: The first node (first computing node) sends the first connection relationship associated with the first computing node to the second node.

[0276] Connection relationships can be represented by an adjacency matrix or a 2 / 3-dimensional array of edges. For example, if the first computing node and the second computing node are connected, the connection relationship between the first computing node and the second computing node can be represented by [1, 2]. Here, "1" represents the identifier of the first computing node and "2" represents the identifier of the second computing node.

[0277] Step 505b: The second computing node sends the first connection relationship associated with the second computing node to the second node.

[0278] Each computing node (including the first node) sends the first connection relationship corresponding to that computing node to the second node.

[0279] Step 506: The second node aggregates (organizes) the first connection relationships from each computing node to obtain the connection relationships between computing nodes.

[0280] For example, delete redundant connections.

[0281] For example, the first connection relationship corresponding to computing node 1 is that computing node 1 is connected to computing nodes 2 and 3; the first connection relationship corresponding to computing node 2 is that computing node 2 is connected to computing nodes 1 and 4; the first connection relationship corresponding to computing node 3 is that computing node 1 is connected to computing node 4; and the first connection relationship corresponding to computing node 4 is that computing node 4 is connected to computing nodes 2 and 3. The fourth connection relationship after the aggregation of the second node is: computing node 1 is connected to computing node 2, computing node 1 is connected to computing node 3, computing node 4 is connected to computing node 2, and computing node 4 is connected to computing node 3.

[0282] Optionally, the second node stores the connection relationships between computing nodes.

[0283] Step 507: The second node sends the connection relationship between the computing nodes to the third node. Optionally, the second node also sends the amount of data transmitted unidirectionally or bidirectionally between computing nodes during a single round of model training to the third node.

[0284] Optionally, the second node does not perform aggregation processing, but directly sends the first connection relationship corresponding to each computing node to the third node. The third node aggregates (organizes) the first connection relationship from each computing node to obtain the connection relationship between computing nodes and saves it.

[0285] After learning about the connection relationships between computing nodes and the amount of data transmitted unidirectionally or bidirectionally between computing nodes during a single round of model training, the third node can optimize the parameters subsequently sent to the second node by combining historical network traffic information. The parameters sent to the second node include, but are not limited to, one or more of the following: the time requirement for transmitting parameters of the first model between computing nodes, the time requirement for training the first model in a single round, the total time requirement for training the first model, the upper limit of the number of computing nodes that a computing node is allowed to connect to (the upper limit of the number of neighbors), the strategy for determining the connection relationship between computing nodes (i.e., the connection relationship between computing nodes is determined by the computing node or the node that manages the computing nodes), and the period for determining the connection relationship between computing nodes.

[0286] In one scenario, the first node manages the computing nodes.

[0287] Referring to step 606 of Figure 6: The first node sends the first connection relationship between the computing nodes to the second node, and the second node receives the first connection relationship accordingly.

[0288] Optionally, the second node may also send the amount of data transmitted unidirectionally or bidirectionally between computing nodes during a single round of model training to the third node.

[0289] The second node can store the first connection relationship, the amount of data transmitted unidirectionally, or the amount of data transmitted bidirectionally.

[0290] After learning about the connection relationships between computing nodes and the amount of data transmitted unidirectionally or bidirectionally between computing nodes during a single round of model training, the second node can optimize the parameters subsequently sent to the first node by combining historical network traffic information. The parameters sent to the first node include, but are not limited to, one or more of the following: the time requirement for transmitting parameters of the first model between computing nodes, the time requirement for training the first model in a single round, the total time requirement for training the first model, the upper limit of the number of computing nodes that a computing node is allowed to connect to (the upper limit of the number of neighbors), the strategy for determining the connection relationship between computing nodes (i.e., the connection relationship between computing nodes is determined by the computing node or the node that manages the computing nodes), and the period for determining the connection relationship between computing nodes.

[0291] Figures 5 and 6 above illustrate how the EMS sends the determined connection relationships between computing nodes used to transmit the parameters of the first model to the NMS. The NMS saves these connection relationships for optimization of subsequent parameters sent to the EMS, and does not affect the connection relationships between computing nodes during this model training.

[0292] In one possible implementation, the EMS sends the determined connection relationships between computing nodes for transmitting parameters of the first model to the NMS. The NMS adjusts these connection relationships and indicates the adjusted relationships to the EMS, which then triggers the establishment of connections between the computing nodes based on the adjusted relationships.

[0293] Figure 7 illustrates a flowchart of a communication method. It should be noted that Figure 7 uses NMS-NMS interaction as an example; the EMS in Figure 7 can also be replaced by a model training provider, and the NMS can be replaced by a model training user.

[0294] Step 701: EMS determines the connection relationships between computing nodes.

[0295] One approach is as follows: a compute node determines the connection relationships associated with it and sends these relationships to the EMS. The EMS then aggregates these connections to obtain the connection relationships between the compute nodes. For example, refer to steps 505a, 505b, and 506 in Figure 5.

[0296] Another approach is for EMS to directly determine the connection relationships between computing nodes. For example, refer to step 603 in Figure 6.

[0297] Step 702: EMS sends the connection relationship between computing nodes to NMS, and the amount of data transmitted unidirectionally or bidirectionally between computing nodes during a single round of model training.

[0298] Alternatively, EMS can send the data volume of the first model's parameters and the model training method to NMS, and NMS can determine the data volume for one-way transmission and two-way transmission based on the data volume of the first model's parameters and the model training method.

[0299] Step 703: The NMS sends an indication message to the EMS, which is used to indicate the third connection relationship between computing nodes.

[0300] The third connection relationship may or may not change compared to the connection relationship sent by EMS to NMS in step 702.

[0301] If the connection remains unchanged, the indication information may include the connection relationship from step 702, or it may be used to indicate that the connection relationship has not changed. For example, a single bit value can be used to represent that the connection relationship has not changed.

[0302] If a change occurs, the indication information may include the changed connections between computing nodes (i.e., both changed and unchanged connections are included), or the indication information may include connections that differ from those before the change.

[0303] Before step 703, NMS determines the third connection relationship. For example, NMS can determine the third connection relationship based on historical network traffic information, the connection relationship between computing nodes in step 702, and the amount of data transmitted unidirectionally (or bidirectionally) between computing nodes in a single round of model training.

[0304] In step 702 of NMS query, the connection relationship between any two computing nodes is determined. Based on historical network traffic information, the historical network bandwidth usage between these two computing nodes is determined. It is then determined whether the historical network bandwidth usage (e.g., historical average available bandwidth) meets the training time requirement proposed by NMS (i.e., the first duration). If it does, the connection between the two computing nodes is retained; otherwise, the connection is deleted. Optionally, a new connection relationship is selected for the computing node whose connection relationship was deleted. For example, this can be based on the historical network bandwidth usage between the computing node and directly connected computing nodes; for instance, the directly connected computing node with the largest historical average available bandwidth is selected as the new neighbor. Based on this process, a third connection relationship is obtained.

[0305] Step 704: EMS triggers the establishment of connections between computing nodes based on the third connection relationship between computing nodes.

[0306] The specific details of triggering the establishment of connections between computing nodes can be found in the previous text, and will not be repeated here.

[0307] Of course, the third connection relationship may change compared to the connection relationship sent by EMS to NMS in step 702. In this embodiment, it is not ruled out that EMS may not adopt the third connection relationship and may still use the connection relationship sent by EMS to NMS in step 702 to trigger the establishment of a connection between computing nodes.

[0308] In one possible implementation, the NMS can send a proposed fourth connection relationship for transmitting parameters of the first model between computing nodes to the EMS. The EMS can determine the connection relationship between computing nodes based on the proposed fourth connection relationship. Then, based on the determined connection relationship, the EMS triggers the establishment of a connection between computing nodes.

[0309] NMS can determine the connection relationship between computing nodes based on the available bandwidth between computing nodes, the direct communication relationship between computing nodes, and the amount of data transmitted unidirectionally (or bidirectionally) between computing nodes in a single round of model training.

[0310] The principle by which NMS determines the fourth connection relationship is the same as that by EMS determining the first connection relationship. The differences include: EMS determines it based on the available bandwidth between computing nodes, while NMS can determine it based on the historical network bandwidth usage between computing nodes (such as the historical average available bandwidth). Other technical details will not be repeated here.

[0311] EMS can determine the connection relationship between compute nodes based on the proposed fourth connection relationship, which is the same principle as NMS determining the third connection relationship in step 702. The area includes: in step 702, NMS determines the connection relationship based on the historical network bandwidth usage between compute nodes (e.g., historical average available bandwidth), while EMS can determine the connection relationship based on the available bandwidth between compute nodes.

[0312] For example, in the fourth step of the EMS query, the connection relationship between any two computing nodes is determined; it is determined whether the available bandwidth between these two computing nodes meets the training time requirement (i.e., the first duration) proposed by NMS. If it does, the connection between these two computing nodes is retained; if not, the connection relationship between these two computing nodes is deleted. Optionally, a new connection relationship is selected for the computing node whose connection relationship has been deleted. For example, the selection is based on the available bandwidth between the computing node and the computing node with the direct communication connection. For example, the computing node with the largest available bandwidth in the direct communication connection is selected as the new neighbor.

[0313] In one possible implementation, after the first model training is completed, the EMS determines the actual training efficiency and the convergence performance of the training, and sends the actual training efficiency and the model convergence performance to the NMS.

[0314] If the connection relationships between computing nodes and the amount of data transmitted unidirectionally or bidirectionally between computing nodes during a single round of model training have not been sent to NMS before, they can be sent to NMS along with the actual training efficiency and the model's convergence performance after model training.

[0315] Actual training efficiency can be expressed in the following ways: whether the first duration is met, the relevant duration of transmitting the parameters of the first model between computing nodes during the entire model training process (e.g., average duration, maximum duration, minimum duration, etc.), the relevant duration of single-round model training during the entire model training process (e.g., average duration, maximum duration, minimum duration, etc.), and the total training time of the first model.

[0316] The convergence performance of a model represents its accuracy, which can be expressed as a loss sequence or value.

[0317] Once NMS learns about the actual training efficiency and model convergence performance, it can combine historical network traffic information to optimize the parameters subsequently sent to the first node. The parameters sent to the first node include, but are not limited to, one or more of the following: the time requirement for transmitting the parameters of the first model between computing nodes, the time requirement for a single round of training the first model, the total time requirement for training the first model, the upper limit of the number of computing nodes that a computing node is allowed to connect to (the upper limit of the number of neighbors), the strategy for determining the connection relationship between computing nodes (i.e., the connection relationship between computing nodes is determined by the computing node or the node that manages the computing nodes), and the period for determining the connection relationship between computing nodes.

[0318] It is understood that, in order to achieve the functions in the above embodiments, the first node and the second node include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0319] Figures 8 and 9 are schematic diagrams of possible communication devices provided in embodiments of this application. These communication devices can be used to implement the functions of the first node or the second node in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments.

[0320] As shown in Figure 8, the communication device 800 includes a processing unit 810 and a transceiver unit 820.

[0321] For example, the communication device 800 is used to implement the functions of the first node or the second node in the method embodiments shown in Figures 4 to 7. The transceiver unit 820 can perform the receiving and sending actions performed by the first node or the second node in the method embodiments. The processing unit 810 can perform other actions besides the sending and receiving actions performed by the first node or the second node in the method embodiments.

[0322] For example, when the communication device 800 is used to implement the function of the first node in the method embodiment shown in FIG4, the transceiver unit 820 is used to: receive information for a first duration. The processing unit 810 is used to trigger the establishment of a connection between the first computing node and the second computing node, etc.

[0323] For example, when the communication device 800 is used to implement the function of the second node in the method embodiment shown in FIG4, the transceiver unit 820 is used to: send information for a first duration.

[0324] A more detailed description of the processing unit 810 and the transceiver unit 820 can be obtained directly from the relevant descriptions in the method embodiments shown in Figures 4 to 7, and will not be repeated here. The processing unit 810 can be implemented by a processor, and the transceiver unit 820 can be implemented by a transceiver.

[0325] It should be understood that the division of units in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, all units in the device can be implemented entirely through software calls from processing elements; all units can be implemented entirely in hardware; or some units can be implemented through software calls from processing elements, and others in hardware. For example, each unit can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as a program in memory, called and executed by a processing element of the device. Moreover, these units can be fully or partially integrated together, or implemented independently. The processing unit here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, the operations of the above methods or the various units can be implemented through integrated logic circuits in the processor element or through software calls from processing elements.

[0326] In one example, a unit in any of the above devices can be one or more integrated circuits configured to implement the methods described above, such as: one or more application-specific integrated circuits (ASICs), or one or more microcontroller units (MCUs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these forms of integrated circuits. As another example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element can be a processor, such as a general-purpose central processing unit (CPU), or other processor capable of calling programs. Furthermore, these units can be integrated together to implement a system-on-a-chip (SOC).

[0327] The receiving unit described above is an interface circuit of this device, used to receive signals from other devices. For example, when the device is implemented as a chip, this unit is an interface circuit for the chip to receive signals from other chips or devices. The transmitting unit described above is an interface circuit of this device, used to transmit signals to other devices. For example, when the device is implemented as a chip, this unit is an interface circuit for the chip to transmit signals to other chips or devices.

[0328] As shown in Figure 9, the communication device 900 includes a processor 910 and an interface circuit 920. The processor 910 and the interface circuit 920 are coupled to each other. It is understood that the interface circuit 920 can be a transceiver or an input / output interface. Optionally, the communication device 900 may also include a memory 930 for storing instructions executed by the processor 910, or storing input data required for the processor 910 to execute instructions, or storing data generated after the processor 910 executes instructions. Sometimes, the interface circuit 920 can also be understood as part of the processor 910, in which case the communication device 900 includes the processor 910.

[0329] When the communication device 900 is used to implement the methods shown in Figures 4 to 7, the processor 910 is used to implement the functions of the processing unit 810, and the interface circuit 920 is used to implement the functions of the transceiver unit 820.

[0330] When the aforementioned communication device is a chip applied to a terminal device, the terminal device chip implements the functions of the terminal device in the above method embodiments. The terminal device chip receives information from a network device, which can be understood as the information being first received by other modules (such as an RF module or antenna) in the terminal device, and then sent to the terminal device chip by these modules. The terminal device chip sends information to a network device, which can be understood as the information being first sent to other modules (such as an RF module or antenna) in the terminal device, and then sent to the network device by these modules.

[0331] When the aforementioned communication device is a chip applied to a network device, the network device chip implements the functions of the network device in the above method embodiments. The network device chip receives information from the terminal device, which can be understood as the information being first received by other modules (such as radio frequency modules or antennas) in the network device, and then sent to the network device chip by these modules. The network device chip sends information to the terminal device, which can be understood as the information being sent down to other modules (such as radio frequency modules or antennas) in the network device, and then sent to the terminal device by these modules. Here, the network device module can be the baseband chip of the network device, or a DU (Digital Unit) or other modules. The DU here can be a DU under the Open Radio Access Network (O-RAN) architecture.

[0332] In this application, entity A sends information to entity B, either directly or indirectly through other entities. Similarly, entity B receives information from entity A, either directly or indirectly through other entities. Entities A and B can be network devices or terminal devices, or modules within network devices or terminal devices. The sending and receiving of information can be between network devices and terminal devices, between two network devices (e.g., CU and DU), or between different modules within a single device (e.g., a terminal device chip and other modules within the terminal device, or a network device chip and other modules within the network device).

[0333] It is understood that the processor in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.

[0334] This application also provides a computer-readable storage medium storing a computer program that, when executed by a computer, enables the computer to perform the aforementioned communication method. Alternatively, the computer program includes instructions for implementing the aforementioned communication method.

[0335] This application also provides a chip including a processor. When the processor executes a computer program or instructions, it implements the communication method provided above. Optionally, the chip may further include a memory. The chip may be composed of a single chip or may include chips and other discrete devices. The memory is used to store computer programs or instructions.

[0336] The application also provides a circuit for performing the communication method described above. This circuit may include a chip circuit. Optionally, the circuit may also be coupled to a memory.

[0337] This application also provides a computer program product, including: computer program code, which, when run on a computer, enables the computer to execute the communication method provided above.

[0338] This application also provides a communication system, which includes a first computing and a second node that execute the above-described communication method.

[0339] This application also provides a communication system, which includes multiple computing nodes for training a first model.

[0340] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a base station or terminal. Of course, the processor and storage medium can also exist as discrete components in the base station or terminal.

[0341] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a first control plane network element, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video optical disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both types of storage media.

[0342] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0343] In this application embodiment, the number of nouns, unless otherwise specified, refers to "singular nouns or plural nouns," that is, "one or more." "At least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A or B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. For example, A / B means: A or B. Expressions such as "at least one of the following" or "one or more of them" refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and / or c, or one or more of a, b and / or c, means: a, b, c, a and b, a and c, b and c, or a and b and c. Each of a, b, and c can be single or multiple.

[0344] The ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects, and are not used to limit the size, content, order, timing, priority, or importance of the multiple objects. Furthermore, such names do not indicate differences in the content, sending / receiving end, sending order, size, application scenario, priority, or importance of the two pieces of information. Additionally, the numbering of steps in the various embodiments described in this application is only to distinguish different steps and is not used to limit the order of steps.

Claims

1. A communication method, characterized in that, Applied to the first node, including: Receive information about the first duration from the second node; wherein the first duration is the time requirement for transmitting parameters of the first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model; Based on the first duration, a connection is established between the first computing node and the second computing node; wherein, the computing nodes for training the first model include at least the first computing node and the second computing node, and the connection between the first computing node and the second computing node is used to transmit the parameters of the first model between the first computing node and the second computing node.

2. The method as described in claim 1, characterized in that, The first node is either the first computing node or the second computing node; or, The first node manages the computing nodes.

3. The method as described in claim 1 or 2, characterized in that, The step of triggering the establishment of a connection between the first computing node and the second computing node based on the first duration includes: Based on the first duration, the available bandwidth between the first computing node and the second computing node, and the data volume of the parameters of the first model, a connection between the first computing node and the second computing node is triggered.

4. The method as described in claim 3, characterized in that, The first duration is the required duration for transmitting parameters of the first model between computing nodes; wherein, given the available bandwidth and the data volume of the parameters of the first model, the first estimated duration for bidirectional transmission of the parameters of the first model between the first computing node and the second computing node satisfies the duration requirement; or, The first duration is the required duration for a single round of training the first model; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the second estimated duration for the first computing node and the second computing node to jointly train the first model for one round satisfies the single-round duration requirement; or... The first duration is the total duration required to train the first model; wherein, given the available bandwidth and the amount of data for the parameters of the first model, the third estimated duration for the first computing node and the second computing node to jointly train the first model satisfies the total duration requirement.

5. The method according to any one of claims 1-4, characterized in that, The step of triggering the establishment of a connection between the first computing node and the second node based on the first duration includes: Based on the first duration and the first quantity, a connection is established between the first computing node and the second computing node; wherein, the first quantity is used to indicate the upper limit of the number of computing nodes that a computing node is allowed to connect to, and the connection is used to transmit parameters of the first model.

6. The method according to any one of claims 1-5, characterized in that, The step of triggering the establishment of a connection between the first computing node and the second computing node based on the first duration includes: Based on the first duration, a first connection relationship between the computing nodes is determined; wherein, the first connection relationship is a connection relationship used for transmitting parameters of the first model between the computing nodes, and the first connection relationship includes the connection relationship between the first computing node and the second computing node; Based on the first connection relationship, a connection between the first computing node and the second computing node is established.

7. The method as described in claim 6, characterized in that, Determining the first connection relationship between the computing nodes based on the first duration includes: Based on the first duration and the first quantity, a first connection relationship between the computing nodes is determined; wherein, the first quantity is used to indicate the upper limit of the number of computing nodes that a computing node is allowed to connect to.

8. The method according to any one of claims 1-7, characterized in that, The first node is the first computing node, and triggering the establishment of a connection between the first computing node and the second computing node includes: The identifier of the second computing node is stored in a first list, which is used to store the identifiers of computing nodes that exchange parameters of the first model with the first computing node; or, The first node is the second computing node, and the step of triggering the establishment of a connection between the first computing node and the second computing node includes: The identifier of the first computing node is stored in the second list. The second list is used to store the identifiers of computing nodes that transmit parameters of the first model to the second computing node.

9. The method according to any one of claims 1-7, characterized in that, The first node manages the computing nodes; The triggering of establishing a connection between the first computing node and the second computing node includes: Send a first indication message to the first computing node, the first indication message being used to instruct the first computing node and the second computing node to establish a connection for transmitting parameters of the first model; And / or, Send a second instruction message to the second computing node, the second instruction message being used to instruct the first computing node and the second computing node to establish a connection for transmitting parameters of the first model.

10. The method according to any one of claims 1-9, characterized in that, The first node is either the first computing node or the second computing node; the method further includes: Receive a first strategy from the second node, the first strategy being used to instruct the computing node to determine the connection relationship for transmitting the parameters of the first model; or, The first node manages the computing nodes; the method further includes: A second strategy is received from the second node, which instructs the node managing the computing nodes to determine the connection relationships between the computing nodes for transmitting parameters of the first model.

11. The method according to any one of claims 1-10, characterized in that, After triggering the establishment of a connection between the first computing node and the second computing node based on the first duration, the method further includes: Based on the first duration, a second connection relationship is determined between the computing nodes, wherein the second connection relationship is a connection relationship used for transmitting parameters of the first model between the computing nodes; Based on the second connection relationship, the connections between the computing nodes are updated.

12. The method as described in claim 11, characterized in that, Determining the second connection relationship between the computing nodes based on the first duration includes: Based on the first cycle and the first duration, a second connection relationship between the computing nodes is determined.

13. The method according to any one of claims 1-12, characterized in that, The method further includes: Receive one or more of the following information from the second node: The first quantity information indicates the upper limit of the number of computing nodes that a computing node is allowed to connect to, the connection being used to transmit parameters of the first model; Information for the first period, wherein the first period is used to indicate the period for determining the connection relationship between the computing nodes, the connection relationship being the connection relationship used for transmitting parameters of the first model between the computing nodes; Available bandwidth between the computing nodes; The third connection relationship between the computing nodes is a connection relationship used for transmitting parameters of the first model between the computing nodes.

14. The method as described in claim 6 or 7, characterized in that, The method further includes: Send one or more of the following messages to the second node: The first connection relationship, the amount of data on the parameters of the first model transmitted bidirectionally between the computing nodes, or the amount of data on the parameters of the first model transmitted unidirectionally between the computing nodes, the convergence performance of the first model, the relevant time for transmitting the parameters of the first model between the computing nodes during the entire model training process, the relevant time for a single round of model training during the entire model training process, and the total training time of the first model.

15. A communication method, characterized in that, Applied to the second node, including: Send information of a first duration to the first node; wherein, the first duration is the time requirement for transmitting parameters of the first model between computing nodes, or the time requirement for a single round of training the first model, or the total time requirement for training the first model; the first duration is used by the first node to trigger the establishment of a connection between the first computing node and the second computing node, the computing nodes for training the first model include at least the first computing node and the second computing node, and the connection between the first computing node and the second computing node is used for transmitting parameters of the first model between the first computing node and the second computing node.

16. The method as described in claim 15, characterized in that, The first node is either the first computing node or the second computing node; or, The first node manages the computing nodes.

17. The method as described in claim 15 or 16, characterized in that, The first node is either the first computing node or the second computing node, and the second node manages the computing node; the method further includes: Receive information about the first duration from the third node; or, The first node manages the computing node, and the second node manages the first node; the method further includes: Determine the first duration.

18. The method according to any one of claims 15-17, characterized in that, The method further includes: Send one or more of the following information to the first node: The first quantity information is used to indicate the upper limit of the number of computing nodes that a computing node is allowed to connect to, the connection being used to transmit parameters of the first model; Information for the first period, wherein the first period is used to indicate the period for determining the connection relationship between the computing nodes, the connection relationship being the connection relationship used for transmitting parameters of the first model between the computing nodes; Available bandwidth between the computing nodes; A first strategy is used to instruct computing nodes to determine the connection relationships used to transmit parameters of the first model. The second strategy is used to instruct the node for managing the computing node to determine the connection relationship between the computing nodes for transmitting parameters of the first model; The third connection relationship between the computing nodes is a connection relationship used for transmitting parameters of the first model between the computing nodes.

19. The method as described in claim 18, characterized in that, The first node is either the first computing node or the second computing node, and the second node manages the computing node; the method further includes: Receive one or more pieces of information from the third node; or, The first node manages the computing node, and the second node manages the first node; the method further includes: To determine one or more of the aforementioned information.

20. The method according to any one of claims 15-19, characterized in that, Also includes: Receive one or more of the following information from the first node: The first connection relationship between computing nodes, the amount of data for bidirectional transmission of the parameters of the first model between computing nodes, or the amount of data for unidirectional transmission of the parameters of the first model between computing nodes, the convergence performance of the first model, the relevant time for transmission of the parameters of the first model between computing nodes during the entire model training process, the relevant time for single-round model training during the entire model training process, and the total training time of the first model.

21. The method as described in claim 20, characterized in that, The first node manages the computing node, and the second node manages the first node; the method further includes: Based on the first parameter, a third connection relationship is determined; wherein the first parameter includes one or more of the following: The historical network traffic of the computing nodes, the first connection relationship between the computing nodes, and the amount of data of parameters of the first model transmitted unidirectionally or bidirectionally between the computing nodes.

22. The method as described in claim 20, characterized in that, The first node is either the first computing node or the second computing node, and the second node manages the computing node; the method further includes: The first connection relationship between the computing nodes is aggregated, and the aggregated connection relationship between the computing nodes is sent to the third node.

23. A communication device, characterized in that, Includes modules for performing the method as described in any one of claims 1-14 or any one of claims 15-22.

24. A communication device, characterized in that, Includes a processor, which is coupled to a memory; The memory is used to store computer programs or instructions; The processor is configured to execute some or all of the computer programs or instructions in the memory, and when the some or all of the computer programs or instructions are executed, to implement the method as described in any one of claims 1-14, or the method as described in any one of claims 15-22.

25. A communication device, characterized in that, Including processor and memory; The memory is used to store computer programs or instructions; The processor is configured to execute some or all of the computer programs or instructions in the memory, and when the some or all of the computer programs or instructions are executed, to implement the method as described in any one of claims 1-14, or the method as described in any one of claims 15-22.

26. A chip system, characterized in that, Includes a processor and interface circuitry, wherein the processor is coupled to a memory; The memory is used to store computer programs or instructions; The processor is configured to execute some or all of the computer programs or instructions in the memory, and when the some or all of the computer programs or instructions are executed, to implement the method as described in any one of claims 1-14, or the method as described in any one of claims 15-22.

27. A computer-readable storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a communication device, implement the method as described in any one of claims 1-14, or the method as described in any one of claims 15-22.

28. A computer program product, characterized in that, The computer program product includes: computer instructions that, when executed on a computer, cause the method as described in any one of claims 1-14, or the method as described in any one of claims 15-22, to be implemented.

Citation Information

Patent Citations

  • Method and system for accelerating distributed machine learning

    CN111210020A

  • Model parameter scheduling method and device

    CN118133928A

  • Systems and methods for distributed learning for wireless edge dynamics

    US20230068386A1

  • Communication method and apparatus

    WO2023231620A1