Distributed training method and apparatus

The difference is characterized by the child node feedback model, and the central node adjusts the training indication, solving the problems of non-independent and homogeneous data and unstable transmission in distributed training, improving model performance and system stability.

WO2025131014A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140796
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In distributed training, since the central node cannot access user data, the data of the child nodes is non-independent and homogeneously distributed, the model performance cannot be guaranteed, and the unstable wireless transmission feedback gradient causes fluctuations in model parameter updates.

Method used

By allowing the child nodes to feedback the characterization differences between models to the central node, the central node conducts model training instructions based on the differences, allowing the child nodes to obtain a high-performance machine learning model in a limited number of communications.

Benefits of technology

The performance of distributed training is improved, the dispersion of the child node update direction is reduced, the impact of non-independent homogeneous distributed data on system performance is reduced, and the deviation caused by packet loss is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140796_26062025_PF_FP_ABST
    Figure CN2024140796_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Artificial intelligence (AI) model distributed training methods and an apparatus. A sub-node participating in training a model of a central node feeds back to the central node a characterizing difference between models and, on the basis of the characterizing difference, the central node performs a training indication of the model, so that a high-performance machine learning model can be obtained in a limited number of communication processes between sub-nodes and the central node, improving the performance of distributed training.
Need to check novelty before this filing date? Find Prior Art

Description

Distributed training method and device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 21, 2023, with application number 202311777211.1 and invention name “Distributed Training Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of communication technology, and in particular to a distributed training method and device. Background Art

[0003] Distributed training refers to dividing the machine training process into multiple sub-computing nodes (referred to as "sub-nodes"). It allows the central node to collect machine learning models trained by multiple sub-nodes to improve the effectiveness of the entire machine learning training.

[0004] Distributed training aims to achieve high-performance machine learning models through a limited number of child nodes communicating with a central node. However, because the central node cannot access user data, model performance cannot be guaranteed when the child node data is not independent and identically distributed. Furthermore, when child nodes feed back gradients via unstable wireless transmission, the number of gradients successfully received by the central node is somewhat random, which can cause fluctuations in the central node's model parameter updates.

[0005] In view of this, how to improve the performance of distributed training is an urgent problem to be solved. Summary of the Invention

[0006] The present application provides a distributed training method and apparatus to improve the performance of distributed training.

[0007] In a first aspect, a distributed training method is provided, the method comprising: receiving first information from a first child node, the first information being used to indicate a first representation difference between a model of a central node and a model of the first child node, the first child node being any one of a plurality of child nodes participating in the model training of the central node; and sending second information to the first child node, the second information being obtained based on the first representation difference, the second information including training indication information of the model of the first child node.

[0008] In this aspect, the central node receives the representation differences between the models fed back by the child nodes participating in the training of the central node model, and the central node gives training instructions for the model based on the representation differences, so that a high-performance machine learning model can be obtained in the process of a limited number of child nodes communicating with the central node, thereby improving the performance of distributed training.

[0009] In combination with the first aspect, in a possible implementation, the method further includes: broadcasting third information, the third information including at least one of the following: model configuration information, a public data set; wherein the model configuration information is used to indicate at least one of the following: the type of the model of the multiple child nodes, the structural information of the model of the multiple child nodes, the model parameters of the model of the multiple child nodes, or the training configuration information of the model of the multiple child nodes.

[0010] In this implementation, the third information is used for model training of the first child node.

[0011] By having each child node calculate the first representation difference based on the same public data set, the representation difference between the model of each child node and the model of the central node, or the representation difference between the models of each child node, can be more accurately compared.

[0012] In a second aspect, a distributed training method is provided, the method comprising: sending first information to a central node, the first information being used to indicate a first representation difference between a model of the central node and a model of a first child node, the first child node being any one of a plurality of child nodes participating in the model training of the central node; and receiving second information from the central node, the second information being obtained based on the first representation difference, the second information including training indication information of the model of the first child node.

[0013] In this aspect, the sub-nodes participating in the training of the central node model feed back the representation differences between the models to the central node, and the central node gives training instructions for the model based on the representation differences, so that a high-performance machine learning model can be obtained in the process of a limited number of communications between the sub-nodes and the central node, thereby improving the performance of distributed training.

[0014] In combination with the second aspect, in a possible implementation, the method further includes: updating the model of the first subnode according to the second information.

[0015] In combination with the second aspect, in another possible implementation, the method also includes: receiving third information, the third information including at least one of the following: model configuration information, a public data set; wherein the model configuration information is used to indicate at least one of the following: the type of the model of the multiple child nodes, the structural information of the model of the multiple child nodes, the model parameters of the model of the multiple child nodes, or the training configuration information of the model of the multiple child nodes.

[0016] In this implementation, the third information is used for model training of the first child node.

[0017] By having each child node calculate the first representation difference based on the same public data set, the representation difference between the model of each child node and the model of the central node, or the representation difference between the models of each child node, can be more accurately compared.

[0018] In combination with the first aspect, the second aspect, or any one of the first and second aspects, in another possible implementation, the first representation difference satisfies a first condition, and the training instruction information is used to instruct the first child node to continue training the model of the first child node. Optionally, the first representation difference satisfies the first condition includes: the first representation difference is less than or equal to a first threshold. Optionally, the first condition, such as the first threshold, is predetermined by the central node or predefined by a protocol.

[0019] In combination with any one of the first aspect, the second aspect, or the first aspect and the second aspect, in another possible implementation, the first characterization difference satisfies the second condition, and the training instruction information is used to instruct the first child node to stop training the model of the first child node. Optionally, the first characterization difference satisfies the first condition including: the first characterization difference is greater than or equal to a second threshold. Optionally, the second threshold is the same as or different from the aforementioned first threshold. Optionally, the second condition, such as the second threshold, is predetermined by the central node or predefined by the protocol.

[0020] In combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the training indication information is also used to indicate the representation difference coefficient in the loss function of the model of the first sub-node.

[0021] In this implementation, the central node can also assign corresponding loss function parameters, such as characterizing the difference coefficient, to sub-nodes with large difference values ​​based on the difference value analysis results.

[0022] In combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the second information also includes a representation difference coefficient in the loss function of the model of the first sub-node.

[0023] In combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the greater the first characterization difference, the greater the characterization difference coefficient.

[0024] In this implementation, the larger the first representation difference, the larger the representation difference coefficient. This is because a larger first representation difference indicates that the model trained by the child node deviates further from the model of the central node, or that the model trained by the child node deviates further from the models trained by other child nodes among all child nodes participating in the central node's model training. Therefore, the central node assigns a representation difference coefficient so that the child node can update its loss function based on the representation difference coefficient, so that subsequent updates to the child node's model have less impact on the training results.

[0025] In combination with the first aspect, the second aspect, or any one implementation of the first aspect and the second aspect, in another possible implementation, the first representation difference is obtained based on a local data set or a public data set of the first child node.

[0026] In this implementation, by having each child node calculate the first representation difference based on the same local data set or public data set, the representation difference between the model of each child node and the model of the central node, or the representation difference between the models of each child node, can be more accurately compared.

[0027] In combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the first characterization difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, and / or the first characterization difference is the difference between the output or intermediate quantity of the models of the multiple child nodes, wherein the output or intermediate quantity is obtained based on the same input.

[0028] In this implementation, by obtaining the difference between the outputs or intermediate quantities of the models based on the same input as the representation difference of the models, the difference between the models can be accurately represented.

[0029] The method of the first aspect described above may be executed by a central node, or by a module (such as a processor, chip, or chip system) applied to the central node, or by a logical node, logical module, or software that can implement all or part of the functions of the central node.

[0030] The method of the second aspect described above may be executed by a sub-node, or by a module (such as a processor, chip, or chip system) applied to the sub-node, or by a logical node, logical module, or software that can implement all or part of the sub-node functions.

[0031] In a third aspect, a distributed training method is provided, which includes: sending first information, the first information including at least one of the following: model configuration information, a public data set, a threshold update rule, and a representation difference coefficient update rule, the threshold is used to compare the representation difference between the model of the central node and the model of the first child node, and the representation difference coefficient is a parameter in the loss function of the model of the first child node; and receiving second information, the second information including model information obtained by training the first child node.

[0032] In this aspect, the central node sends model configuration information, public data sets, threshold update rules, and representation difference coefficient update rules to the child nodes. The child nodes can calculate the representation difference between the central node model and the child node model by themselves, and compare the representation difference and the threshold to determine whether to continue or stop model training. If the model training continues, the model parameters are updated using the loss function with the representation difference coefficient. By constraining the model representation difference, the dispersion of the child node update direction can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the deviation caused by packet loss can be reduced, the final performance of the central node model can be improved, and the performance of distributed training can be improved.

[0033] In combination with the third aspect, in a possible implementation, the method further includes: updating a model of the central node according to the second information.

[0034] In a fourth aspect, a distributed training method is provided, which includes: calculating a first representation difference between a model of a central node and a model of the first child node based on a local data set or a public data set of the first child node; when the first representation difference is less than or equal to a first threshold, continuing the training of the model of the first child node; and updating the model of the first child node using the local data set and a loss function with a representation difference coefficient; wherein, the greater the first representation difference, the greater the representation difference coefficient.

[0035] In this regard, the child node can calculate the representation difference between the model of the central node and the model of the child node by itself, and compare the representation difference with the threshold to determine whether to continue or stop the model training. If the model training continues, the model parameters are updated using the loss function with the representation difference coefficient. By constraining the model representation difference, the dispersion of the child node update direction can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the deviation caused by packet loss can be reduced, the final performance of the central node model can be improved, and the performance of distributed training can be improved.

[0036] In combination with the fourth aspect, in a possible implementation, the method further includes: receiving first information, the first information including at least one of the following: model configuration information, a public data set, a threshold update rule, and an update rule for the characterization difference coefficient.

[0037] In this implementation, the update rule of the threshold value can be pre-established by the central node. For example, the update rule is established as follows: the threshold value decreases as the number of communications between the first child node and the central node increases, and the threshold value of the first child node remains unchanged before the first child node communicates with the central node. Among them, the central node sends a new model to the first child node once, which is one communication. It can be understood that as the child nodes are trained and the central node aggregates the models of the child nodes, the models of the central node and the child nodes gradually converge. Therefore, as the number of communications between the first child node and the central node increases, the representation difference between the model of the central node and the model of the first child node will gradually decrease. Therefore, a threshold update rule can be established in which the threshold value decreases as the number of communications between the first child node and the central node increases.

[0038] The updating rule of the characterization difference coefficient is used to instruct the child node to obtain an updated characterization difference coefficient based on the initial characterization difference coefficient and the calculated characterization difference after the characterization difference is calculated. The updating rule of the characterization difference coefficient is pre-established by the central node. For example, the updating rule of the characterization difference coefficient is: the linear or nonlinear scaling of the characterization difference coefficient of the previous update is used as the update amount of the characterization difference coefficient compared with the previous characterization difference coefficient. It can be understood that the larger the characterization difference is, the larger the characterization difference coefficient is. Because the larger the characterization difference is, it indicates that the model trained by the child node deviates far from the model of the central node, or the model trained by the child node deviates far from the model trained by other child nodes among all the child nodes participating in the model training of the central node. Therefore, the child node obtains an updated characterization difference coefficient, so that the child node can update its loss function according to the characterization difference coefficient, so that the update of the model of the subsequent child node has less impact on the training results.

[0039] In combination with the fourth aspect, in another possible implementation, the method further includes: updating the first threshold based on an update rule of the threshold to obtain a second threshold.

[0040] In combination with the fourth aspect, in another possible implementation, the method further includes: updating the characterization difference coefficient based on an update rule of the characterization difference coefficient.

[0041] In combination with the fourth aspect, in another possible implementation, the method further includes: when the first representation difference is greater than the first threshold, stopping the training of the model of the first sub-node.

[0042] In combination with the third aspect or the fourth aspect, in another possible implementation, the model configuration information is used to indicate at least one of the following: the types of models of multiple child nodes participating in the training of the model of the central node, structural information of the models of the multiple child nodes, model parameters of the models of the multiple child nodes, or training configuration information of the models of the multiple child nodes.

[0043] In yet another possible implementation, the threshold value is updated as follows: the threshold value decreases as the number of communications between the first sub-node and the central node increases.

[0044] In another possible implementation, the updating rule of the characterization difference coefficient is to use a linear or nonlinear scaling of the characterization difference coefficient of the previous update as the update amount of the characterization difference coefficient compared with the previous characterization difference coefficient.

[0045] In yet another possible implementation, the greater the representation difference, the greater the representation difference coefficient.

[0046] In a fifth aspect, a distributed training device is provided for implementing the distributed training method in any one of the implementations of the first aspect, the third aspect, or the first aspect and the third aspect. The device can be a central node, or a module applied to a central node (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of a central node. In one implementation, the distributed training device may include a sending unit, a receiving unit, and may also include a processing unit. The sending unit and the receiving unit may be independent or combined together (which may be referred to as a "transceiver unit").

[0047] In the sixth aspect, a distributed training device is provided for implementing the distributed training method in any one of the implementations of the second aspect, the fourth aspect, or the second aspect and the fourth aspect. The device can be a sub-node, or a module applied to a sub-node (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of a sub-node. In one implementation, the distributed training device may include a sending unit, a receiving unit, and may also include a processing unit. The sending unit and the receiving unit may be independent or combined together (which may be referred to as a "transceiver unit").

[0048] In a possible implementation, the distributed training device in the fifth to sixth aspects includes a module for respectively executing the method in any one of the first to second aspects or any one of the implementations.

[0049] In another possible implementation, the distributed training device in the fifth to sixth aspects above includes a processing circuit coupled to a memory; the processing circuit is configured to enable the device to perform the corresponding functions in the above-mentioned distributed training method. The memory is used to couple with the processing circuit, which stores the necessary programs (instructions) and / or data for the device. Optionally, the distributed training device may further include a communication interface for enabling communication between the device and other network elements. Optionally, the memory may be located inside the distributed training device or outside the distributed training device. Exemplarily, the processing circuit may be a processor or a circuit in a processor for processing.

[0050] When the distributed training device in the fifth and sixth aspects is a chip, the sending unit may be an output unit, such as an output circuit or a communication interface; the receiving unit may be an input unit, such as an input circuit or a communication interface. When the distributed training device is a terminal, the sending unit may be a transmitter or a transmitter; and the receiving unit may be a receiver or a receiver.

[0051] In a seventh aspect, a computer-readable storage medium is provided, in which a computer program or instruction is stored. When the computer program or instruction is executed, the methods described in the above aspects are implemented.

[0052] In an eighth aspect, a computer program product comprising instructions is provided, which, when executed on a distributed training device, causes the distributed training device to execute the methods described in the above aspects.

[0053] In a ninth aspect, a distributed training system is provided, which includes the distributed training device described in the fifth aspect and the distributed training device described in the sixth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] FIG1 is a schematic diagram of the architecture of a distributed training system provided in an embodiment of the present application;

[0055] FIG2 is a simplified schematic diagram of a wireless communication system provided by an embodiment of the present application;

[0056] FIG3 is a schematic diagram of the architecture of another distributed training system provided by the present application;

[0057] 4A to 4D are schematic diagrams of a network architecture according to an embodiment of the present application;

[0058] FIG5 is a schematic diagram of a neuron structure;

[0059] Figure 6 is a schematic diagram of a neural network;

[0060] Figure 7 is a schematic diagram of an AI application framework;

[0061] FIG8 is a schematic diagram of the architecture of another communication system provided in an embodiment of the present application;

[0062] FIG9 is a flow chart of a distributed training method according to an embodiment of the present application;

[0063] FIG10 is a flow chart of another distributed training method provided in an embodiment of the present application;

[0064] FIG11 is a schematic diagram of a calculation method for characterizing differences according to an embodiment of the present application;

[0065] FIG12 is a system block diagram of a distributed training example according to an embodiment of the present application;

[0066] FIG13 is a flow chart of another distributed training method provided in an embodiment of the present application;

[0067] FIG14 is a schematic structural diagram of a distributed training device provided in an embodiment of the present application;

[0068] FIG15 is a schematic structural diagram of another distributed training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0069] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0070] Embodiments of the present application can be applied to a distributed training system as shown in Figure 1 , which includes a central node and K child nodes, where K is a positive integer. For example, the distributed training system can be a federated learning system or a gossip learning system. The central node and each child node can transmit data, model information, and the like.

[0071] The machine learning models trained by this distributed training system can be for non-wireless communication services, such as image recognition, natural language processing, etc., or for wireless communication services, such as beam selection based on environmental information.

[0072] The technology provided by this application can be applied to various communication systems. For example, the communication system can be a fourth generation (4G) th generation, 4G) communication systems (such as long term evolution (LTE) systems), fifth generation (5 thgeneration (5G) communication systems, worldwide interoperability for microwave access (WiMAX), wireless local area network (WLAN) systems, satellite communication systems, integrated systems of multiple systems, or future communication systems such as the sixth generation (6 th generation, 6G) communication system, etc. Among them, the 5G communication system can also be called a new radio (NR) system.

[0073] A network element in a communication system can send a signal to another network element or receive a signal from another network element. The signal may include information, signaling, or data, etc. The network element can also be replaced by an entity, a network entity, a device, a terminal device, a communication module, a node, a communication node, etc. The present application uses the network element as an example for description. For example, the communication system may include at least one terminal device and at least one network device. The network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the network device. In addition, it can be understood that if the communication system includes multiple terminal devices, the multiple terminal devices can also send signals to each other, that is, the signal sending network element and the signal receiving network element can both be terminal devices.

[0074] Refer to Figure 2, which is a simplified schematic diagram of a wireless communication system provided in an embodiment of the present application. As shown in Figure 2, the wireless communication system includes a wireless access network 100. The wireless access network 100 can be a next-generation (e.g., 6G or higher) wireless access network, or a traditional (e.g., 5G, 4G) wireless access network. One or more terminal devices (120a-120j, collectively referred to as 120) can be connected to each other, or connected to one or more network devices (110a, 110b, collectively referred to as 110) in the wireless access network 100. Optionally, Figure 2 is only a schematic diagram, and the wireless communication system may also include other devices, such as core network devices, wireless relay devices and / or wireless backhaul devices, which are not shown in Figure 2.

[0075] Optionally, in actual applications, the wireless communication system may include multiple network devices (also called access network devices) and multiple terminal devices at the same time. A network device can serve one or more terminal devices at the same time. A terminal device can also access one or more network devices at the same time. The embodiments of the present application do not limit the number of terminal devices and network devices included in the wireless communication system.

[0076] The network device may be an entity on the network side for transmitting or receiving signals. The network device may be an access device for a terminal device to access the wireless communication system in a wireless manner, such as a base station. The base station can broadly cover various names as follows, or be replaced with the following names, such as: radio access network (RAN) node, NodeB, evolved NodeB (eNB), next generation NodeB (gNB), network equipment in open radio access network (O-RAN), relay station, access point, transmission point (TRP), transmitting point (TP), master eNB (MeNB), secondary eNB (SeNB), multi-standard radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, building baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), centralized unit (CU), distributed unit (DRU), etc. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network.The network device can support networks with the same or different access technologies. The embodiments of the present application do not limit the specific technology and specific device form used by the network device.

[0077] Network devices can be fixed or mobile. For example, base stations 110a and 110b are stationary and are responsible for wireless transmission and reception in one or more cells from terminal device 120. The helicopter or drone 120i shown in Figure 2 can be configured to act as a mobile base station, and one or more cells can move according to the location of the mobile base station 120i. In other examples, the helicopter or drone (120i) can be configured to act as a terminal device communicating with base station 110b.

[0078] In this application, the communication device used to implement the above-mentioned network access function can be a network device, or a network device with partial network access functions, or a device capable of supporting the implementation of the network access function, such as a chip system, a hardware circuit, a software module, or a hardware circuit plus a software module. The device can be installed in the network device or used in combination with the network device. In the method of this application, the communication device used to implement the network device function is described as an example of a network device.

[0079] A terminal device may be an entity on the user side for receiving or transmitting signals, such as a mobile phone. The terminal device may be used to connect people, objects, and machines. The terminal device may communicate with one or more core networks through a network device. The terminal device includes a handheld device with wireless connection capabilities, other processing devices connected to a wireless modem, or a vehicle-mounted device. The terminal device may be a portable, pocket-sized, handheld, computer-built-in, or vehicle-mounted mobile device. The terminal device 120 may be widely used in various scenarios, such as cellular communication, D2D, V2X, point-to-point (P2P), machine-to-machine (M2M), machine type communication (MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, telemedicine, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.Some examples of the terminal device 120 include: user equipment (UE) of the 3GPP standard, fixed equipment, mobile equipment, handheld equipment, wearable equipment, cellular phones, smart phones, session initiated protocol (SIP) phones, laptops, personal computers, smart books, vehicles, satellites, global positioning system (GPS) equipment, target tracking equipment, drones, helicopters, aircraft, ships, remote control equipment, smart home equipment, industrial equipment, personal communication service (PCS) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), wireless network cameras, tablet computers, handheld computers, mobile internet devices (MIDs), wearable devices such as smart watches, VR devices, AR devices, wireless terminals in industrial control, terminals in vehicle networking systems, wireless terminals in self-driving cars, wireless terminals in smart grids, wireless terminals in transportation safety, and smart cities. The terminal device 120 may be a wireless terminal in a city such as a smart gas pump, a terminal device on a high-speed rail, and a wireless terminal in a smart home, such as a smart speaker, a smart coffee machine, a smart printer, etc. The terminal device 120 may be a wireless device in the above various scenarios or a device for being set in a wireless device, for example, a communication module, a modem or a chip in the above device. The terminal device may also be referred to as a terminal, a terminal device, a UE, a mobile station (MS), a mobile terminal (MT), etc. The terminal device may also be a terminal device in a future wireless communication system. The terminal device may be used in a dedicated network device or a general device. The embodiments of the present application do not limit the specific technology and specific device form adopted by the terminal device.

[0080] Alternatively, a terminal device can function as a base station. For example, a UE can act as a dispatching entity, providing sidelink signals between UEs in V2X, D2D, or P2P scenarios. As shown in Figure 2, a cell phone 120a and a car 120b communicate with each other using sidelink signals. Cell phone 120a and smart home device 120e communicate without relaying the communication signal through base station 110b.

[0081] In this application, the communication device used to implement the functions of the terminal device can be a terminal device, or a terminal device with some of the functions of the above terminal devices, or a device that can support the implementation of the functions of the above terminal devices, such as a chip system, which can be installed in the terminal device or used in combination with the terminal device. In this application, the chip system can be composed of chips, or it can include chips and other discrete devices. In the technical solution provided in this application, the communication device is described as a terminal device or UE as an example.

[0082] Optionally, a wireless communication system is typically composed of cells, with base stations providing cell management and communication services to multiple mobile stations (MS) in the cell. The base station includes a baseband unit (BBU) and a remote radio unit (RRU). The BBU and RRU can be placed in different locations, for example: the RRU is remote and placed in an area with high traffic volume, while the BBU is placed in a central computer room. The BBU and RRU can also be placed in the same computer room. The BBU and RRU can also be different components under the same rack. Optionally, a cell can correspond to a carrier or component carrier.

[0083] In some deployments, the network device referred to in the embodiments of the present application may include a CU, a DU, or both a CU and a DU, or a control plane CU node (centralized unit-control plane (CU-CP)), a user plane CU node (centralized unit-user plane (CU-UP)), and a DU node. For example, the network device may include a gNB-CU-CP, a gNB-CU-UP, and a gNB-DU.

[0084] In some deployments, multiple RAN nodes collaborate to assist terminals in achieving wireless access, with different RAN nodes implementing portions of the base station's functionality. For example, a RAN node can be a CU, DU, CU-CP, CU-UP, or RU. The CU and DU can be separate or included in the same network element, such as the BBU. The RU can be included in a radio frequency device or radio unit, such as an RRU, AAU, or RRH.

[0085] The RAN node may support one or more types of fronthaul interfaces, and different fronthaul interfaces correspond to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more baseband functions, and the RU is configured to implement one or more radio frequency functions. If the fronthaul interface between the DU and the RU is another type of interface, relative to the CPRI, some of the downlink and / or uplink baseband functions, such as precoding, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) for downlink, are moved from the DU to the RU for implementation; and for uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / cyclic prefix removal, are moved from the DU to the RU for implementation. In one possible implementation, the interface may be an enhanced common public radio interface (eCPRI). In the eCPRI architecture, the division between the DU and RU is different, corresponding to different types (category, Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, and F.

[0086] Taking eCPRI Cat A as an example, for downlink transmission, based on layer mapping, the DU is configured to implement layer mapping and one or more functions preceding it (i.e., one or more of coding, rate matching, scrambling, modulation, and layer mapping). Other functions after layer mapping (e.g., resource element (RE) mapping, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) addition) are moved to the RU for implementation. For uplink transmission, based on RE demapping, the DU is configured to implement demapping and one or more functions preceding it (i.e., one or more of decoding, rate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, and RE demapping). Other functions after demapping (e.g., one or more of digital BF or FFT / CP removal) are moved to the RU for implementation. It is understandable that for the functional description of DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol, which will not be described in detail here.

[0087] In one possible design, the processing unit for implementing baseband functions in the BBU is called a baseband high layer (BBH) unit, and the processing unit for implementing baseband functions in the RRU / AAU / RRH is called a baseband low layer (BBL) unit.

[0088] In different systems, CU (or CU-CP and CU-UP), DU or RU may also have different names, but those skilled in the art can understand their meanings. For example, in an open radio access network (ORAN) system, CU may also be referred to as O-CU (open CU), DU may also be referred to as O-DU, CU-CP may also be referred to as O-CU-CP, CU-UP may also be referred to as O-CU-UP, and RU may also be referred to as O-RU. Any of the CU (or CU-CP, CU-UP), DU and RU in this application may be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.

[0089] In the embodiments of the present application, the device for implementing the functions of the network device can be a network device; it can also be a device that can support the network device to implement the functions, such as a chip system, a hardware circuit, a software module, or a hardware circuit and a software module. The device can be installed in the network device or used in conjunction with the network device. In the embodiments of the present application, only the device for implementing the functions of the network device is used as an example to illustrate, and does not constitute a limitation on the solutions of the embodiments of the present application.

[0090] It is understandable that the present application can be applied between network devices and terminal devices.

[0091] Protocol layer structure between network devices and terminal devices:

[0092] The communication between the network device and the terminal device follows a certain protocol layer structure. The protocol layer structure may include a control plane protocol layer structure and a user plane protocol layer structure. For example, the control plane protocol layer structure may include the functions of the radio resource control (RRC) layer, the packet data convergence protocol (PDCP) layer, the radio link control (RLC) layer, the medium access control (MAC) layer, and the physical layer. For example, the user plane protocol layer structure may include the functions of the PDCP layer, the RLC layer, the MAC layer, and the physical layer. In one possible implementation, a service data adaptation protocol (SDAP) layer may also be included above the PDCP layer.

[0093] Optionally, the protocol layer structure between the network device and the terminal device may further include an artificial intelligence (AI) layer for transmitting data related to AI functions.

[0094] Taking data transmission between network devices and terminal devices as an example, data transmission needs to pass through the user plane protocol layers, such as the SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer. The SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer can also be collectively referred to as the access layer. Data transmission is divided into sending or receiving based on the direction of transmission, and each of these layers is further divided into a sending part and a receiving part. Taking downlink data transmission as an example, after the PDCP layer obtains data from the upper layer, it transmits the data to the RLC layer and MAC layer. The MAC layer then generates a transport block, which is then wirelessly transmitted through the physical layer. Data is encapsulated accordingly in each layer. For example, data received by a layer from the layer above it is considered a service data unit (SDU) of that layer. After encapsulation by that layer, it becomes a protocol data unit (PDU) and is then passed to the next layer.

[0095] For example, a terminal device may also include an application layer and a non-access layer. The application layer can be used to provide services to applications installed in the terminal device. For example, downlink data received by the terminal device can be sequentially transmitted from the physical layer to the application layer, which then provides it to the application. For another example, the application layer can obtain data generated by the application and sequentially transmit the data to the physical layer for transmission to other communication devices. The non-access layer can be used to forward user data, such as forwarding uplink data received from the application layer to the SDAP layer, or forwarding downlink data received from the SDAP layer to the application layer.

[0096] It should be understood that the number and type of each device in the communication system shown in Figure 2 are for illustration only, and the present application is not limited to this. In actual applications, the communication system may also include more terminal devices, more network devices, and other network elements, such as core network devices, and / or network elements for implementing artificial intelligence functions.

[0097] It is understandable that all or part of the functions implemented by one or more of the terminal devices, network devices, core network devices, or network elements for implementing artificial intelligence functions can be virtualized, that is, implemented by one or more of the proprietary processors or general-purpose processors and the corresponding software modules. Among them, since the terminal devices and network devices involve interfaces for air interface transmission, the transceiver functions of the interfaces can be implemented by hardware. Core network devices, such as operation administration and maintenance (OAM) network elements, can be virtualized. Optionally, one or more functions of the virtualized terminal devices, network devices, core network devices, or network elements for implementing artificial intelligence functions can be implemented by cloud devices, such as cloud devices in over the top (OTT) systems.

[0098] In an embodiment of the present application, when the central node is a network device and the child node is a terminal device (such as a UE), the network device and UE1 to UE5 can form a distributed AI training system as shown in Figure 3. In this communication system, UE1 to UE5 can send data to the network device, and the network device needs to receive uplink data sent by UE1 to UE5. The uplink data can be the characterization difference or model parameter calculated by the child node, or it can be the feedback amount containing its status information. At the same time, the network device can send configuration information to UE1-UE5. The configuration information can be the model parameter data used by the central node to synchronize each child node, or it can be control data that indicates the training method of the child node. The data between the network equipment and the UE can be carried on physical channels, such as the physical downlink control channel (PDCCH), the physical downlink shared channel (PDSCH), the physical uplink shared channel (PUSCH) or the physical uplink control channel (PUCCH); for example, the physical sidelink control channel (PSCCH) and the physical sidelink shared channel (PSSCH).

[0099] In order to support AI technology in wireless networks, AI nodes may also be introduced into the network.

[0100] Optionally, the AI ​​node can be deployed in one or more of the following locations in the communication system: network equipment, terminal equipment, or core network equipment. Alternatively, the AI ​​node can be deployed separately, for example, in a location other than any of the above devices, such as a host or cloud server in an over-the-top (OTT) system. The AI ​​node can communicate with other devices in the communication system, such as one or more of the following: network equipment, terminal equipment, or core network elements.

[0101] It is understood that this application does not limit the number of AI nodes. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on function, such as different AI nodes are responsible for different functions.

[0102] It can also be understood that AI nodes can be independent devices, or they can be integrated into the same device to implement different functions, or they can be network elements in hardware devices, or they can be software functions running on dedicated hardware, or they can be virtualized functions instantiated on a platform (for example, a cloud platform). This application does not limit the specific form of the above-mentioned AI nodes.

[0103] An AI node can be an AI network element or an AI module.

[0104] One or more AI modules are provided in one or more of these network element nodes, such as core network equipment, access network nodes (RAN nodes), terminals or OAM devices. The access network node can be a separate RAN node, or it can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be provided with one or more AI modules. Optionally, the CU can also be split into CU-CP and CU-UP. One or more AI models are provided in the CU-CP and / or CU-UP.

[0105] The AI ​​module is used to implement the corresponding AI function. The AI ​​modules deployed in different network elements can be the same or different. The model of the AI ​​module can implement different functions according to different parameter configurations. The model of the AI ​​module can be configured based on one or more of the following parameters: structural parameters (such as the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of the neuron, the activation function of the neuron, or at least one of the bias in the activation function), input parameters (such as the type of input parameters and / or the dimension of the input parameters), or output parameters (such as the type of output parameters and / or the dimension of the output parameters). Among them, the bias in the activation function can also be called the bias of the neural network.

[0106] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or on the same node or device.

[0107] The communication system includes a RAN intelligent controller (RIC). For example, the RIC can be the above-mentioned AI module, which is used to implement AI-related functions. The RIC includes a near-real-time RIC (near-real time RIC, near-RT RIC) and a non-real-time RIC (non-real time RIC, Non-RT RIC). Among them, the non-real-time RIC mainly processes non-real-time information, such as data that is not sensitive to delay, and the delay of the data can be in the order of seconds. The real-time RIC mainly processes near-real-time information, such as data that is relatively sensitive to delay, and the delay of the data is in the order of tens of milliseconds.

[0108] Near real-time RIC is used for model training and reasoning. For example, it is used to train an AI model and use the AI ​​model for reasoning. Near real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or reasoning data. Optionally, near real-time RIC can deliver the reasoning results to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and DU, and / or between the DU and RU. For example, the near real-time RIC delivers the reasoning results to the DU, and the DU sends it to the RU.

[0109] Non-real-time RIC is also used for model training and reasoning. For example, it is used to train AI models and use the models for reasoning. Non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU and / or RU) and / or terminals. This information can be used as training data or reasoning data, and the reasoning results can be submitted to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and DU, and / or between the DU and RU. For example, the non-real-time RIC submits the reasoning results to the DU, and the DU sends it to the RU.

[0110] The near-real-time RIC and non-real-time RIC can also be set up as separate network elements. Optionally, the near-real-time RIC and non-real-time RIC can also be part of other devices. For example, the near-real-time RIC is set up in a RAN node (e.g., a CU or DU), while the non-real-time RIC is set up in an OAM, a cloud server, a core network device, or other network devices.

[0111] For example, the configuration of near real-time RIC and non-real-time RIC in the network architecture may be as shown in FIG4A to FIG4D :

[0112] As shown in (a) of FIG4A , in a first possible implementation, the network device includes a near real-time RIC module for performing model learning and / or reasoning.

[0113] As shown in (b) of FIG4A , in a second possible implementation, in a communication system, a non-real-time RIC may be included outside the network device. Optionally, the non-real-time RIC may be located in the OAM or in a core network device.

[0114] As shown in (c) of Figure 4A, in a third possible implementation, the network device includes a near real-time RIC and a non-real-time RIC is also included outside the network device. Optionally, the non-real-time RIC can be located in the OAM or the core network device.

[0115] Compared to (c) in Figure 4A, the CU is separated into CU-CP and CU-UP in Figure 4B. The settings of near real-time RIC and non-real-time RIC are the same as those in (c) in Figure 4A.

[0116] As shown in Figure 4C, optionally, the network device includes one or more AI entities, and the function of the AI ​​entity is similar to the above-mentioned near real-time RIC. Optionally, the OAM includes one or more AI entities, and the function of the AI ​​entity is similar to the above-mentioned non-real-time RIC. Optionally, the core network device includes one or more AI entities, and the function of the AI ​​entity is similar to the above-mentioned non-real-time RIC. When both the OAM and the core network device include AI entities, the models trained by their respective AI entities are different, and / or the models used for reasoning are different. In the present application, the difference in models may include at least one of the following differences: structural parameters of the model (such as the number of layers of the model, and / or weights, etc.), input parameters of the model, or output parameters of the model.

[0117] Relative to Figure 4C, the network device in Figure 4D is separated into CU and DU. Optionally, the CU may include an AI entity, and the function of the AI ​​entity is similar to the above-mentioned near real-time RIC. Optionally, the DU may include an AI entity, and the function of the AI ​​entity is similar to the above-mentioned near real-time RIC. When both the CU and the DU include AI entities, the models trained by their respective AI entities are different, and / or the models used for reasoning are different. Optionally, the CU in Figure 4D can be further split into CU-CP and CU-UP. Optionally, one or more AI models can be deployed in the CU-CP. And / or, one or more AI models can be deployed in the CU-UP. Optionally, in Figure 4C or Figure 4D, the OAM of the network device and the OAM of the core network device can be deployed separately and independently.

[0118] For ease of understanding, the following first introduces the AI ​​technology involved in this application. It should be understood that this introduction does not limit this application.

[0119] (1) AI Model

[0120] AI refers to the intelligence exhibited by machines created by humans. Generally, AI refers to the technology that replicates human intelligence through ordinary computer programs. AI can be defined as machines or computers that mimic humans and possess cognitive functions associated with human thinking, such as learning and problem-solving. AI is able to learn from past experiences, make rational decisions, and respond quickly. The goal of AI is to understand intelligence by building computer programs capable of symbolic reasoning or deduction.

[0121] Machine learning (ML) is a path to artificial intelligence (AI), specifically using machine learning to solve AI problems. Machine learning theory primarily involves the design and analysis of algorithms that enable computers to automatically "learn." Machine learning algorithms automatically analyze data to identify patterns and use these patterns to make predictions about unknown data. Because learning algorithms involve extensive statistical theory, machine learning is particularly closely linked to inferential statistics, also known as statistical learning theory.

[0122] Machine learning can be divided into supervised learning, unsupervised learning, and reinforcement learning.

[0123] Supervised learning uses a machine learning algorithm to learn the mapping relationship between sample values ​​and sample labels based on collected sample values ​​and sample labels. This learned mapping relationship is then expressed using a machine learning model. The process of training a machine learning model is the process of learning this mapping relationship. For example, in signal detection, a noisy received signal is a sample, and the true constellation point corresponding to this signal is the label. Through training, machine learning aims to learn the mapping relationship between samples and labels, essentially enabling the machine learning model to become a signal detector. During training, the model parameters are optimized by calculating the error between the model's predicted values ​​and the true labels. Once the mapping relationship is learned, it can be used to predict the label of each new sample. The mapping relationship learned by supervised learning can include linear and nonlinear mappings. Learning tasks can be categorized into classification and regression tasks based on the type of label.

[0124] Unsupervised learning relies solely on collected sample values, using algorithms to discover inherent patterns within them. One type of unsupervised learning algorithm uses the samples themselves as supervisory signals, meaning the model learns the mapping from one sample to another. This is called self-supervised learning. During training, the model parameters are optimized by calculating the error between the model's predictions and the samples themselves. Self-supervised learning can be used in signal compression and decompression recovery applications. Common algorithms include autoencoders and generative adversarial networks.

[0125] Reinforcement learning, unlike supervised learning, is a type of algorithm that learns problem-solving strategies through interaction with the environment. Unlike supervised and unsupervised learning, reinforcement learning problems lack explicit label data for "correct" actions. Instead, the algorithm must interact with the environment to obtain reward signals from the environment, and then adjust its decision-making actions to maximize the reward signal value. For example, in downlink power control, the reinforcement learning model adjusts the downlink transmit power of each user based on the overall system throughput fed back by the wireless network, hoping to achieve higher system throughput. The goal of reinforcement learning is also to learn the mapping between environmental states and optimal decision-making actions. However, because the labels for "correct actions" are not available in advance, network optimization cannot be achieved by calculating the error between actions and "correct actions." Reinforcement learning training is achieved through iterative interaction with the environment.

[0126] An AI model is an algorithm or computer program that implements AI functions. It is the concrete implementation of AI technology functions. An AI model represents the mapping relationship between the model's input and output. AI models can be neural networks, linear regression models, decision tree models, support vector machines (SVMs), Bayesian networks, Q-learning models, or other machine learning models.

[0127] (2) Deep neural network (DNN)

[0128] Deep neural networks are a specific implementation of AI or machine learning technology. According to the universal approximation theorem, neural networks can theoretically approximate any continuous function, enabling them to learn arbitrary mappings. Traditional communication systems require extensive expert knowledge to design communication modules. However, deep learning communication systems based on DNNs can automatically discover implicit patterns in massive data sets and establish mapping relationships between data, achieving performance superior to traditional modeling methods.

[0129] The idea of ​​DNN is derived from the neuronal structure of the brain. For example, each neuron performs a weighted sum operation on its input values ​​and outputs the result through an activation function. Figure 5 shows a schematic diagram of the neuron structure. Assume that the input of the neuron is x = [x0, x1, ..., xn ], and the weights corresponding to each input are w=[w0,w1,…,w n ], where w i As x i The weight of x i Weighted. The bias of the weighted sum of the input values ​​according to the weight is, for example, b. The activation function can take many forms. Assuming that the activation function of a neuron is: y = f(z) = max(0,z), then the output of the neuron is:

[0130] For another example, if the activation function of a neuron is: y = f(z) = z, then the output of the neuron is:

[0131] Among them, b, w i 、x i It can be a decimal, an integer (such as 0, a positive integer or a negative integer), or a complex number. The activation functions of different neurons in a neural network can be the same or different.

[0132] A neural network generally includes multiple layers, each of which may include one or more neurons. By increasing the depth and / or width of a neural network, its expressive power can be improved, providing more powerful information extraction and abstract modeling capabilities for complex systems. The depth of a neural network can refer to the number of layers it comprises, while the number of neurons in each layer can be referred to as the width of that layer. In one implementation, a neural network includes an input layer and an output layer. The input layer processes the input information received by the neural network through neurons, passes the processing results to the output layer, and the output layer obtains the output of the neural network. In another implementation, the neural network includes an input layer, a hidden layer, and an output layer. The neural network schematic diagram in Figure 6 is provided. The input layer processes the input information received by the neural network through neurons, passes the processing results to an intermediate hidden layer, which performs calculations on the received processing results to obtain a calculation result. The hidden layer then passes the calculation results to the output layer or an adjacent hidden layer, and the output layer ultimately obtains the output of the neural network. A neural network can include one hidden layer or multiple hidden layers connected in sequence, without limitation.

[0133] Depending on how the network is constructed, DNNs can include feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). Figure 6 shows an FNN network, which is characterized by complete connectivity between neurons in adjacent layers. This typically requires a large amount of storage space and results in high computational complexity.

[0134] CNNs are neural networks specifically designed to process data with a grid-like structure. For example, time series data and image data can both be considered grid-like. CNNs don't use all input information at once for computation. Instead, they use a fixed-size window to intercept a portion of the information for convolution operations, significantly reducing the computational complexity of model parameters. Furthermore, depending on the type of information intercepted by the window (e.g., people and objects in an image represent different types of information), different convolution kernels can be used for each window, enabling CNNs to better extract features from the input data.

[0135] RNNs are a type of DNN that utilizes feedback time series information. Their input consists of a new input value at the current moment and their own output value at the previous moment. RNNs are suitable for capturing temporally correlated sequence features and are particularly well-suited for applications such as speech recognition and channel coding.

[0136] The above-mentioned FNN, CNN, and RNN are common neural network structures, which are all constructed based on neurons. As mentioned above, each neuron performs a weighted sum operation on its input values, and the weighted summation result generates an output through a nonlinear function. We call the weights of the weighted summation operation of neurons in the neural network and the nonlinear function the parameters of the neural network. Taking the neuron with max{0,x} as the nonlinear function as an example, The parameters of the neuron to be operated are weights w=[w0,…,w n ], the weighted sum bias is b, and the nonlinear function max{0,x}. The parameters of all neurons in a neural network constitute the parameters of the neural network.

[0137] (3) Training dataset and inference data

[0138] The training dataset is used to train the AI ​​model. The training dataset may include the input of the AI ​​model, or the input and target output of the AI ​​model. The training dataset includes one or more training data. The training data may be a training sample input to the AI ​​model or the target output of the AI ​​model. The target output may also be referred to as a label or a label sample. The training dataset is an important part of machine learning. Model training is essentially learning certain features from the training data so that the output of the AI ​​model is as close as possible to the target output, such as the difference between the output of the AI ​​model and the target output is as small as possible. The composition and selection of the training dataset can, to a certain extent, determine the performance of the trained AI model.

[0139] In addition, during the training process of an AI model (such as a neural network), a loss function can be defined. The loss function describes the gap or difference between the output value of the AI ​​model and the target output value. This application does not limit the specific form of the loss function. The training process of the AI ​​model is a process of adjusting the model parameters of the AI ​​model so that the value of the loss function is less than the threshold, or the value of the loss function meets the target requirements. For example, the AI ​​model is a neural network, and adjusting the model parameters of the neural network includes adjusting at least one of the following parameters: the number of layers, width, weights of neurons, or parameters in the activation function of neurons.

[0140] Inference data can be used as input to a trained AI model for inference. During the model inference process, the inference data is input into the AI ​​model, and the corresponding output is the inference result.

[0141] (4) AI model design

[0142] The design of an AI model primarily includes a data collection phase (e.g., collecting training data and / or inference data), a model training phase, and a model inference phase. It may further include an inference result application phase. See Figure 7, which illustrates an AI application framework. In the aforementioned data collection phase, a data source is used to provide training data sets and inference data. In the model training phase, an AI model is obtained by analyzing or training the training data provided by the data source. The AI ​​model represents the mapping relationship between the model's input and output. Learning an AI model through model training nodes is equivalent to learning the mapping relationship between the model's input and output using training data. In the model inference phase, the AI ​​model trained in the model training phase is used to perform inference based on the inference data provided by the data source to obtain an inference result. This phase can also be understood as: inputting inference data into the AI ​​model, obtaining an output through the AI ​​model, and the output being the inference result. The inference result may indicate: configuration parameters used (executed) by the execution object, and / or the operation performed by the execution object. Inference results are published during the application phase. For example, the inference results can be centrally planned by an execution entity (actor). For example, the execution entity can send the inference results to one or more execution targets (e.g., core network equipment, network equipment, or terminal devices) for execution. Furthermore, the execution entity can provide feedback on model performance to the data source to facilitate subsequent model update and training.

[0143] It is understood that a communication system may include network elements with artificial intelligence capabilities. The aforementioned AI model design-related steps can be performed by one or more network elements with AI capabilities. In one possible design, AI capabilities (such as AI modules or AI entities) can be configured within existing network elements in the communication system to implement AI-related operations, such as AI model training and / or inference. For example, these existing network elements can be network devices (such as gNBs), terminal devices, core network devices, or network management systems. Network management systems can categorize network management tasks into three categories based on the actual needs of the operator's network operations: operations, administration, and maintenance. Network management systems are also referred to as Operational and Administered Management (OAM) network elements, or OAM for short. Operations primarily perform routine network and service analysis, forecasting, planning, and configuration; maintenance primarily involves daily operational activities such as testing and fault management of the network and its services. Network management systems can monitor network operating status, optimize network connectivity and performance, improve network stability, and reduce network maintenance costs. Alternatively, in another possible design, independent network elements can be introduced into the communication system to perform AI-related operations, such as training AI models. The independent network element can be called an AI network element or an AI node, etc., and this application does not limit this name. The AI ​​network element can be directly connected to the network equipment in the communication system, or it can be indirectly connected to the network equipment through a third-party device. Among them, the third-party device can be a core network element such as an authentication management function (AMF) network element, a user plane function (UPF) network element, an OAM, a cloud server or other network elements, without limitation. For example, referring to Figure 8, the communication system includes a network device 810, terminal devices 820 and 830, and an AI network element 840 is also introduced in the communication system.

[0144] In this application, a model can be inferred to obtain a single parameter or multiple parameters. The training process of different models can be deployed on different devices or nodes, or on the same device or node. The inference process of different models can be deployed on different devices or nodes, or on the same device or node.

[0145] Among them, the model parameters may include one or more of the following structural parameters of the model (such as the number of layers and / or weights of the model, etc.), the input parameters of the model (such as input dimension, number of input ports), or the output parameters of the model (such as output dimension, number of output ports). It can be understood that the input dimension may refer to the size of an input data. For example, when the input data is a sequence, the input dimension corresponding to the sequence may indicate the length of the sequence. The number of input ports may refer to the number of input data. Similarly, the output dimension may refer to the size of an output data. For example, when the output data is a sequence, the output dimension corresponding to the sequence may indicate the length of the sequence. The number of output ports may refer to the number of output data.

[0146] (5) Centralized training

[0147] Over the past decade, the number of smart devices, such as mobile phones and wearables, has continued to increase. It is foreseeable that billions of IoT devices will be deployed across communication networks in the near future, enabling the automation and intelligence of social operations. The performance of current intelligent services based on advanced machine learning is likely to benefit from the explosive growth of data on these devices, as well as the computing power available in these UEs. Most machine learning techniques, such as deep neural network-based learning algorithms, require centralized training using all available data. However, centralized training requires the collection of sufficient data, often originating from UEs. This data must be uploaded by the UEs, which incurs significant upload overhead. Furthermore, the concentration of massive amounts of training data in a single node limits training efficiency due to storage space constraints. Furthermore, collecting UE-side data may infringe on user privacy. Storing massive amounts of data for centralized training can raise significant concerns about privacy leaks. On the other hand, if users train machine learning algorithms solely using their own data, training effectiveness is often limited by the limited amount of local data.

[0148] (6) Distributed training

[0149] Distributed training is an effective solution to these challenges. This technology allows the machine learning process to be divided among multiple sub-nodes on the user side, achieving scalability of learning algorithms. It allows a cloud or server, acting as a central node, to collect machine learning models trained by multiple sub-nodes. The central node then integrates these models to improve the overall machine learning training results. Because the training data is always stored on the sub-nodes, distributed learning technology has the potential to achieve the same performance as centralized training while utilizing the user end's data and / or computing power, while protecting user data privacy.

[0150] To achieve this goal, the field of machine learning has integrated computing and communication technologies to design a simple distributed training framework called a parameter server. A parameter server primarily consists of a central node and child nodes. The central node is responsible for storing and / or updating parameters, while the child nodes are responsible for training. Simply put, the basic idea behind parameter server training is to introduce multiple child nodes for simultaneous training. The child nodes synchronize their model parameters by using the central node as a medium for parameter exchange. Each child node first receives training data and / or broadcast model parameters distributed by the central node. It then uses the received data to execute a stochastic gradient algorithm to calculate gradients and upload the gradients to the central node. The central node receives the gradients uploaded by the child nodes, aggregates the gradients to derive new model parameters, and broadcasts these to the child nodes again.

[0151] However, the above training process faces the following two problems:

[0152] 1. Each child node needs to communicate with the central node every time it calculates one or more gradients. Because machine learning algorithms, such as deep learning algorithms, require a large number of gradient updates to converge, the above distributed training framework has extremely high communication overhead.

[0153] 2. The data for child nodes is distributed by the central node, meaning the central node can coordinate the distribution of training data for the child nodes to ensure that the data distribution is consistent across them. In machine learning, independent and identically distributed training data is crucial for ensuring unbiased estimation of the gradients of stochastic gradient algorithms. However, in real-world scenarios, child nodes may be user terminals, and user terminal data is inaccessible to the central node. This means the central node loses control over the distribution of child node data. This can lead to significant disparity between child nodes and compromise overall performance.

[0154] To address the above issues, one approach is to use a federated learning algorithm. Federated learning allows child nodes to perform multi-step gradient updates using local datasets, and then feeds the updated model back to the central node to reduce the communication overhead caused by the need for child nodes to interact with the central node for each gradient update. Specifically, in the federated learning framework, K child nodes and the central server node collaborate to train a global shared model. The server aggregates the models sent by the child nodes, and one or more child node models are trained using local private data for multiple rounds. One round refers to a traversal of all available local data. Assuming that the local dataset contains N training samples, and b samples are randomly selected each time during training to execute the stochastic gradient descent algorithm, then one round of calculation is subgradient, where = represents rounding down. The central node then performs model aggregation. After aggregation is complete, the central node shares the aggregated model with its child nodes, and the above steps are repeated until the model converges. Since all local models are trained based on locally stored data, the data privacy of the child nodes is protected. The entire process of this framework includes local training, communication between child nodes and the central node, and model aggregation. During the local training phase, child nodes train their own models independently and in parallel. The models used by child nodes for local training can be any type of machine learning model. Currently, due to the proven power of deep neural networks in various fields, DNNs are typically used as local models for child nodes to achieve optimal performance. During the communication phase, the data transmitted between child nodes and the central node is only the model trained by the child node or the model aggregated by the central node. To improve efficiency, federated learning typically selects a subset of child nodes to upload models. The central node aggregates the received models within a set time window. For aggregation, model averaging is commonly used in federated learning, where the central node performs a weighted average of the received child node models. The above training, communication, and aggregation steps are repeated multiple times until the central node's model converges.

[0155] However, although federated learning can protect user privacy and reduce the communication overhead of traditional parameter server technology, it has the following disadvantages because the central node cannot access user data:

[0156] 1. When the data distribution of child nodes is not independent and identically distributed (IID), performance cannot be guaranteed. In practice, the amount of data between child nodes may be unbalanced. Due to different user preferences, the data on user devices may not be IID. The distribution of a single child node dataset may not represent the overall data distribution. In this case, the gradients obtained by the stochastic gradient descent algorithm on the child node may be biased, ultimately leading to poor distributed training performance.

[0157] 2. Performance fluctuations are significant when encountering unstable feedback links. When user subnodes feed back gradients via unstable wireless transmission, the number of subnode gradients successfully received by the server is somewhat random, which will cause fluctuations in the model parameter updates of the central server.

[0158] The above problems will cause the sub-nodes to perform multiple gradient updates locally, resulting in large differences between the model and the central node, which will eventually lead to a decline in distributed training performance.

[0159] Another approach is the Gossip learning method. This learning framework manages a set of child computing nodes, one or more of which owns a machine learning model and performs two iterations: local gradient update and inter-node aggregation update. Specifically, one or more child nodes implement multiple gradient updates to the local machine learning model in the gradient update step. The child node then shares its model parameters with another randomly selected child node in the aggregation update step. These steps are repeated until all child nodes converge on a consensus model. Like federated learning, this technology allows child nodes to perform multiple steps of gradient calculations before communicating, so frequent communication is not required. At the same time, distributed learning without a fixed central node can be achieved, so the Gossip learning method has stronger scalability.

[0160] Similar to the federated learning algorithm, the Gossip learning method still faces the problem of non-independent and identically distributed data. In addition, in the process of model parameter interaction between nodes, the node receiving the parameters can also be regarded as the central node, and the node sending the parameters can be regarded as the sub-computing node. Therefore, the Gossip learning framework will also encounter feedback instability. These problems will lead to large differences between the models of the computing nodes, and ultimately make it difficult to converge to a consensus model with good performance.

[0161] As can be seen above, to ensure user privacy, distributed training requires that the central node not access the data of child nodes. Existing federated learning and gossip learning technologies alleviate the frequent communication issues of distributed training systems by allowing child nodes to perform multiple gradient updates locally before communicating. However, these technologies face challenges with non-IID (independent and identically distributed) data distributions and unstable wireless feedback. This leads to significant differences between child nodes and between child nodes and the central node, ultimately resulting in the distributed system's performance failing to meet service requirements. Therefore, there is a need to improve the performance of distributed training.

[0162] To address the above-mentioned issues, the present application provides a distributed training method, in which the sub-nodes participating in the training of the central node model feed back the representation differences between the models to the central node, and the central node gives training instructions for the model based on the representation differences, so that a high-performance machine learning model can be obtained in the process of a limited number of communications between the sub-nodes and the central node, thereby improving the performance of distributed training.

[0163] The distributed training method provided by the embodiment of the present application is described in detail below. It will be understood that the present application uses the central node and the sub-node as an example to illustrate the execution subject of the interactive diagram, but the present application does not limit the execution subject of the interactive diagram. For example, the central node in the method provided by the present application may also be a chip, chip system, circuit or processor applied to the central node, or a logical node, logic module or software that can realize all or part of the central node; the sub-node in the method provided by the present application may also be a chip, chip system, circuit or processor applied to the sub-node, or a logical node, logic module or software that can realize all or part of the sub-node functions.

[0164] As shown in Figure 9, a flow chart of a distributed training method provided in an embodiment of the present application is provided. Exemplarily, the method may include the following steps:

[0165] S901. A first child node sends first information to a central node. Correspondingly, the central node receives the first information.

[0166] The distributed training system shown in Figure 1 may include a central node and multiple subnodes. In this embodiment, the first subnode is any one of the multiple subnodes participating in the model training of the central node. The interactive operations between the central node and the first subnode, as well as the operations performed within the first subnode itself, described in this embodiment, can also be applied to other subnodes.

[0167] In order to reduce the frequency of communication between child nodes and central nodes, child nodes have the ability to perform multi-step gradient calculations and / or model parameter updates locally. When the data distribution between child nodes is independent and identically distributed, the gradient obtained by the child node using the stochastic gradient descent (SGD) method is unbiased. At this time, the distributed training system can show good performance. However, in practice, since the training data of the child nodes may come from users in different geographical locations, different time periods, and different preferences, the assumption of independent and identical distribution is often not applicable to the actual system. This results in the model differences between child nodes and the model differences between child nodes and the central node increasing as the training progresses when the child nodes use local data to update the model parameters, resulting in the overall performance of distributed training being degraded.

[0168] During distributed training, the central node sends the initial model information to each child node participating in the model training. Each child node trains the model based on local data and performs local model updates.

[0169] In this embodiment, each time the local model of the first child node is updated, the first child node can calculate the first representation difference between the model of the central node and the model of the first child node, or calculate the first representation difference between the models of all child nodes participating in the model training of the central node, and send first information to the central node. The first information is used to indicate the first representation difference between the model of the central node and the model of the first child node, or the first representation difference between the models of all child nodes. Exemplarily, the first child node can send the first information to the central node via wireless transmission.

[0170] For example, if the first characterization difference is the Euclidean norm (abbreviated as "L2 norm") distance (abbreviated as "L2 distance") or the cosine distance, then the first information includes the L2 distance between the model of the central node and the model of the first child node, or includes the cosine distance between the model of the central node and the model of the first child node, or includes the L2 distance between the models of each child node, or includes the cosine distance between the models of each child node. At the beginning of model training, the first child node receives the initialization model information of the central node, and after the global model is updated, the first child node receives the updated model information of the central node. The first child node can calculate the L2 distance or cosine distance between the model of the central node and the model of the first child node based on the received model information of the central node and its own model information. In another possible implementation, the first child node can receive the model information of each child node, and calculate the L2 distance or cosine distance between the model of the first child node and the models of other child nodes based on the received model information of other child nodes and its own model information. Among them, the other child nodes can be one or more child nodes.

[0171] Among them, the L2 distance is defined as:

[0172] L2(w o ,w k )=||w k -w o ||2, where w o is the parameter of the central node model, w k is the model parameter of the first child node, ||·||2 is the L2 distance operator, for vector x=(x1,x2,…,x n ),

[0173] The cosine distance is defined as:

[0174] where w o is the parameter of the central node model, w k is the model parameter of the first child node, represents w k The transpose of .

[0175] It is understandable that the first subnode may also send the received initialization model information of the central node to a third-party device, and the third-party device may perform model training based on the local data of the first subnode, and perform model updates, and calculate the first representation difference between the model of the central node and the model of the first subnode, or calculate the first representation difference between the model of the other subnodes participating in the model training of the central node and the model of the first subnode, and send the first representation difference to the first subnode, and the first subnode may send the first representation difference to the central node. The other subnodes may be one or more subnodes. In one possible implementation, the other subnodes include all subnodes participating in the model training of the central node.

[0176] S902: The central node sends second information to the first child node according to the first representation difference. Correspondingly, the first child node receives the second information.

[0177] After receiving the first representation differences sent by all child nodes participating in the model training of the central node, the central node performs difference value analysis, generates training instruction information for each child node, and sends second information to each child node. The second information includes the training instruction information of the model of each child node.

[0178] If the first representation difference between the central node's model and the first child node's model is too large, or the first representation difference between the first child node and the models of other child nodes is too large, indicating that the model trained by the first child node deviates far from the central node's model, or the model trained by the first child node deviates far from the models trained by other child nodes, the central node sends a training instruction message, which is used to instruct the first child node to stop training the first child node's model. This can timely prevent the model difference between the first child node and the central node (or other child nodes) from increasing continuously as training progresses, resulting in the overall performance degradation of distributed training.

[0179] Furthermore, the central node can also assign corresponding loss function parameters to child nodes with large difference values ​​based on the difference value analysis results, so that the child nodes can update their models according to the parameters, so that the subsequent updates of the child node models have less impact on the training results.

[0180] If the first representation difference between the model of the central node and the model of the first child node is small, or the first representation difference between the model of the first child node and other child nodes is small, it indicates that the model trained by the first child node does not deviate far from the model of the central node, or the model trained by the first child node does not deviate far from the model trained by other child nodes. The central node sends training instruction information, which is used to instruct the first child node to continue training the model of the first child node.

[0181] According to a distributed training method provided in an embodiment of the present application, the sub-nodes participating in the training of the central node model feed back the representation differences between the models to the central node, and the central node gives training instructions for the model based on the representation differences, so that a high-performance machine learning model can be obtained in the process of a limited number of communications between the sub-nodes and the central node, thereby improving the performance of distributed training.

[0182] As shown in Figure 10, it is a flowchart of another distributed training method provided in an embodiment of the present application. Exemplarily, the method may include the following steps:

[0183] S1001. The central node broadcasts third information. Correspondingly, the first child node receives the third information.

[0184] The distributed training system shown in Figure 1 may include a central node and multiple subnodes. In this embodiment, the first subnode is any one of the multiple subnodes participating in the model training of the central node. The interactive operations between the central node and the first subnode, as well as the operations performed within the first subnode itself, described in this embodiment, can also be applied to other subnodes.

[0185] The central node usually sends the same third information to all child nodes participating in the model training of the central node in a broadcast manner.

[0186] The third information includes at least one of the following: model configuration information and a public dataset. During initial training, the central node typically broadcasts the same model configuration information and / or public dataset to one or more child nodes participating in the central node's model training. The central node itself also stores the same model configuration information. Optionally, the one or more child nodes may include all child nodes participating in the central node's model training.

[0187] The model configuration information is used to indicate at least one of the following: the model type of the multiple sub-nodes, the structural information of the models of the multiple sub-nodes, the model parameters of the models of the multiple sub-nodes, or the training configuration information of the models of the multiple sub-nodes. Optionally, the models of the multiple sub-nodes are the same, that is, the model of the central node.

[0188] For example, the model types of the child nodes include DNN, CNN, random forest model, etc.

[0189] For example, the structural information of the subnode model includes the number of hidden layers of the DNN, the neuron data of each layer or part of the layer, and the activation function.

[0190] For example, the type and structure information of the models of multiple child nodes sent down can be reflected in the form of configuration text, or it can be a code script that compiles the corresponding machine learning model.

[0191] The model parameters of the sub-node model are generated by the central node through a certain strategy, including but not limited to random generation, pre-training generation, or acquisition from other third-party entities.

[0192] Among them, the training configuration information of the subnode model includes the optimizer used by the subnode to perform gradient updates (such as SGD, root mean square propagation (RMSprop), adaptive moment estimation (Adam), etc.), the loss function of the task (such as the cross entropy loss function for classification tasks and the mean square error loss function for regression tasks), the regularization penalty term (such as the L2 penalty term), the initial learning rate, the gradient update batch size and the maximum number of training rounds for the subnode, etc. It is agreed that the loss function used for local training of the subnode is the loss function of the task plus the representation difference penalty term. For example, if the task loss is the classification cross entropy CrossEntropy(), then for a batch of training samples The loss function used for local training of child node k is: CrossEntropy(D)+A*e k (D), where x i is the input of the model; i is the label of the model; N is the number of training samples; i is the i-th training sample; A is the representation difference coefficient; e k To represent the differences.

[0193] S1002. The first sub-node calculates a first representation difference between the model of the central node and the model of the first sub-node based on the local dataset or the public dataset of the first sub-node.

[0194] For example, each time the local model of the first child node is updated, each child node of the first child node may be triggered to calculate the first representation difference based on its own local dataset or public dataset. The local dataset may be pre-stored in each child node at the factory. The local datasets of each child node are the same.

[0195] As mentioned above, the public data set is sent by the central node to each child node. Each child node receives the same public data set.

[0196] By having each child node calculate the first representation difference based on the same local data set or public data set, the representation difference between the model of each child node and the model of the central node, or the representation difference between the models of each child node, can be more accurately compared.

[0197] The first representation difference refers to the difference between the outputs or intermediate quantities calculated by the machine learning models of each child node and / or the central node using the same input. That is, the first representation difference is the difference between the outputs or intermediate quantities of the model of the central node and the model of the first child node, and / or the first representation difference is the difference between the outputs or intermediate quantities of the models of multiple child nodes, wherein the outputs or intermediate quantities are obtained based on the same input. The same input is the local dataset or public dataset of the above-mentioned child node.

[0198] As shown in Figure 11, it is a schematic diagram of a calculation of a characterization difference of an example of an embodiment of the present application. The machine learning model is a neural network, and its structure includes a feature extractor, a classifier or a regressor. For this model, its characterization difference calculation follows the following steps: put the central node model into verification mode (turn off the randomness of the modules with randomness in the model, such as turning off the dropout function in the DNN), randomly select a batch of training samples from the local data set or the public data set, send the samples to the central node model and obtain the feature z extracted by the model o ; Put the subnode local model into training mode (turn on the randomness of the random modules in the model, such as turning on the Dropout function in DNN), send the sample to the subnode model and obtain the feature z extracted by the model k ; Calculate feature z o and feature z k Some measure between dis(z o ,z k ) (such as L2 distance or cosine distance), this metric is used to represent the difference e k It is understandable that the example in FIG11 uses the difference between the intermediate quantities of the model of the central node and the model of the first child node as the representation difference, and the representation difference may also be the difference between the outputs of the model of the central node and the model of the first child node.

[0199] It can be understood that Figure 11 uses the L2 distance or cosine distance between the output of the last layer of the feature extractor of the central node and the output of the last layer of the feature extractor of the child node as the characterization difference. In addition, the L2 distance or cosine distance between the output of any layer of the feature extractor of the central node and the output of any layer of the feature extractor of the child node can also be calculated as the characterization difference. Alternatively, the L2 distance or cosine distance between the output of any layer of the classifier or regressor of the central node and the output of any layer of the classifier or regressor of the child node can also be calculated as the characterization difference. This application does not impose any restrictions on this.

[0200] Based on the process shown in FIG11 , the representation difference between the model of the central node and the model of the child nodes can be calculated.

[0201] In another scenario, the model of the central node may be replaced with the model of another child node (reference node model), and the representation difference between each child node model and the reference node model may be calculated.

[0202] S1003: The first child node sends first information to the central node. Correspondingly, the central node receives the first information.

[0203] After calculating the first representation difference, the first child node sends first information to the central node. Exemplarily, the first child node may send the first information to the central node via wireless transmission. The first information indicates the first representation difference between the central node's model and the first child node's model, or indicates the first representation difference between the models of the child nodes.

[0204] For example, if the first characterization difference is L2 distance or cosine distance, the first information includes the L2 distance between the model of the central node and the model of the first child node, or includes the cosine distance between the model of the central node and the model of the first child node, or includes the L2 distance between the models of each child node, or includes the cosine distance between the models of each child node.

[0205] S1004. The central node aggregates the first representation differences of each sub-node and generates first training instruction information R1 for each sub-node.

[0206] After receiving the first representation differences sent by all the child nodes participating in the model training of the central node, the central node performs difference value analysis to generate first training instruction information R1 for each child node.

[0207] The central node can set a first condition and a second condition to determine whether the first representation difference of each child node satisfies the first condition or the second condition. If the first representation difference of a child node satisfies the first condition, the first training instruction information is used to instruct the first child node to continue training the first child node's model; if the first representation difference of a child node satisfies the second condition, the first training instruction information is used to instruct the first child node to stop training the first child node's model. Exemplarily, the first representation difference satisfying the first condition includes: the first representation difference being less than or equal to a first threshold. The central node can compare the first representation difference of each child node with the first threshold. If the first representation difference of any child node is less than or equal to the first threshold, the first training instruction information is used to instruct the child node to continue training the child node's model. The central node can also compare the first representation difference of each child node with a second threshold. If the first representation difference of any child node is greater than or equal to the second threshold, the first training instruction information is used to instruct the child node to stop training the child node's model. By selecting a unified condition to determine whether the first representation difference of each child node satisfies the above conditions, the central node can improve the efficiency and accuracy of difference value analysis. Optionally, the first condition, such as the first threshold, is predetermined by the central node or predefined by the protocol; and the second condition, such as the second threshold, is predetermined by the central node or predefined by the protocol. Exemplarily, the second threshold may be the same as or different from the first threshold.

[0208] Furthermore, the central node can also assign corresponding training function parameters, such as a characterization difference coefficient A, to child nodes with large difference values ​​based on the difference value analysis results. It can be understood that the larger the first characterization difference, the larger the characterization difference coefficient. Because the larger the first characterization difference, it indicates that the model trained by the child node deviates further from the model of the central node, or the model trained by the child node deviates further from the model trained by other child nodes among all child nodes participating in the model training of the central node. Therefore, the central node assigns a characterization difference coefficient A so that the child node can update its loss function according to the characterization difference coefficient A, so that the update of the model of the subsequent child node has less impact on the training results.

[0209] S1005. The central node sends the second information to the first child node. Correspondingly, the first child node receives the second information.

[0210] The second information may be implemented in the following ways:

[0211] In one implementation, the second information includes first training instruction information R1 of the model of the first sub-node.

[0212] According to the above difference value analysis results, if the first representation difference of the first subnode meets the first condition, for example, the first representation difference of the first subnode is less than or equal to the above first threshold, then the first training indication information is used to instruct the first subnode to continue the training of the model of the first subnode.

[0213] Furthermore, the first training indication information may also be used to indicate a representation difference coefficient in a loss function of the model of the first sub-node.

[0214] According to the above-mentioned difference value analysis results, if the first representation difference of the first subnode meets the second condition, for example, the first representation difference of the first subnode is greater than or equal to the above-mentioned second threshold, then the first training indication information is used to instruct the first subnode to stop training the model of the first subnode.

[0215] For example, the first training instruction information may include one or more bits, wherein one bit is used to instruct the first child node to continue or stop training the first child node's model. For example, if the value of this bit is "1," it is used to instruct the first child node to continue training the first child node's model; if the value of this bit is "0," it is used to instruct the first child node to stop training the first child node's model. The reverse is also possible.

[0216] If the value of this bit is "1", it is used to instruct the first sub-node to continue the training of the model of the first sub-node, and the remaining bits of the first training indication information can be used to indicate the representation difference coefficient in the loss function of the model of the first sub-node.

[0217] In another implementation, the second information includes first training instruction information R1 of the model of the first sub-node, and the first training instruction information is used to instruct the first sub-node to continue or stop training of the model of the first sub-node.

[0218] And if the first representation difference of the first subnode satisfies the first condition, the second information also includes a representation difference coefficient in the loss function of the model of the first subnode.

[0219] S1006. The first child node updates its model according to the second information.

[0220] After the first subnode receives the second information, if the first training indication information is used to instruct the first subnode to continue the training of the model of the first subnode, and is also used to indicate the representation difference coefficient in the loss function of the model of the first subnode; or, the second information also includes the representation difference coefficient in the loss function of the model of the first subnode, then the first subnode uses local data to update one or more step model parameters according to the second information, including using the representation difference coefficient to update the loss function of the model.

[0221] This step assumes that the first training instruction information is used to instruct the first child node to continue training the model of the first child node. If the first training instruction information is used to instruct the first child node to stop training the model of the first child node, it will be described in detail later.

[0222] S1007. The first sub-node calculates a second representation difference between the model of the central node and the model of the first sub-node based on the local dataset or the public dataset of the first sub-node.

[0223] After the first child node updates its local model based on the second information, it can trigger another calculation based on the first child node's local dataset or public dataset to calculate a second representation difference between the central node's model and the first child node's model. This calculation process can be referred to the description of step S1002 and will not be repeated here. The second representation difference calculated by the first child node can be the same as or different from the first representation difference.

[0224] S1008. The first sub-node sends fourth information to the central node. Correspondingly, the central node receives the fourth information.

[0225] After calculating the second representation difference, the first child node sends fourth information to the central node, where the fourth information is used to indicate the second representation difference between the central node's model and the first child node's model. The specific implementation of this step can be referred to the description of step S1003 and will not be repeated here.

[0226] S1009. The central node aggregates the second representation differences of each sub-node and generates second training instruction information R2 for each sub-node.

[0227] The specific implementation of this step can be referred to the description of step S1004 and will not be repeated here. It is understandable that the second training instruction information R2 of each sub-node generated by the central node can be the same as or different from the first training instruction information R1 of the corresponding sub-node.

[0228] S1010: The central node sends fifth information to the first child node. Correspondingly, the first child node receives the fifth information.

[0229] After generating the second training instruction information R2 for each child node, the central node sends fifth information to the first child node, wherein the fifth information includes the second training instruction information R2 of the model of the first child node.

[0230] Regarding the fifth information, there are several possible implementations:

[0231] In one implementation, the fifth information includes second training instruction information R2 of the model of the first sub-node.

[0232] According to the above difference value analysis result, if the second representation difference of the first sub-node is greater than or equal to the above threshold, the second training instruction information is used to instruct the first sub-node to stop the training of the model of the first sub-node.

[0233] Another implementation is that the fifth information includes second training instruction information R2 of the model of the first sub-node, and the second training instruction information is used to instruct the first sub-node to stop training of the model of the first sub-node.

[0234] And if the second representation difference of the first sub-node is less than or equal to the above threshold, the fifth information also includes the representation difference coefficient in the loss function of the model of the first sub-node.

[0235] S1011. The first child node stops model training.

[0236] If the first representation difference of the first child node is greater than or equal to the threshold, the first training instruction information is used to instruct the first child node to stop training the model of the first child node. After receiving the fifth information, the first child node stops model training.

[0237] Alternatively, when the first child node reaches the maximum number of training rounds indicated by the above training configuration information, the model training is stopped.

[0238] S1012: The first sub-node sends sixth information to the central node. Correspondingly, the central node receives the sixth information.

[0239] After the first child node stops model training, it can send the model information obtained through training to the central node. For example, the first child node can send sixth information to the central node. The sixth information includes the model information obtained through training by the first child node.

[0240] S1013. The central node updates the model of the central node according to the sixth information.

[0241] After receiving the sixth information, the central node may update the model of the central node according to the sixth information.

[0242] After a child node stops model training, the central node receives the model information obtained by the child node training, and can update the model of the central node based on the model information obtained by the child node training; or it can wait until the model information obtained by the training of all child nodes participating in the training of the central node model is received, and then uniformly update the model of the central node.

[0243] The central node updates the model in a manner including but not limited to weighted averaging of model parameters.

[0244] It is understood that the above steps S1002 to S1006 and steps S1007 to S1013 can be implemented independently or in combination. The above steps S1002 to S1006 can also be performed once or multiple times until the maximum number of training rounds for the first child node is reached, or because the calculated representation difference is greater than or equal to a threshold, the central node is instructed to stop training the model of the first child node.

[0245] S1014. The central node sends a new central node model. Correspondingly, the first node receives the new central node model sent by the central node.

[0246] After all child nodes participating in the training of the central node model have reached the maximum number of training rounds, or the central node has instructed to stop training the child node model because the calculated representation difference is greater than or equal to a threshold, the central node receives the model information obtained from the training of all child nodes participating in the training of the central node model, uses the model information of these child nodes to update the central node model, and broadcasts the new model configuration information of the central node to all child nodes participating in the training again, starting a new round of training. The new round of training process can repeat the above steps S1001 to S1014.

[0247] Optionally, the central node is a third-party device that performs the aforementioned central node-related actions. For example, the above steps S1001, S1003-S1005, S1008-S1010, and S1012-S1014 are all performed by a third-party device.

[0248] Optionally, the first sub-node is a third-party device that performs the aforementioned actions related to the first sub-node. For example, the above steps S1001-S1003, S1005-S1008, S1010-S1012, and S1014 are all performed by a third-party device.

[0249] Optionally, the central node is a network device. In this case, the network device can complete the model training. For example, the above steps S1001, S1003-S1005, S1008-S1010, and S1012-S1014 are all performed by the network device.

[0250] Optionally, the first child node is a terminal device. In this case, the terminal device can complete the training of the model. For example, the above steps S1001-S1003, S1005-S1008, S1010-S1012, and S1014 are all performed by the terminal device.

[0251] Optionally, the central node includes a network device and a third-party device. In one example, one or more of steps S1004, S1009, and S1013 can be performed by a third-party device, such as an OTT or cloud server, and one or more of steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 can be performed by the network device. In addition, the network device and the third-party device can also communicate with each other to transmit the content transmitted by one or more of steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014.

[0252] Optionally, the first sub-node includes a terminal device and a third-party device. In one example, one or more of steps S1002, S1006, S1007, and S1011 may be performed by a third-party device, such as an OTT or cloud server, and one or more of steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 may be performed by the terminal device. In addition, the terminal device and the third-party device may also communicate with each other to transmit the content transmitted by one or more of steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014.

[0253] According to a distributed training method provided in an embodiment of the present application, the sub-nodes participating in the training of the central node model feed back the representation differences between the models to the central node, and the central node gives training instructions for the model based on the representation differences, so that a high-performance machine learning model can be obtained in the process of a limited number of communications between the sub-nodes and the central node, thereby improving the performance of distributed training.

[0254] By selecting appropriate physical quantities to characterize the differences in the models of distributed nodes (the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, or the difference between the output or intermediate quantity of the models of each child node, such as L2 distance or cosine distance), and instructing the training of the child node model based on the characterization difference, the model training of the distributed training system can be better coordinated and the overall performance can be improved.

[0255] However, due to the lack of monitoring of the status of sub-nodes in the existing technology, there are large differences between sub-nodes and between sub-nodes and central nodes, which leads to performance degradation.

[0256] The embodiments of the present application can effectively reduce the dispersion of sub-node update directions by constraining model characterization differences, reduce the impact of non-independent and identically distributed data on the performance of the distributed training system and the deviation caused by packet loss, and improve the final performance of the central node model.

[0257] In the embodiments of the present application, when the representation difference is large, the central node instructs the child node to stop updating the model, which can reduce unnecessary computational overhead for the child node. Alternatively, the central node can provide feedback to the child node on a large representation difference coefficient, thereby limiting the representation difference during the child node training process. This can improve the performance of the child node during joint training with the central node. If the child node loses packets when feeding back model parameters, the child node and the central node's models will still maintain a high degree of synchronization.

[0258] As shown in FIG12, a block diagram of a distributed training system according to an embodiment of the present application is shown. The distributed training system includes a central node and K child nodes, where K is a positive integer greater than or equal to 1. The model of the central node is f o (.,w o ), the model of child node k is model f k (.,w k ). Where k∈{1,2,…,K}. The central node broadcasts the model configuration information and the public dataset. After receiving the information broadcast by the central node, each child node initializes its local model. Each child node calculates the representation difference e between the central node model and the child node model based on its local dataset or public dataset. k Each child node feeds back the representation difference e to the central node k . The central node aggregates the representation differences of each child node and generates training instruction information for each child node. The central node sends training instruction information to each child node respectively. Each child node performs local training based on the received training instruction information. During the training process, each child node uses local data to calculate the loss, and adds the representation difference coefficient as a penalty term to the loss function, and then calculates the gradient of the model parameters and performs local training. If a child node meets the training stop condition, the training is stopped and the trained model information is fed back to the central node. The central node receives the model information from the child node and updates its own parameters w based on the received model. o , and feed back the new model to one or more child nodes, and re-execute the aforementioned representation difference calculation and reporting process. For example, if the task loss is the classification cross entropy CrossEntropy(), then for a batch of training samples The loss function used for local training of child node k is: CrossEntropy(D)+A*e k (D), where x i is the input of the model; i is the label of the model; N is the number of training samples; i is the i-th training sample; A is the representation difference coefficient; e kThe larger the first representation difference, the larger the representation difference coefficient. Updating the loss function according to the representation difference coefficient A can reduce the impact of subsequent sub-node model updates on the training results.

[0259] In the above process, the child nodes can feed back model information to the central node after completing the multi-step gradient update, which can alleviate the communication overhead between the child nodes and the central node. In addition, each child node uses a local data set or a public data set to calculate the representation difference. By making each child node calculate the first representation difference based on the same local data set or public data set, the representation difference between the model of each child node and the model of the central node, or the representation difference between the models of each child node can be more accurately compared. By aggregating the representation differences of each child node and sending targeted training instruction information to each child node, the central node can ensure that the representation differences between the models of the child nodes and between the models of the child nodes and the model of the central node are not too large, thereby improving the performance of distributed training.

[0260] The above embodiment describes a scenario where subnodes feed back representation differences to a central node, which then aggregates the representation differences of each subnode and issues training instructions. The following embodiment describes a scenario where subnodes calculate the representation differences themselves and decide whether to continue or stop model training:

[0261] As shown in Figure 13, a flow chart of another distributed training method provided in an embodiment of the present application is provided. Exemplarily, the method may include the following steps:

[0262] S1301: The central node broadcasts first information, and the first child node receives the first information accordingly.

[0263] The first information includes at least one of the following: model configuration information, a public data set, a threshold update rule, and a characterization difference coefficient update rule.

[0264] For the meanings of the model configuration information and the public data set, reference may be made to the relevant description in step S1001 in the embodiment shown in FIG10 .

[0265] Among them, the update rule of the threshold can be pre-established by the central node. For example, the update rule is established as follows: the threshold decreases as the number of communications between the first child node and the central node increases, and the threshold of the first child node remains unchanged before the first child node communicates with the central node. Among them, the central node sends a new model to the first child node once, which is one communication. It can be understood that as the child nodes are trained and the central node aggregates the models of the child nodes, the models of the central node and the child nodes gradually converge. Therefore, as the number of communications between the first child node and the central node increases, the representation difference between the model of the central node and the model of the first child node will gradually decrease. Therefore, a threshold update rule can be established in which the threshold decreases as the number of communications between the first child node and the central node increases.

[0266] Optionally, the first information or the threshold update rule may further include an initial threshold.

[0267] Among them, the updating rule of the characterization difference coefficient is used to instruct the child node to obtain an updated characterization difference coefficient based on the initial characterization difference coefficient and the calculated characterization difference after the characterization difference is calculated. The updating rule of the characterization difference coefficient is pre-established by the central node. For example, the updating rule of the characterization difference coefficient is: the linear or nonlinear scaling of the characterization difference coefficient of the previous update is used as the update amount of the characterization difference coefficient compared with the previous characterization difference coefficient. It can be understood that the larger the characterization difference, the larger the characterization difference coefficient. Because the larger the characterization difference, it indicates that the model trained by the child node deviates far from the model of the central node, or the model trained by the child node deviates far from the model trained by other child nodes among all the child nodes participating in the model training of the central node. Therefore, the child node obtains an updated characterization difference coefficient, so that the child node can update its loss function according to the characterization difference coefficient, so that the update of the model of the subsequent child node has less impact on the training results.

[0268] Optionally, the first information or the updating rule of the characterization difference coefficient may further include an initial characterization difference coefficient.

[0269] S1302. The first sub-node calculates a first representation difference between the model of the central node and the model of the first sub-node based on the local dataset or the public dataset of the first sub-node.

[0270] For the specific implementation of this step, please refer to the relevant description of step S1002 in the embodiment shown in FIG10 .

[0271] S1303. When the first sub-node determines that the first representation difference is less than the first threshold, continue training the model of the first sub-node.

[0272] The first threshold may be pre-stored in the first child node, or sent by the central node to the first child node in step S1301, for example, the initial threshold carried in the first information or the threshold update rule.

[0273] The first child node compares the calculated first representation difference with the first threshold. If the first representation difference is greater than the first threshold, the first child node stops the model training; if the first representation difference is less than the first threshold, the first child node continues the model training; if the first representation difference is equal to the first threshold, the first child node continues or stops the model training.

[0274] This step assumes that the first representation difference is less than or equal to the threshold, and the first child node continues model training. As will be described in detail later, if the first representation difference is greater than the threshold, the first child node stops model training.

[0275] S1304. The first child node calculates a gradient using the local data set and a loss function with a coefficient representing the difference, and updates the model parameters.

[0276] The first child node uses the local dataset to update the model parameters for one or more steps. The update uses a loss function with a coefficient representing the variance to calculate the gradient. The meaning of the coefficient representing the variance can be found in the description above.

[0277] The characterization difference coefficient may be obtained according to the characterization difference coefficient update rule issued by the central node in step S1301.

[0278] S1305. The first child node updates the threshold according to the threshold updating rule, and updates the characterization difference coefficient according to the characterization difference coefficient updating rule.

[0279] The first child node updates the first threshold according to the threshold update rule received in step S1301 to obtain the second threshold, and updates the representation difference coefficient according to the representation difference coefficient update rule received in step S1301.

[0280] This step is optional, and the threshold update rule / characterization difference coefficient update rule may not stipulate that the threshold and / or characterization difference coefficient must be updated after each local model update of the first child node.

[0281] S1306. The first sub-node calculates a second representation difference between the model of the central node and the model of the first sub-node based on the local dataset or the public dataset of the first sub-node.

[0282] After the first child node updates its local model based on the second information, it can trigger another calculation based on the first child node's local dataset or public dataset to calculate a second representation difference between the central node's model and the first child node's model. This calculation process can be referred to the description of step S1302 and will not be repeated here. The second representation difference calculated by the first child node can be the same as or different from the first representation difference.

[0283] S1307. If the first sub-node determines that the second representation difference is greater than the second threshold, or reaches the maximum number of training rounds, the training of the model of the first sub-node is stopped.

[0284] Optionally, when the first sub-node determines that the second representation difference is equal to a second threshold, the training of the model of the first sub-node is stopped.

[0285] S1308. The first child node sends the second information to the central node. Correspondingly, the central node receives the second information.

[0286] After the first child node stops model training, it can send the model information obtained through training to the central node. For example, the first child node can send second information to the central node. The second information includes the model information obtained through training by the first child node.

[0287] S1309. The central node updates the model of the central node according to the second information.

[0288] After receiving the second information, the central node may update the model of the central node according to the second information.

[0289] After a child node stops model training, the central node receives the model information obtained by the child node training, and can update the model of the central node based on the model information obtained by the child node training; or it can wait until the model information obtained by the training of all child nodes participating in the training of the central node model is received, and then uniformly update the model of the central node.

[0290] The central node updates the model in a manner including but not limited to weighted averaging of model parameters.

[0291] It is understood that the above steps S1302-S1305 and steps S1306-S1307 can be implemented independently or in combination. The above steps S1302-S1305 can also be performed once or multiple times until the maximum number of training rounds for the first child node is reached, or because the calculated representation difference is greater than or equal to a threshold, the training of the model of the first child node is stopped.

[0292] S1310: The central node sends a new central node model. Accordingly, the first node receives the new central node model sent by the central node.

[0293] After all child nodes participating in the training of the central node model reach the maximum number of training rounds, or the training of the child node model is stopped because the calculated representation difference is greater than or equal to the threshold, the central node receives the model information obtained from the training of all child nodes participating in the training of the central node model, uses the model information of these child nodes to update the central node model, and broadcasts the new model configuration information of the central node to the child nodes again, starting a new round of training. The new round of training process can repeat the above steps S1301 to S1309.

[0294] Optionally, the central node is a third-party device that performs the aforementioned central node-related actions. For example, the above steps S1301, S1308-S1310 are all performed by a third-party device.

[0295] Optionally, the first sub-node is a third-party device that performs the aforementioned actions related to the first sub-node. For example, the above steps S1301-S1308 and S1310 are all performed by a third-party device.

[0296] Optionally, the central node is a network device. For example, the above steps S1301, S1308-S1310 are all performed by the network device.

[0297] Optionally, the first child node is a terminal device. For example, the above steps S1301-S1308 and S1310 are all performed by the terminal device.

[0298] Optionally, the central node includes a network device and a third-party device. In one example, step S1309 can be performed by a third-party device, such as an OTT or cloud server, and one or more of steps S1301, S1308, and S1310 can be performed by the network device. Furthermore, the network device and the third-party device can communicate with each other to transmit the content transmitted in one or more of steps S1301, S1308, and S1310.

[0299] Optionally, the first subnode includes a terminal device and a third-party device. In one example, one or more of steps S1302-S1307 can also be performed by a third-party device, such as an OTT or cloud server, and one or more of steps S1301, S1308, and S1310 can be performed by the terminal device. Furthermore, the terminal device and the third-party device can also communicate with each other to transmit the content transmitted in one or more of steps S1301, S1308, and S1310.

[0300] According to a distributed training method provided in an embodiment of the present application, the central node sends model configuration information, a public data set, a threshold update rule, and a representation difference coefficient update rule to the child node. The child node can calculate the representation difference between the central node model and the child node model by itself, and compare the representation difference and the threshold to determine whether to continue or stop model training. If the model training continues, the model parameters are updated using a loss function with a representation difference coefficient. By constraining the model representation difference, the dispersion of the child node update direction can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the deviation caused by packet loss can be reduced, the final performance of the central node model is improved, and the performance of distributed training is improved.

[0301] In this application, "sending information to... (e.g., a child node)" or the related illustrations in the accompanying drawings can be understood as the destination end of the information being a child node. This can include sending information to a child node directly or indirectly. "Receiving information from... (e.g., a child node)" or "receiving information from... (e.g., a child node)", or the related illustrations in the accompanying drawings can be understood as the source end of the information being a child node, which can include receiving information from a child node directly or indirectly. The information may be processed as necessary between the source end and the destination end of the information transmission, such as format changes, etc., but the destination end can understand the valid information from the source end. Similar expressions in this application can be understood similarly and will not be repeated here.

[0302] The above mainly introduces the solutions provided by the embodiments of the present application from the perspective of the interaction between various nodes. Accordingly, the embodiments of the present application also provide a distributed training device, which is used to implement the various methods described above. The distributed training device can be the central node in the above method embodiments, or a component that can be used for the central node; or, the distributed training device can be a sub-node in the above method embodiments, or a component that can be used for the sub-node. It is understandable that in order to implement the above functions, the distributed training device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0303] The embodiments of the present application can divide the functional modules of the distributed training device according to the above-mentioned method embodiments. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing unit. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical functional division. In actual implementation, other division methods can be used.

[0304] Based on the same concept of the above-mentioned distributed training method, this application also provides the following distributed training device:

[0305] FIG14 is a schematic diagram of the structure of a distributed training device provided in an embodiment of the present application. The distributed training device 1400 includes a transceiver unit 1401 and a processing unit 1402.

[0306] When the distributed training device is used to implement the functions of the first subnode in the above-mentioned method embodiment, the transceiver unit 1401 is used to perform one or more of the operations of the first subnode in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 of the embodiment shown in FIG10 , and the processing unit 1402 is used to perform one or more of steps S1002, S1006, S1007, and S1011 of the embodiment shown in FIG10 ; or, the transceiver unit 1401 is used to perform one or more of the operations of the first subnode in steps S1301, S1308, and S1310 of the embodiment shown in FIG13 , and the processing unit 1402 is used to perform one or more of steps S1302 to S1307 of the embodiment shown in FIG13 . Optionally, the distributed training device can be a terminal device, or a third-party device such as an OTT or cloud server, or a system consisting of a terminal device and a third-party device.

[0307] When the distributed training device is used to implement the functions of the central node in the above-mentioned method embodiment, the transceiver unit 1401 is used to perform one or more of the operations of the central node in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 of the embodiment shown in FIG10 , and the processing unit 1402 is used to perform one or more of steps S1004, S1009, and S1013 of the embodiment shown in FIG10 ; or, the transceiver unit 1401 is used to perform one or more of the operations of the central node in steps S1301, S1308, and S1310 of the embodiment shown in FIG13 , and the processing unit 1402 is used to perform one or more of step S1309 of the embodiment shown in FIG13 . Optionally, the distributed training device can be a network device, or a third-party device such as an OTT or cloud server, or a system consisting of a network device and a third-party device.

[0308] For the specific implementation of the above-mentioned transceiver unit 1401 and the processing unit 1402, reference may be made to the description in the above-mentioned method embodiment. In addition, it should be noted that the above-mentioned transceiver unit and / or processing unit can be implemented by a virtual module. For example, the processing unit can be implemented by a software function unit or a virtual device, and the transceiver unit can be implemented by a software function or a virtual device. Alternatively, the processing unit or the transceiver unit can also be implemented by a physical circuit. For example, if the device is implemented using a chip / chip circuit, the transceiver unit can be an input and output circuit and / or a communication interface to perform input operations (corresponding to the above-mentioned receiving operations) and output operations (corresponding to the above-mentioned sending operations); the processing unit is a processing circuit, such as an integrated processor or microprocessor or integrated circuit.

[0309] The division of modules in this application is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the examples of this application may be integrated into a single processor, exist physically as separate modules, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in either hardware or software functional modules.

[0310] As shown in Figure 15, it is a structural diagram of another distributed training device provided in an embodiment of the present application. The distributed training device 1500 includes one or more processing circuits 1501 (one processing circuit is illustrated in the figure). Optionally, the distributed training device 1500 may also include a memory 1503 (indicated by a dotted line in the figure). The memory 1503 is used to store instructions executed by the processing circuit 1501, or to store input data required for the processing circuit 1501 to run instructions, or to store data generated after the processing circuit 1501 runs instructions. Optionally, the distributed training device 1500 may also include an interface circuit 1502 (indicated by a dotted line in the figure), and the processing circuit 1501 and the interface circuit 1502 are coupled to each other. It will be understood that the interface circuit 1502 can be a transceiver or an input-output interface.

[0311] The processing circuit may be a processor or a circuit in a processor used for processing.

[0312] When the distributed training device is used to implement the functions of the first subnode in the above method embodiment, the interface circuit 1502 is used to perform one or more of the operations of the first subnode in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 of the embodiment shown in FIG10 , and the processing circuit 1501 is used to perform one or more of steps S1002, S1006, S1007, and S1011 of the embodiment shown in FIG10 ; or, the interface circuit 1502 is used to perform one or more of the operations of the first subnode in steps S1301, S1308, and S1310 of the embodiment shown in FIG13 , and the processing circuit 1501 is used to perform one or more of steps S1302 to S1307 of the embodiment shown in FIG13 . Optionally, the distributed training device can be a terminal device, or a third-party device such as an OTT or cloud server, or a system consisting of a terminal device and a third-party device.

[0313] When the distributed training device is used to implement the functions of the central node in the above-mentioned method embodiment, the interface circuit 1502 is used to perform one or more of the operations of the central node in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 of the embodiment shown in FIG10 , and the processing circuit 1501 is used to perform one or more of steps S1004, S1009, and S1013 of the embodiment shown in FIG10 ; or, the interface circuit 1502 is used to perform one or more of the operations of the central node in steps S1301, S1308, and S1310 of the embodiment shown in FIG13 , and the processing circuit 1501 is used to perform one or more of step S1309 of the embodiment shown in FIG13 . Optionally, the distributed training device can be a network device, or a third-party device such as an OTT or cloud server, or can be a system composed of a network device and a third-party device.

[0314] When the above-mentioned distributed training device is a chip applied to the central node, the chip implements the function of the central node in the above-mentioned method embodiment. The chip receives information from other modules in the central node, and the information is sent by the child node to the central node; or, the chip sends information to other modules in the central node, and the information is sent by the central node to the child node. When the central node is a network device, the module of the central node here can be the baseband chip of the central node, or it can be a CU, DU or other module, or it can be a device under the open radio access network (O-RAN) architecture, such as an open CU, open DU and other devices. When the central node is a third-party device, the module of the child node here can be a processing chip of the third-party device. Among them, the processing chip can be used to implement AI training.

[0315] When the above-mentioned distributed training device is a chip applied to a sub-node, the chip implements the functions of the sub-node in the above-mentioned method embodiment. The chip receives information from other modules in the sub-node, and the information is sent by the central node to the sub-node; or, the chip sends information to other modules in the sub-node, and the information is sent by the sub-node to the central node. When the sub-node is a terminal device, the module of the sub-node here can be the baseband chip of the sub-node, or, the baseband chip and the processing chip. Among them, the processing chip can be used to implement AI training. When the sub-node is a third-party device, the module of the sub-node here can be the processing chip of the third-party device. Among them, the processing chip can be used to implement AI training.

[0316] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed, the method in the above embodiment is implemented.

[0317] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method in the above embodiment.

[0318] An embodiment of the present application also provides a distributed training system, including the above-mentioned distributed training device.

[0319] The present application also provides a circuit, which is coupled to a memory and is used to execute the method shown in the above embodiment. The circuit may include a chip circuit.

[0320] Optionally, an embodiment of the present application further provides a chip system, comprising: at least one processor and an interface, wherein the at least one processor is coupled to a memory via the interface, and when the at least one processor executes a computer program or instruction in the memory, the chip system executes the method in any of the above method embodiments. Optionally, the chip system may be composed of a chip, or may include a chip and other discrete devices, which is not specifically limited in the embodiments of the present application.

[0321] The memory in the present application may also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data. A memory is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. For example, the memory may be a non-volatile memory, such as a digital versatile disc (DVD), a hard disk drive (HDD), or a solid-state drive (SSD), or a volatile memory, such as a random-access memory (RAM).

[0322] As used in the following description of this application, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0323] It should be understood that in the description of this application, unless otherwise specified, " / " indicates that the objects associated with each other are in an "or" relationship. For example, A / B can mean A or B; where A and B can be singular or plural. Also, in the description of this application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural. In addition, to facilitate the clear description of the technical solutions of the embodiments of this application, in the embodiments of this application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean different. At the same time, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0324] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0325] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art can understand and implement other changes to the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0326] It is understood that the various numbers used in the embodiments of this application are merely for ease of description and are not intended to limit the scope of the embodiments of this application. The order of the sequence numbers of the above-mentioned processes does not necessarily imply a specific order of execution; the order of execution of the processes should be determined by their functions and inherent logic.

[0327] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0328] The components in the device of the embodiment of the present application can be merged, divided, or deleted according to actual needs. Those skilled in the art can combine or combine the different embodiments and features of the different embodiments described in this specification.

[0329] In this application, under the premise of no logical contradiction, the examples can reference each other, for example, the methods and / or terms between method embodiments can reference each other, for example, the functions and / or terms between device embodiments can reference each other, for example, the functions and / or terms between device examples and method examples can reference each other.

Claims

1. A distributed training method, characterized in that: The method comprises: Receiving first information from a first child node, the first information being used to indicate a first representation difference between a model of a central node and a model of the first child node, the first child node being any one of a plurality of child nodes participating in model training of the central node; Sending second information to the first sub-node, where the second information is obtained based on the first representation difference, and the second information includes training indication information of the model of the first sub-node.

2. The method according to claim 1, characterized in that The first representation difference satisfies a first condition, and the first training instruction information is used to instruct the first sub-node to continue training of the model of the first sub-node.

3. The method according to claim 2, characterized in that The training indication information is also used to indicate the representation difference coefficient in the loss function of the model of the first sub-node.

4. The method according to claim 2, characterized in that The second information also includes a representation difference coefficient in the loss function of the model of the first subnode.

5. The method according to claim 3 or 4, characterized in that The greater the first characterization difference, the greater the characterization difference coefficient.

6. The method according to claim 1, characterized in that The first representation difference satisfies a second condition, and the first training instruction information is used to instruct the first sub-node to stop training of the model of the first sub-node.

7. The method according to any one of claims 1 to 6, characterized in that: The first representation difference is obtained based on a local data set or a public data set of the first child node.

8. The method according to any one of claims 1 to 7, characterized in that The first characterization difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, and / or the first characterization difference is the difference between the output or intermediate quantity of the models of the multiple child nodes, wherein the output or intermediate quantity is obtained based on the same input.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Broadcast third information, the third information comprising at least one of the following: model configuration information, a public data set; wherein the model configuration information is used to indicate at least one of the following: types of models of the multiple child nodes, structural information of the models of the multiple child nodes, model parameters of the models of the multiple child nodes, or training configuration information of the models of the multiple child nodes.

10. A distributed training method, characterized in that: The method comprises: Sending first information to a central node, where the first information is used to indicate a first representation difference between a model of the central node and a model of a first child node, where the first child node is any one of a plurality of child nodes participating in model training of the central node; Receive second information from the central node, where the second information is obtained based on the first representation difference, and the second information includes training indication information of the model of the first sub-node.

11. The method according to claim 10, characterized in that The first representation difference satisfies a first condition, and the training instruction information is used to instruct the first sub-node to continue training of the model of the first sub-node.

12. The method according to claim 11, characterized in that The method further comprises: The model of the first child node is updated according to the second information.

13. The method according to claim 11 or 12, characterized in that The training indication information is also used to indicate the representation difference coefficient in the loss function of the model of the first sub-node.

14. The method according to claim 11 or 12, characterized in that: The second information also includes a representation difference coefficient in the loss function of the model of the first subnode.

15. The method according to claim 13 or 14, characterized in that The greater the first characterization difference, the greater the characterization difference coefficient.

16. The method according to claim 10, characterized in that The first representation difference satisfies a second condition, and the training instruction information is used to instruct the first sub-node to stop training of the model of the first sub-node.

17. The method according to any one of claims 10 to 16, characterized in that: The first representation difference is obtained based on a local data set or a public data set of the first child node.

18. The method according to any one of claims 10 to 17, characterized in that: The first characterization difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, and / or the first characterization difference is the difference between the output or intermediate quantity of the models of the multiple child nodes, wherein the output or intermediate quantity is obtained based on the same input.

19. The method according to any one of claims 10 to 18, characterized in that: The method further comprises: Receive third information, the third information comprising at least one of the following: model configuration information, a public data set; wherein the model configuration information is used to indicate at least one of the following: the type of the model of the multiple sub-nodes, the structural information of the model of the multiple sub-nodes, the model parameters of the model of the multiple sub-nodes, or the training configuration information of the model of the multiple sub-nodes.

20. A distributed training method, characterized in that: The method comprises: Sending first information, the first information including at least one of the following: model configuration information, a public data set, a threshold update rule, and a representation difference coefficient update rule, the threshold is used to compare the representation difference between the model of the central node and the model of the first child node, and the representation difference coefficient is a parameter in the loss function of the model of the first child node; Second information is received, where the second information includes model information obtained by training the first sub-node.

21. The method of claim 20, wherein: The model configuration information is used to indicate at least one of the following: the types of models of multiple child nodes participating in the training of the model of the central node, structural information of the models of the multiple child nodes, model parameters of the models of the multiple child nodes, or training configuration information of the models of the multiple child nodes.

22. The method according to claim 20 or 21, characterized in that The updating rule of the threshold is that the threshold decreases as the number of communications between the first sub-node and the central node increases.

23. The method according to any one of claims 20 to 22, characterized in that The updating rule of the characterization difference coefficient is to use the linear or nonlinear scaling of the characterization difference coefficient updated last time as the update amount of the characterization difference coefficient compared with the characterization difference coefficient of the last time.

24. The method according to any one of claims 20 to 23, characterized in that The greater the characterization difference, the greater the characterization difference coefficient.

25. A distributed training method, characterized in that: The method comprises: Calculating a first representation difference between the model of the central node and the model of the first sub-node based on the local data set or the public data set of the first sub-node; When the first representation difference is less than or equal to a first threshold, continuing training of the model of the first subnode; The model of the first child node is updated using the local data set and a loss function with a representation difference coefficient; wherein the greater the first representation difference, the greater the representation difference coefficient.

26. The method of claim 25, wherein: The method further comprises: First information is received, where the first information includes at least one of the following: model configuration information, a public data set, a threshold update rule, and a characterization difference coefficient update rule.

27. The method according to claim 25 or 26, characterized in that The method further comprises: The first threshold is updated based on the threshold update rule to obtain a second threshold.

28. The method according to any one of claims 25 to 27, characterized in that The method further comprises: The characterization difference coefficient is updated based on an update rule of the characterization difference coefficient.

29. The method according to any one of claims 25 to 28, characterized in that The method further comprises: When the first representation difference is greater than the first threshold, the training of the model of the first subnode is stopped.

30. The method according to any one of claims 25 to 29, characterized in that The model configuration information is used to indicate at least one of the following: the types of models of multiple child nodes participating in the training of the model of the central node, structural information of the models of the multiple child nodes, model parameters of the models of the multiple child nodes, or training configuration information of the models of the multiple child nodes.

31. The method according to any one of claims 25 to 30, characterized in that The updating rule of the threshold is that the threshold decreases as the number of communications between the first sub-node and the central node increases.

32. The method according to any one of claims 25 to 31, characterized in that The updating rule of the characterization difference coefficient is to use the linear or nonlinear scaling of the characterization difference coefficient updated last time as the update amount of the characterization difference coefficient compared with the characterization difference coefficient of the last time.

33. The method according to any one of claims 25 to 32, characterized in that The greater the characterization difference, the greater the characterization difference coefficient.

34. A distributed training device, characterized in that: Comprising a module for executing the method as claimed in any one of claims 1 to 9, or comprising a module for executing the method as claimed in any one of claims 10 to 19, or comprising a module for executing the method as claimed in any one of claims 20 to 24, or comprising a module for executing the method as claimed in any one of claims 25 to 33.

35. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program or instruction. When the computer program or instruction is executed, it implements the method as described in any one of claims 1 to 9, or implements the method as described in any one of claims 10 to 19, or implements the method as described in any one of claims 20 to 24, or implements the method as described in any one of claims 25 to 33.

36. A computer program product, characterized in that The computer program product comprises program instructions, which, when executed, implement the method according to any one of claims 1 to 9, or the method according to any one of claims 10 to 19, or the method according to any one of claims 20 to 24, or the method according to any one of claims 25 to 33.

Citation Information

Patent Citations

  • Distributed training method and device

    CN120197733A

  • Recommendation method and system based on graph convolutional neural network, and storage medium

    CN114021018A

  • Federal learning model training method for large-scale industrial chain privacy calculation

    CN114169412A

  • Federal learning method and device, equipment and medium

    CN115511103A

  • Intelligent model training method and device

    CN116362334A