Distributed training method and device

The difference is characterized by the child node feedback model and the model training instructions are carried out based on the difference, which solves the problem of non-independent and homogeneous data distribution and unstable transmission in distributed training, and improves the model performance and stability.

CN120197733APending Publication Date: 2025-06-24HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311777211.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In distributed training, since the central node cannot access user data, the data of the child nodes is non-independent and homogeneously distributed, the model performance cannot be guaranteed, and the unstable wireless transmission feedback gradient causes fluctuations in model parameter updates.

Method used

By letting children to feedback the characterization differences between models to the central node, the central node performs model training instructions based on this difference, ensuring that a high-performance machine learning model is obtained when a finite secondary child node communicates with the central node.

Benefits of technology

The performance of distributed training is improved, ensuring the stability and efficiency of the model under non-independent homogeneous data and unstable transmission conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197733A_ABST
    Figure CN120197733A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence AI model distributed training method and device. According to the technical scheme, the child nodes participating in training of the center node model feed back the characterization difference between the models to the center node, and the center node carries out model training indication according to the characterization difference, so that a high-performance machine learning model can be obtained in the process of communication between the child nodes and the center node in the finite number, and the performance of distributed training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a distributed training method and apparatus. Background Art

[0002] Distributed training refers to dividing the machine training process among multiple sub-compute nodes (abbreviated as "sub-nodes"), which allows a central node to collect machine learning models trained by multiple sub-nodes to improve the effect of the entire machine learning training.

[0003] Distributed training aims to obtain a high-performance machine learning model during a limited number of communications between sub-nodes and the central node. However, since the central node cannot access user data, when the data of sub-nodes is non-independent and identically distributed, the performance of the model cannot be guaranteed. Moreover, when sub-nodes feedback gradients through unstable wireless transmissions, the number of gradients of sub-nodes successfully received by the central node has a certain randomness, which will bring fluctuations to the model parameter updates of the central node.

[0004] In view of this, how to improve the performance of distributed training is an urgent problem to be solved. Summary of the Invention

[0005] This application provides a distributed training method and apparatus to improve the performance of distributed training.

[0006] In a first aspect, a distributed training method is provided. The method includes: receiving first information from a first sub-node, where the first information is used to indicate a first representation difference between a model of a central node and a model of the first sub-node, and the first sub-node is any one of multiple sub-nodes participating in the training of the model of the central node; and sending second information to the first sub-node, where the second information is obtained according to the first representation difference, and the second information includes training instruction information for the model of the first sub-node.

[0007] In this aspect, the central node receives the representation differences between the models fed back by the sub-nodes participating in the training of the central node model, and the central node performs model training instructions based on the representation differences, so that a high-performance machine learning model can be obtained during a limited number of communications between sub-nodes and the central node, improving the performance of distributed training.

[0008] In combination with the first aspect, in a possible implementation, the method further includes: broadcasting third information, where the third information includes at least one of the following: model configuration information, a common data set; where the model configuration information is used to indicate at least one of the following: the type of the models of the multiple sub-nodes, the structural information of the models of the multiple sub-nodes, the model parameters of the models of the multiple sub-nodes, or the training configuration information of the models of the multiple sub-nodes.

[0009] In this implementation, the third information is used for the model training of the first child node.

[0010] By enabling each child node to calculate the first representation difference based on the same common data set, the representation differences between the models of each child node and the model of the central node, or the representation differences between the models of each child node, can be more accurately compared.

[0011] In a second aspect, a distributed training method is provided. The method includes: sending first information to a central node, where the first information is used to indicate a first representation difference between the model of the central node and the model of a first child node, and the first child node is any one of a plurality of child nodes participating in the model training of the central node; and receiving second information from the central node, where the second information is obtained based on the first representation difference, and the second information includes training instruction information for the model of the first child node.

[0012] In this aspect, the child nodes participating in the training of the central node model feedback the representation differences between the models to the central node, and the central node gives training instructions for the model based on the representation differences, so that a high-performance machine learning model can be obtained during a limited number of communications between the child nodes and the central node, improving the performance of distributed training.

[0013] In combination with the second aspect, in a possible implementation, the method further includes: updating the model of the first child node according to the second information.

[0014] In combination with the second aspect, in another possible implementation, the method further includes: receiving third information, where the third information includes at least one of the following: model configuration information, common data set; where the model configuration information is used to indicate at least one of the following: the type of the models of the plurality of child nodes, the structural information of the models of the plurality of child nodes, the model parameters of the models of the plurality of child nodes, or the training configuration information of the models of the plurality of child nodes.

[0015] In this implementation, the third information is used for the model training of the first child node.

[0016] By enabling each child node to calculate the first representation difference based on the same common data set, the representation differences between the models of each child node and the model of the central node, or the representation differences between the models of each child node, can be more accurately compared.

[0017] Implementing in combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the first representation difference satisfies a first condition, and the training indication information is used to instruct the first child node to continue training the model of the first child node. Optionally, the first representation difference satisfying the first condition includes: the first representation difference is less than or equal to a first threshold. Optionally, the first condition, such as the first threshold, is pre-determined by the central node or pre-defined by the protocol.

[0018] Implementing in combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the first representation difference satisfies a second condition, and the training indication information is used to instruct the first child node to stop training the model of the first child node. Optionally, the first representation difference satisfying the first condition includes: the first representation difference is greater than or equal to a second threshold. Optionally, the second threshold is the same as or different from the aforementioned first threshold. Optionally, the second condition, such as the second threshold, is pre-determined by the central node or pre-defined by the protocol.

[0019] Implementing in combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the training indication information is further used to indicate the representation difference coefficient in the loss function of the model of the first child node.

[0020] In this implementation, the central node can also allocate the parameters of the corresponding loss function, such as the representation difference coefficient, to the child node with a large difference value according to the difference value analysis result.

[0021] Implementing in combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the second information further includes the representation difference coefficient in the loss function of the model of the first child node.

[0022] Implementing in combination with the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the greater the first representation difference, the greater the representation difference coefficient.

[0023] In this implementation, the greater the first representation difference, the greater the representation difference coefficient. Because the greater the first representation difference, it indicates that the model trained by this child node deviates further from the model of the central node, or among all the child nodes participating in the model training of the central node, the model trained by this child node deviates further from the models trained by other child nodes. Therefore, the central node allocates a representation difference coefficient, enabling the child node to update its loss function according to this representation difference coefficient, so that the impact of the update of the subsequent child node's model on the training result is smaller.

[0024] In combination with the implementation of the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the first representation difference is obtained based on the local data set or the common data set of the first child node.

[0025] In this implementation, by enabling each child node to calculate the first representation difference based on the same local data set or common data set, the representation differences between the models of each child node and the model of the central node, or the representation differences between the models of each child node, can be more accurately compared.

[0026] In combination with the implementation of the first aspect, the second aspect, or any one of the first aspect and the second aspect, in another possible implementation, the first representation difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, and / or the first representation difference is the difference between the output or intermediate quantity of the models of the multiple child nodes, where the output or intermediate quantity is obtained based on the same input.

[0027] In this implementation, by obtaining the difference between the output or intermediate quantity of the models based on the same input as the representation difference of the models, the differences between the models can be accurately represented.

[0028] The method of the first aspect above can be executed by the central node, or by a module applied to the central node (such as a processor, a chip, or a chip system, etc.), or can also be implemented by a logical node, a logical module, or software that can implement all or part of the functions of the central node.

[0029] The method of the second aspect above can be executed by the child node, or by a module applied to the child node (such as a processor, a chip, or a chip system, etc.), or can also be implemented by a logical node, a logical module, or software that can implement all or part of the functions of the child node.

[0030] In a third aspect, a distributed training method is provided. The method includes: sending first information, where the first information includes at least one of the following: model configuration information, a common data set, an update rule for a threshold, and an update rule for a representation difference coefficient, the threshold being used for comparing the representation difference between the model of the central node and the model of the first child node, and the representation difference coefficient being a parameter in the loss function of the model of the first child node; and receiving second information, where the second information includes the model information obtained by training the first child node.

[0031] In this aspect, the central node issues model configuration information, a common data set, an update rule for a threshold value, and an update rule for a representation difference coefficient to child nodes. The child nodes can calculate the representation difference between the model of the central node and their own models, compare the representation difference with the threshold value, and determine whether to continue or stop the training of the model. If the training of the model continues, the model parameters are updated using a loss function with a representation difference coefficient. By constraining the model representation difference, the divergence of the update directions of the child nodes can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the bias caused by packet loss can be alleviated, the final performance of the central node model can be improved, and the performance of distributed training can be enhanced.

[0032] Combined with the third aspect, in a possible implementation, the method further includes: updating the model of the central node according to the second information.

[0033] In a fourth aspect, a distributed training method is provided. The method includes: calculating a first representation difference between the model of the central node and the model of the first child node based on the local data set or the common data set of the first child node; continuing the training of the model of the first child node when the first representation difference is less than or equal to a first threshold value; and updating the model of the first child node using the local data set and a loss function with a representation difference coefficient; wherein, the greater the first representation difference, the greater the representation difference coefficient.

[0034] In this aspect, the child nodes can calculate the representation difference between the model of the central node and their own models, compare the representation difference with the threshold value, and determine whether to continue or stop the training of the model. If the training of the model continues, the model parameters are updated using a loss function with a representation difference coefficient. By constraining the model representation difference, the divergence of the update directions of the child nodes can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the bias caused by packet loss can be alleviated, the final performance of the central node model can be improved, and the performance of distributed training can be enhanced.

[0035] Combined with the fourth aspect, in a possible implementation, the method further includes: receiving first information, where the first information includes at least one of the following: model configuration information, a common data set, an update rule for a threshold value, an update rule for the representation difference coefficient.

[0036] In this implementation, the update rule of the threshold can be pre - defined by the central node. For example, the update rule is defined as follows: the threshold decreases as the number of communications between the first child node and the central node increases, and the threshold of the first child node remains unchanged before the first child node communicates with the central node. Here, when the central node sends a new model to the first child node, it is considered as one communication. It can be understood that as the child nodes train and the central node aggregates the models of the child nodes, the models of both the central node and the child nodes gradually converge. Therefore, as the number of communications between the first child node and the central node increases, the representational difference between the model of the central node and the model of the first child node will gradually decrease. Therefore, a threshold update rule can be defined such that the threshold decreases as the number of communications between the first child node and the central node increases.

[0037] The update rule of the representational difference coefficient is used to indicate that after the child node calculates the representational difference, an updated representational difference coefficient is obtained based on the initial representational difference coefficient and the calculated representational difference. The update rule of the representational difference coefficient is pre - defined by the central node. For example, the update rule of the representational difference coefficient is: the linear or non - linear scaling of the previous updated representational difference coefficient is used as the update amount of the representational difference coefficient compared to the previous representational difference coefficient. It can be understood that the greater the representational difference, the greater the representational difference coefficient. Because the greater the representational difference indicates that the model trained by this child node deviates further from the model of the central node, or among all the child nodes participating in the model training of the central node, the model trained by this child node deviates further from the models trained by other child nodes. Therefore, the child node obtains an updated representational difference coefficient, enabling the child node to update its loss function according to this representational difference coefficient, making the impact of the subsequent update of the model of the child node on the training result smaller.

[0038] Combined with the fourth aspect, in another possible implementation, the method further includes: updating the first threshold based on the update rule of the threshold to obtain a second threshold.

[0039] Combined with the fourth aspect, in yet another possible implementation, the method further includes: updating the representational difference coefficient based on the update rule of the representational difference coefficient.

[0040] Combined with the fourth aspect, in yet another possible implementation, the method further includes: stopping the training of the model of the first child node when the first representational difference is greater than the first threshold.

[0041] Combined with the third aspect or the fourth aspect, in yet another possible implementation, the model configuration information is used to indicate at least one of the following: the types of the models of multiple child nodes participating in the model training of the central node, the structural information of the models of the multiple child nodes, the model parameters of the models of the multiple child nodes, or the training configuration information of the models of the multiple child nodes.

[0042] In yet another possible implementation, the update rule of the threshold is that the threshold decreases as the number of communications between the first child node and the central node increases.

[0043] In yet another possible implementation, the update rule of the representation difference coefficient is that the linear or non-linear scaling of the representation difference coefficient updated last time is used as the update amount of the representation difference coefficient compared with the representation difference coefficient of the previous time.

[0044] In yet another possible implementation, the greater the representation difference, the greater the representation difference coefficient.

[0045] In a fifth aspect, a distributed training device is provided for implementing the distributed training method in any implementation of the above first aspect, third aspect, or the first aspect and the third aspect. The device may be a central node, or a module applied to the central node (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of the central node. In one implementation, the distributed training device may include a sending unit, a receiving unit, and may further include a processing unit. The sending unit and the receiving unit may be independent or combined together (which may be referred to as a "transceiving unit").

[0046] In a sixth aspect, a distributed training device is provided for implementing the distributed training method in any implementation of the above second aspect, fourth aspect, or the second aspect and the fourth aspect. The device may be a child node, or a module applied to the child node (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of the child node. In one implementation, the distributed training device may include a sending unit, a receiving unit, and may further include a processing unit. The sending unit and the receiving unit may be independent or combined together (which may be referred to as a "transceiving unit").

[0047] In a possible implementation manner, the distributed training devices in the above fifth aspect to the sixth aspect include modules for respectively executing the methods in any aspect or any implementation of the above first aspect to the second aspect.

[0048] In another possible implementation, the distributed training device in the fifth to sixth aspects described above includes a processing circuit coupled to a memory; the processing circuit is configured to implement the corresponding functions of the device in executing the distributed training method described above. The memory is used to be coupled to the processing circuit and stores the necessary programs (instructions) and / or data of the device. Optionally, the distributed training device may further include a communication interface for implementing communication between the device and other network elements. Optionally, the memory may be located inside the distributed training device or outside the distributed training device. Exemplarily, the processing circuit may be a processor or a circuit for processing in the processor.

[0049] When the distributed training device in the fifth to sixth aspects described above is a chip, the sending unit may be an output unit, such as an output circuit or a communication interface; the receiving unit may be an input unit, such as an input circuit or a communication interface. When the distributed training device is a terminal, the sending unit may be a transmitter or a transceiver; the receiving unit may be a receiver or a receiver.

[0050] In a seventh aspect, a computer-readable storage medium is provided, in which a computer program or instruction is stored, and when the computer program or instruction is executed, the methods described in the above aspects are implemented.

[0051] In an eighth aspect, a computer program product containing instructions is provided, and when the instructions run on a distributed training device, the distributed training device is caused to execute the methods described in the above aspects.

[0052] In a ninth aspect, a distributed training system is provided, and the distributed training system includes the distributed training device described in the fifth aspect and the distributed training device described in the sixth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 FIG. is a schematic diagram of the architecture of a distributed training system provided by an embodiment of the present application;

[0054] Figure 2 FIG. is a simplified schematic diagram of a wireless communication system provided by an embodiment of the present application;

[0055] Figure 3 FIG. is a schematic diagram of the architecture of another distributed training system provided by the present application;

[0056] Figures 4A to 4D FIG. is a schematic diagram of the network architecture of an embodiment of the present application;

[0057] Figure 5 FIG. is a schematic diagram of a neuron structure;

[0058] Figure 6Schematic diagram of a neural network;

[0059] Figure 7 Schematic diagram of an AI application framework;

[0060] Figure 8 Schematic diagram of the architecture of another communication system provided by an embodiment of the present application;

[0061] Figure 9 Schematic diagram of the process of a distributed training method provided by an embodiment of the present application;

[0062] Figure 10 Schematic diagram of the process of another distributed training method provided by an embodiment of the present application;

[0063] Figure 11 Schematic diagram of the calculation of a representation difference exemplified by an embodiment of the present application;

[0064] Figure 12 System block diagram of a distributed training exemplified by an embodiment of the present application;

[0065] Figure 13 Schematic diagram of the process of yet another distributed training method provided by an embodiment of the present application;

[0066] Figure 14 Schematic diagram of the structure of a distributed training device provided by an embodiment of the present application;

[0067] Figure 15 Schematic diagram of the structure of another distributed training device provided by an embodiment of the present application. Detailed implementation manners

[0068] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0069] The embodiments of the present application can be applied to a distributed training system as shown in Figure 1 The distributed training system includes a central node and K sub-nodes, where K is a positive integer. Exemplarily, the distributed training system can be a federated learning system or a Gossip learning system. Data and model information, etc., can be transmitted between the central node and each sub-node.

[0070] The machine learning model trained by the distributed training system can be for non-wireless communication services, such as image recognition, natural language processing, etc., or for wireless communication services, such as beam selection based on environmental information.

[0071] The technology provided by the present application can be applied to various communication systems. For example, the communication system can be the fourth generation (4 thgeneration, 4G) communication systems (such as Long Term Evolution (LTE) systems), fifth-generation (5 th generation, 5G) communication systems, Worldwide Interoperability for Microwave Access (WiMAX), Wireless Local Area Network (WLAN) systems, satellite communication systems, integrated systems of multiple systems, or future communication systems, such as sixth-generation (6 th generation, 6G) communication systems, etc. Among them, the 5G communication system can also be referred to as the New Radio (NR) system.

[0072] A network element in a communication system can send a signal to another network element or receive a signal from another network element. The signal can include information, signaling, data, etc. Herein, the network element can also be replaced with an entity, a network entity, a device, a terminal device, a communication module, a node, a communication node, etc. In this application, the network element is used as an example for description. For example, a communication system can include at least one terminal device and at least one network device. The network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the network device. In addition, it can be understood that if there are multiple terminal devices in the communication system, the multiple terminal devices can also send signals to each other, that is, both the signal sending network element and the signal receiving network element can be terminal devices.

[0073] See Figure 2 , Figure 2 is a simplified schematic diagram of the wireless communication system provided by the embodiment of this application. As Figure 2 shown, the wireless communication system includes a radio access network 100. The radio access network 100 can be a next-generation (such as 6G or higher) radio access network or a traditional (such as 5G, 4G) radio access network. One or more terminal devices (120a - 120j, collectively referred to as 120) can be connected to each other or connected to one or more network devices (110a, 110b, collectively referred to as 110) in the radio access network 100. Optionally, Figure 2 This is just a schematic diagram, and the wireless communication system may also include other devices, such as a core network device, a wireless relay device, and / or a wireless backhaul device, etc., which are not drawn in Figure 2 .

[0074] Optionally, in practical applications, the wireless communication system may include multiple network devices (also referred to as access network devices) at the same time, or may include multiple terminal devices at the same time. A network device may serve one or more terminal devices at the same time. A terminal device may also access one or more network devices at the same time. The embodiments of the present application do not limit the number of terminal devices and network devices included in the wireless communication system.

[0075] Among them, the network device can be an entity on the network side for transmitting or receiving signals. The network device can be an access device for the terminal device to access the wireless communication system wirelessly. For example, the network device can be a base station. The base station can generally cover various names in the following, or be replaced with the following names. For example: radio access network (RAN) node, Node B, evolved Node B (eNB), next generation Node B (gNB), network device in open radio access network (O-RAN), relay station, access point, transmitting and receiving point (TRP), transmitting point (TP), master eNB (MeNB), secondary eNB (SeNB), multi-standard radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, building baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), centralized unit (CU), distributed unit (DU), radio unit (RU), centralized unit control plane (CU controlplane, CU-CP) node, centralized unit user plane (CU user plane, CU-UP) node, positioning node, RAN intelligent controller (RIC), etc. The base station can be a macro base station, micro base station, relay node, donor node or the like, or a combination thereof. The network device can also refer to a communication module, a modem or a chip used in the foregoing devices or apparatuses. The network device can also be a mobile switching center and a device that undertakes the function of a base station in device-to-device (D2D), vehicle-to-everything (V2X), machine-to-machine (M2M) communication, a network-side device in a 6G network, a device that undertakes the function of a base station in a future communication system, etc.The network device may support networks with the same or different access technologies. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the network device.

[0076] The network device may be fixed or mobile. For example, base stations 110a and 110b are stationary and are responsible for wireless transmission and reception in one or more cells from the terminal device 120. Figure 2 The helicopter or drone 120i shown in may be configured to act as a mobile base station, and one or more cells may move according to the position of the mobile base station 120i. In other examples, the helicopter or drone (120i) may be configured to be used as a terminal device communicating with the base station 110b.

[0077] In the present application, the communication device for implementing the above access network function may be a network device, or a network device with partial functions of accessing the network, or a device capable of supporting the implementation of the access network function, such as a chip system, a hardware circuit, a software module, or a combination of a hardware circuit and a software module. This device may be installed in the network device or used in combination with the network device. In the method of the present application, the communication device for implementing the network device function is described by taking the network device as an example.

[0078] A terminal device can be an entity on the user side for receiving or transmitting signals, such as a mobile phone. The terminal device can be used to connect people, things, and machines. The terminal device can communicate with one or more core networks through network devices. The terminal device includes a handheld device with wireless connection capabilities, other processing devices connected to a wireless modem, or in-vehicle devices, etc. The terminal device can be a portable, pocket-sized, handheld, computer-integrated, or in-vehicle mobile device. The terminal device 120 can be widely applied in various scenarios, such as cellular communication, D2D, V2X, end-to-end (point-to-point, P2P), machine-to-machine (M2M), machine type communication (MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, telemedicine, smart grid, smart furniture, smart office, smart wearables, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.Some examples of the terminal device 120 are: user equipment (UE) compliant with the 3GPP standard, fixed devices, mobile devices, handheld devices, wearable devices, cellular phones, smart phones, session initiation protocol (SIP) phones, laptop computers, personal computers, smart books, vehicles, satellites, global positioning system (GPS) devices, target tracking devices, drones, helicopters, aircraft, ships, remote control devices, smart home devices, industrial devices, personal communication service (PCS) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), wireless network cameras, tablet computers, palmtop computers, mobile internet devices (MIDs), wearable devices such as smart watches, VR devices, AR devices, wireless terminals in industrial control, terminals in vehicle-to-everything (V2X) systems, wireless terminals in self-driving, wireless terminals in smart grid, wireless terminals in transportation safety, wireless terminals in smart city such as smart fuel dispensers, terminal devices on high-speed trains, and wireless terminals in smart home, such as smart speakers, smart coffee machines, smart printers, etc. The terminal device 120 can be a wireless device in the above various scenarios or a device for being disposed in a wireless device, for example, a communication module, a modem, or a chip in the above devices. The terminal device can also be referred to as a terminal, a terminal device, a UE, a mobile station (MS), a mobile terminal (MT), etc. The terminal device can also be a terminal device in a future wireless communication system. The terminal device can be used in a dedicated network device or a general device. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0079] Optionally, the terminal device can be used to act as a base station. For example, a UE can act as a scheduling entity that provides sidelink signals between UEs in V2X, D2D, or P2P, etc. As Figure 2 shown, the cellular phone 120a and the vehicle 120b communicate with each other using sidelink signals. The cellular phone 120a communicates with the smart home device 120e without relaying communication signals through the base station 110b.

[0080] In this application, the communication device for implementing the functions of the terminal device may be the terminal device itself, or a terminal device with some of the functions of the above terminal device, or a device capable of supporting the implementation of the functions of the above terminal device, such as a chip system. This device may be installed in the terminal device or used in conjunction with the terminal device. In this application, the chip system may be composed of chips or may include chips and other discrete devices. In the technical solution provided in this application, the communication device is described by taking the terminal device or UE as an example.

[0081] Optionally, a wireless communication system usually consists of cells, and the base station provides the management of the cells. The base station provides communication services to multiple mobile stations (MS) in the cell. The base station includes a baseband unit (BBU) and a remote radio unit (RRU). The BBU and the RRU can be placed in different locations. For example, the RRU is remote and placed in a high-traffic area, and the BBU is placed in the central computer room. The BBU and the RRU can also be placed in the same computer room. The BBU and the RRU can also be different components under the same rack. Optionally, one cell may correspond to one carrier or member carrier.

[0082] In some deployments, the network device mentioned in the embodiments of this application may be a device including a CU, or a DU, or a device including a CU and a DU, or a control plane CU node (centralized unit-control plane, CU-CP) and a user plane CU node (centralized unit-user plane, CU-UP) and a DU node. For example, the network device may include a gNB-CU-CP, a gNB-CU-UP, and a gNB-DU.

[0083] In some deployments, multiple RAN nodes cooperate to assist the terminal in achieving wireless access, and different RAN nodes respectively implement some functions of the base station. For example, the RAN node may be a CU, a DU, a CU-CP, a CU-UP, or an RU, etc. The CU and the DU may be set separately or may also be included in the same network element, such as in the BBU. The RU may be included in the radio frequency device or radio frequency unit, such as included in the RRU, AAU, or RRH.

[0084] The RAN node can support one or more types of fronthaul interfaces. Different fronthaul interfaces respectively correspond to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is the Common Public Radio Interface (CPRI), the DU is configured to implement one or more of the baseband functions, and the RU is configured to implement one or more of the radio frequency functions. If the fronthaul interface between the DU and the RU is another interface, compared with the CPRI, some of the downlink and / or uplink baseband functions, for example, for the downlink, one or more of precoding, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / cyclic prefix (CP) addition, are moved from the DU to the RU for implementation. For the uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / cyclic prefix removal, are moved from the DU to the RU for implementation. In a possible implementation, this interface can be the Enhanced Common Public Radio Interface (eCPRI). Under the eCPRI architecture, different splitting methods between the DU and the RU correspond to different categories (Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, F.

[0085] Taking eCPRI Cat A as an example, for downlink transmission, with layer mapping as the segmentation, the DU is configured to implement one or more functions before layer mapping (i.e., one or more of encoding, rate matching, scrambling, modulation, layer mapping), while other functions after layer mapping (e.g., one or more of resource element (RE) mapping, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / adding cyclic prefix (CP)) are moved to the RU for implementation. For uplink transmission, with de-RE mapping as the segmentation, the DU is configured to implement one or more functions before demapping (i.e., one or more of decoding, derate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, de-RE mapping), while other functions after demapping (e.g., one or more of digital BF or FFT / removing CP) are moved to the RU for implementation. It can be understood that for the function descriptions of the DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol and will not be elaborated here.

[0086] In a possible design, the processing unit in the BBU for implementing baseband functions is called the baseband high (BBH) unit, and the processing unit in the RRU / AAU / RRH for implementing baseband functions is called the baseband low (BBL) unit.

[0087] In different systems, the CU (or CU-CP and CU-UP), DU, or RU may also have different names, but those skilled in the art can understand their meanings. For example, in an open radio access network (ORAN) system, the CU can also be called O-CU (open CU), the DU can also be called O-DU, the CU-CP can also be called O-CU-CP, the CU-UP can also be called O-CU-UP, and the RU can also be called O-RU. Any of the CU (or CU-CP, CU-UP), DU, and RU units in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.

[0088] In the embodiments of the present application, the device for implementing the functions of a network device may be the network device; it may also be a device capable of supporting the network device to implement such functions, such as a chip system, a hardware circuit, a software module, or a combination of a hardware circuit and a software module. This device may be installed in the network device or used in combination with the network device. In the embodiments of the present application, only the case where the device for implementing the functions of the network device is the network device is taken as an example for illustration, which does not limit the solutions of the embodiments of the present application.

[0089] It can be understood that the present application can be applied between a network device and a terminal device.

[0090] Protocol layer structure between a network device and a terminal device:

[0091] The communication between a network device and a terminal device follows a certain protocol layer structure. This protocol layer structure may include a control plane protocol layer structure and a user plane protocol layer structure. For example, the control plane protocol layer structure may include functions of protocol layers such as a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, a radio link control (RLC) layer, a medium access control (MAC) layer, and a physical layer. For example, the user plane protocol layer structure may include functions of protocol layers such as the PDCP layer, the RLC layer, the MAC layer, and the physical layer. In a possible implementation, a service data adaptation protocol (SDAP) layer may also be included above the PDCP layer.

[0092] Optionally, the protocol layer structure between a network device and a terminal device may also include an artificial intelligence (AI) layer for transmitting data related to AI functions.

[0093] Taking the data transmission between a network device and a terminal device as an example, the data transmission needs to go through the user plane protocol layer, such as the SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer. Among them, the SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer can also be collectively referred to as the access layer. According to the data transmission direction, it is divided into transmission or reception, and each of the above layers is further divided into a transmission part and a reception part. Taking the downlink data transmission as an example, after the PDCP layer obtains data from the upper layer, it transmits the data to the RLC layer and the MAC layer. Then, the MAC layer generates a transport block, and then performs wireless transmission through the physical layer. The data is encapsulated correspondingly in each layer. For example, the data received by a certain layer from the upper layer of this layer is regarded as the service data unit (SDU) of this layer. After being encapsulated by this layer, it becomes a protocol data unit (PDU), and then is passed to the next layer.

[0094] Exemplarily, the terminal device may also have an application layer and a non-access layer. Among them, the application layer can be used to provide services for the application programs installed in the terminal device. For example, the downlink data received by the terminal device can be sequentially transmitted from the physical layer to the application layer, and then provided to the application program by the application layer; for another example, the application layer can obtain the data generated by the application program and sequentially transmit the data to the physical layer to be sent to other communication devices. The non-access layer can be used to forward user data, such as forwarding the uplink data received from the application layer to the SDAP layer, or forwarding the downlink data received from the SDAP layer to the application layer.

[0095] It should be understood that Figure 2 The number and type of each device in the shown communication system are only for illustration, and this application is not limited thereto. In actual applications, the communication system may also include more terminal devices, more network devices, and may also include other network elements, such as core network devices, and / or network elements for implementing artificial intelligence functions.

[0096] It can be understood that all or part of the functions implemented by one or more of the terminal device, network device, core network device, or network element for implementing artificial intelligence functions can be virtualized, that is, implemented by one or more of a dedicated processor or a general-purpose processor and corresponding software modules. Among them, since the terminal device and the network device involve the interface for air interface transmission, the transceiver functions of this interface can be implemented by hardware. Core network devices, such as operation administration and maintenance (OAM) network elements, can all be virtualized. Optionally, one or more functions of the virtualized terminal device, network device, core network device, or network element for implementing artificial intelligence functions can be implemented by a cloud device, such as a cloud device in an over the top (OTT) system.

[0097] In the embodiments of the present application, when the central node is a network device and the sub-node is a terminal device (such as a UE), the network device and UE1 to UE5 can form a distributed AI training system as shown in Figure 3 In this communication system, UE1 to UE5 can send data to the network device. The network device needs to receive the uplink data sent by UE1 to UE5. The uplink data can be the representation difference or model parameters calculated by the sub-node, or it can be a feedback quantity containing its status information. At the same time, the network device can send configuration information to UE1 - UE5. The configuration information can be the model parameter data used by the central node to synchronize each sub-node, or it can be control data indicating the training method of the sub-node. The data between the network device and the UE can be carried on a physical channel, such as a physical downlink control channel (PDCCH), a physical downlink shared channel (PDSCH), a physical uplink shared channel (PUSCH), or a physical uplink control channel (PUCCH), etc.; for another example, a physical sidelink control channel (PSCCH), a physical sidelink shared channel (PSSCH).

[0098] In order to support AI technology in a wireless network, an AI node may also be introduced into the network.

[0099] Optionally, the AI node can be deployed at one or more of the following positions in the communication system: network devices, terminal devices, core network devices, etc. Alternatively, the AI node can also be deployed independently. For example, it can be deployed at a position outside any of the above devices, such as in the host of an over the top (OTT) system or a cloud server. The AI node can communicate with other devices in the communication system, and the other devices can be, for example, one or more of the following: network devices, terminal devices, or network elements of the core network, etc.

[0100] It can be understood that the number of AI nodes in this application is not limited. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on functions. For example, different AI nodes are responsible for different functions.

[0101] It can also be understood that the AI node can be an independent device, or can be integrated into the same device to implement different functions, or can be a network element in a hardware device, or can be a software function running on dedicated hardware, or can be a virtualized function instantiated on a platform (such as a cloud platform). This application does not limit the specific form of the above AI node.

[0102] The AI node can be an AI network element or an AI module.

[0103] One or more AI modules are provided in one or more of these network element nodes, such as core network devices, access network nodes (RAN nodes), terminals, or OAM. The access network node can be a separate RAN node or can include multiple RAN nodes. For example, it includes a CU and a DU. One or more AI modules can also be provided in the CU and / or the DU. Optionally, the CU can also be split into a CU-CP and a CU-UP. One or more AI models are provided in the CU-CP and / or the CU-UP.

[0104] The AI module is used to implement the corresponding AI function. The AI modules deployed in different network elements can be the same or different. The model of the AI module can be configured according to different parameters, and the AI module can implement different functions. The model of the AI module can be configured based on one or more of the following parameters: structural parameters (such as at least one of the number of neural network layers, the width of the neural network, the connection relationship between layers, the weights of neurons, the activation function of neurons, or the bias in the activation function), input parameters (such as the type and / or dimension of the input parameters), or output parameters (such as the type and / or dimension of the output parameters). Among them, the bias in the activation function can also be referred to as the bias of the neural network.

[0105] An AI module can have one or more models. A model can infer an output that includes one or more parameters. The learning process, training process, or inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.

[0106] A radio access network (RAN) intelligent controller (RIC) is included in the communication system. For example, the RIC can be the above-mentioned AI module for implementing AI-related functions. The RIC includes a near-real time RIC (near-RT RIC) and a non-real time RIC (Non-RT RIC). Among them, the non-real time RIC mainly processes non-real time information, such as data that is not sensitive to latency, and the latency of this data can be in seconds. The real time RIC mainly processes near-real time information, such as data that is relatively sensitive to latency, and the latency of this data is in tens of milliseconds.

[0107] The near-real time RIC is used for model training and inference. For example, it is used to train an AI model and perform inference using this AI model. The near-real time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data. Optionally, the near-real time RIC can deliver the inference result to the RAN node and / or the terminal. Optionally, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the near-real time RIC delivers the inference result to the DU, and the DU sends it to the RU.

[0108] The non-real time RIC is also used for model training and inference. For example, it is used to train an AI model and perform inference using this model. The non-real time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data, and the inference result can be delivered to the RAN node and / or the terminal. Optionally, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the non-real time RIC delivers the inference result to the DU, and the DU sends it to the RU.

[0109] The near-real time RIC and the non-real time RIC can also be separately set as a network element. Optionally, the near-real time RIC and the non-real time RIC can also be part of other devices. For example, the near-real time RIC is set in a RAN node (such as in the CU or DU), while the non-real time RIC is set in the OAM, cloud server, core network device, or other network devices.

[0110] Exemplarily, the settings of the near-real-time RIC and non-real-time RIC in the network architecture can be as Figures 4A to 4D shown below:

[0111] As Figure 4A shown in (a) of [], in the first possible implementation, the network device includes a near-real-time RIC module for model learning and / or inference.

[0112] As Figure 4A shown in (b) of [], in the second possible implementation, in the communication system, a non-real-time RIC can be included outside the network device. Optionally, the non-real-time RIC can be located in the OAM or the core network device.

[0113] As Figure 4A shown in (c) of [], in the third possible implementation, the network device includes a near-real-time RIC, and a non-real-time RIC is also included outside the network device. Optionally, the non-real-time RIC can be located in the OAM or the core network device.

[0114] Relative to Figure 4A in (c), Figure 4B in [], the CU is separated into CU-CP and CU-UP. The settings of the near-real-time RIC and non-real-time RIC are the same as those in Figure 4A in (c).

[0115] As Figure 4C shown, optionally, the network device includes one or more AI entities, and the functions of the AI entities are similar to those of the above-mentioned near-real-time RIC. Optionally, the OAM includes one or more AI entities, and the functions of the AI entities are similar to those of the above-mentioned non-real-time RIC. Optionally, the core network device includes one or more AI entities, and the functions of the AI entities are similar to those of the above-mentioned non-real-time RIC. When both the OAM and the core network device include AI entities, the models trained by their respective AI entities are different, and / or the models used for inference are different. In this application, the difference in models can include at least one of the following: the structural parameters of the model (such as the number of layers of the model, and / or weights, etc.), the input parameters of the model, or the output parameters of the model.

[0116] Relative to Figure 4C , Figure 4D in [], the network device is separated into CU and DU. Optionally, the CU can include an AI entity, and the function of the AI entity is similar to that of the above-mentioned near-real-time RIC. Optionally, the DU can include an AI entity, and the function of the AI entity is similar to that of the above-mentioned near-real-time RIC. When both the CU and the DU include AI entities, the models trained by their respective AI entities are different, and / or the models used for inference are different. Optionally, [] can be further Figure 4DThe CU is split into a CU-CP and a CU-UP. Optionally, one or more AI models may be deployed in the CU-CP. And / or, one or more AI models may be deployed in the CU-UP. Optionally, Figure 4C or Figure 4D in, the OAM of the network device and the OAM of the core network device can be deployed separately and independently.

[0117] For ease of understanding, the AI technology involved in this application will be introduced first below. It can be understood that this introduction is not a limitation on this application.

[0118] (1) AI model

[0119] AI refers to the intelligence exhibited by machines made by humans. Generally, artificial intelligence refers to the technology of presenting human intelligence through ordinary computer programs. Artificial intelligence can be defined as a machine or computer that imitates humans and has cognitive functions related to human thinking, such as learning and problem-solving. Artificial intelligence can learn from past experiences, make reasonable decisions, and respond quickly. The goal of artificial intelligence is to understand intelligence by constructing computer programs with symbolic reasoning or inference.

[0120] Machine learning (ML) is a way to achieve artificial intelligence, that is, to use machine learning as a means to solve problems in artificial intelligence. Machine learning theory mainly designs and analyzes some algorithms that allow computers to automatically "learn". Machine learning algorithms are a class of algorithms that automatically analyze and obtain rules from data and use the rules to predict unknown data. Because a large amount of statistical theory is involved in the learning algorithms, machine learning is particularly closely related to inferential statistics and is also known as statistical learning theory.

[0121] Machine learning can be divided into supervised learning, unsupervised learning, and reinforcement learning.

[0122] Supervised learning is based on the collected sample values and sample labels, uses machine learning algorithms to learn the mapping relationship from sample values to sample labels, and uses a machine learning model to express the learned mapping relationship. The process of training a machine learning model is the process of learning this mapping relationship. For example, in signal detection, the received signal with noise is the sample, and the true constellation point corresponding to this signal is the label. Machine learning expects to learn the mapping relationship between the sample and the label through training, that is, to make the machine learning model learn a signal detector. During training, the model parameters are optimized by calculating the error between the predicted value of the model and the true label. Once the mapping relationship is learned, the learned mapping can be used to predict the label of each new sample. The mapping relationship learned by supervised learning can include linear mapping and non-linear mapping. According to the type of label, the learning tasks can be divided into classification tasks and regression tasks.

[0123] Unsupervised learning only relies on the collected sample values and uses algorithms to discover the internal patterns of the samples by itself. In unsupervised learning, there is a type of algorithm that uses the samples themselves as the supervision signal, that is, the model learns the mapping relationship from samples to samples, which is called self-supervised learning. During training, the model parameters are optimized by calculating the error between the predicted value of the model and the sample itself. Self-supervised learning can be used in applications such as signal compression and decompression recovery. Common algorithms include autoencoders and generative adversarial networks, etc.

[0124] Reinforcement learning is different from supervised learning. It is a type of algorithm that learns strategies to solve problems by interacting with the environment. Different from supervised and unsupervised learning, there is no clear "correct" action label data in reinforcement learning problems. The algorithm needs to interact with the environment to obtain the reward signal feedback from the environment, and then adjust the decision-making actions to obtain a larger reward signal value. For example, in downlink power control, the reinforcement learning model adjusts the downlink transmission power of each user according to the total system throughput rate feedback by the wireless network, and then expects to obtain a higher system throughput rate. The goal of reinforcement learning is also to learn the mapping relationship between the environmental state and the optimal decision-making action. However, because the "correct action" label cannot be obtained in advance, the network cannot be optimized by calculating the error between the action and the "correct action". The training of reinforcement learning is achieved through iterative interaction with the environment.

[0125] An AI model is an algorithm or computer program that can implement AI functions. It is the specific implementation of the AI technology function. The AI model represents the mapping relationship between the input and output of the model. The types of AI models can be neural networks, linear regression models, decision tree models, support vector machines (SVM), Bayesian networks, Q-learning models, or other machine learning models.

[0126] (2) Deep neural network (DNN)

[0127] The deep neural network is a specific implementation form of AI or machine learning technology. According to the universal approximation theorem, neural networks can theoretically approximate any continuous function, thus enabling neural networks to have the ability to learn any mapping. Traditional communication systems need to rely on rich expert knowledge to design communication modules, while deep learning communication systems based on DNN can automatically discover implicit pattern structures from large datasets, establish the mapping relationship between data, and obtain performance superior to traditional modeling methods.

[0128] The idea of DNN comes from the neuron structure of the brain tissue. For example, each neuron performs a weighted summation operation on its input value and outputs the operation result through an activation function. Such as Figure 5As shown, it is a schematic diagram of a neuron structure. Assume the input of the neuron is \(x = [x_0, x_1, \ldots, x n \), and the weights corresponding to each input are \(w = [w_0, w_1, \ldots, w n \), where \(w i \) is the weight of \(x i \) and is used to weight \(x i \). The bias for weighted summation of the input values according to the weights is, for example, \(b\). The form of the activation function can be various. Assume the activation function of a neuron is: \(y = f(z)=\max(0, z)\), then the output of this neuron is:

[0129] For another example, the activation function of a neuron is: \(y = f(z)=z\), then the output of this neuron is:

[0130] where \(b\), \(w i \), \(x i \) can be various possible values such as decimals, integers (such as 0, positive integers or negative integers), or complex numbers. The activation functions of different neurons in a neural network can be the same or different.

[0131] A neural network generally includes multiple layers, and each layer can include one or more neurons. By increasing the depth and / or width of the neural network, the expressive ability of the neural network can be improved, providing a more powerful information extraction and abstract modeling ability for complex systems. Among them, the depth of the neural network can refer to the number of layers included in the neural network, and the number of neurons included in each layer can be called the width of this layer. In one implementation, the neural network includes an input layer and an output layer. The input layer of the neural network processes the received input information through neurons and passes the processing result to the output layer, and the output layer obtains the output result of the neural network. In another implementation, the neural network includes an input layer, a hidden layer, and an output layer, and reference can be made to the schematic diagram of the neural network in Figure 6 . The input layer of the neural network processes the received input information through neurons and passes the processing result to the intermediate hidden layer. The hidden layer calculates the received processing result to obtain a calculation result, and the hidden layer passes the calculation result to the output layer or an adjacent hidden layer, and finally the output layer obtains the output result of the neural network. Among them, a neural network can include one hidden layer, or include multiple sequentially connected hidden layers, without limitation.

[0132] Depending on the construction method of the network, DNN can include feedforward neural network (FNN), convolutional neural networks (CNN), and recurrent neural network (RNN). Figure 6 Shown is an FNN network, which is characterized by the fact that neurons in adjacent layers are fully connected pairwise, which usually requires a large amount of storage space and leads to a high computational complexity for FNN.

[0133] CNN is a neural network specifically designed to process data with a similar grid structure. For example, time series data and image data can both be considered data with a similar grid structure. Instead of using all the input information for calculation at once, CNN uses a window of a fixed size to intercept part of the information for convolution operation, which greatly reduces the computational amount of model parameters. Additionally, depending on the type of information intercepted by the window (such as people and objects in the same picture being different types of information), each window can use different convolution kernels for operation, which enables CNN to better extract the features of the input data.

[0134] RNN is a type of DNN network that utilizes feedback time series information. Its input includes the new input value at the current moment and its own output value at the previous moment. RNN is suitable for obtaining sequence features that are relevant in time and is particularly applicable to applications such as speech recognition and channel coding and decoding.

[0135] The above FNN, CNN, and RNN are common neural network structures, and these network structures are all constructed based on neurons. As introduced above, each neuron performs a weighted sum operation on its input value, and the weighted sum result passes through a non - linear function to generate an output. Then, we call the weights of the weighted sum operation of neurons in the neural network and the non - linear function the parameters of the neural network. Taking the neuron with the non - linear function max{0,x} as an example, for the neuron performing the n operation, the parameters are the weights w = [w0,…,w , the bias of the weighted sum is b, and the non - linear function max{0,x}. The parameters of all neurons in a neural network constitute the parameters of this neural network.

[0136] (3) Training dataset and inference data

[0137] The training dataset is used for training the AI model. The training dataset can include the input of the AI model, or include the input and target output of the AI model. Among them, the training dataset includes one or more training data, and the training data can be the training samples input to the AI model or the target output of the AI model. Among them, the target output can also be referred to as a label or a label sample. The training dataset is one of the important parts of machine learning. The model training essentially learns certain features from the training data so that the output of the AI model is as close as possible to the target output, such as the difference between the output of the AI model and the target output is as small as possible. The composition and selection of the training dataset can, to a certain extent, determine the performance of the trained AI model.

[0138] In addition, during the training process of an AI model (such as a neural network), a loss function can be defined. The loss function describes the gap or difference between the output value of the AI model and the target output value. The present application does not limit the specific form of the loss function. The training process of the AI model is a process of adjusting the model parameters of the AI model so that the value of the loss function is less than a threshold or the value of the loss function meets the target requirements. For example, if the AI model is a neural network, adjusting the model parameters of the neural network includes adjusting at least one of the following parameters: the number of layers of the neural network, the width, the weights of the neurons, or the parameters in the activation function of the neurons.

[0139] The inference data can be used as the input of the trained AI model for inference of the AI model. During the model inference process, inputting the inference data into the AI model can obtain the corresponding output, which is the inference result.

[0140] (4) Design of the AI model

[0141] The design of the AI model mainly includes a data collection link (such as collecting training data and / or inference data), a model training link, and a model inference link. Further, it can also include an inference result application link. See Figure 7, which illustrates an AI application framework. In the aforementioned data collection phase, the data source is used to provide training datasets and inference data. In the model training phase, an AI model is obtained by analyzing or training the training data provided by the data source. Among them, the AI model represents the mapping relationship between the input and output of the model. Obtaining the AI model through the model training node is equivalent to learning the mapping relationship between the input and output of the model using the training data. In the model inference phase, the AI model trained in the model training phase is used to perform inference based on the inference data provided by the data source, and an inference result is obtained. This phase can also be understood as: inputting the inference data into the AI model and obtaining the output through the AI model, and this output is the inference result. The inference result can indicate: configuration parameters used (executed) by the execution object, and / or operations performed by the execution object. In the inference result application phase, the inference result is published. For example, the inference result can be uniformly planned by the execution entity (actor). For example, the execution entity can send the inference result to one or more execution objects (such as core network devices, network devices, or terminal devices, etc.) for execution. Another example is that the execution entity can also feedback the performance of the model to the data source to facilitate subsequent implementation of model update training.

[0142] It can be understood that a network element with artificial intelligence capabilities may be included in a communication system. The above-mentioned steps related to AI model design can be executed by one or more network elements with artificial intelligence capabilities. In one possible design, AI capabilities (such as an AI module or AI entity) can be configured within an existing network element in the communication system to implement AI-related operations, such as training and / or inference of an AI model. For example, the existing network element can be a network device (such as a gNB), a terminal device, a core network device, or a network management device, etc. Among them, the network management device can divide the network management work into three categories according to the actual needs of the operator network operation: operation, administration, and maintenance. The network management device can also be called an OAM network element, abbreviated as OAM. Operation mainly completes the analysis, prediction, planning, and configuration work for daily networks and services; maintenance mainly conducts daily operation activities such as testing and fault management of the network and its services. The network management device can detect the network operation status, optimize network connections and performance, improve network operation stability, and reduce network maintenance costs. Or in another possible design, an independent network element can also be introduced in the communication system to execute AI-related operations, such as training an AI model. This independent network element can be called an AI network element or an AI node, etc., and this application does not limit this name. This AI network element can be directly connected to the network devices in the communication system, or can be indirectly connected to the network devices through a third-party device. Among them, the third-party device can be a core network network element such as an authentication management function (AMF) network element or a user plane function (UPF) network element, OAM, a cloud server, or other network elements, without limitation. Exemplarily, referring to Figure 8 , this communication system includes network device 810, terminal devices 820, 830, and also introduces an AI network element 840 into this communication system.

[0143] In this application, a model can infer one parameter, or infer multiple parameters. The training processes of different models can be deployed in different devices or nodes, or can also be deployed in the same device or node. The inference processes of different models can be deployed in different devices or nodes, or can also be deployed in the same device or node.

[0144] Among them, the model parameters may include one or more of the following: structural parameters of the model (such as the number of layers of the model, and / or weights, etc.), input parameters of the model (such as input dimension, number of input ports), or output parameters of the model (such as output dimension, number of output ports). It can be understood that the input dimension may refer to the size of an input data. For example, when the input data is a sequence, the input dimension corresponding to the sequence may indicate the length of the sequence. The number of input ports may refer to the number of input data. Similarly, the output dimension may refer to the size of an output data. For example, when the output data is a sequence, the output dimension corresponding to the sequence may indicate the length of the sequence. The number of output ports may refer to the number of output data.

[0145] (5) Centralized training

[0146] In the past decade, the number of smart devices such as mobile terminals and wearable devices has been increasing continuously. It can be foreseen that in the near future, billions of Internet of Things devices will be deployed throughout the communication network to achieve the automation and intelligence of social operations. The performance of current intelligent services based on advanced machine learning may benefit from the explosively growing data on these devices and the computing power available on these UEs themselves. Most machine learning techniques, for example, learning algorithms based on deep neural networks, require all available data to be centralized for training. However, since centralized training requires collecting enough data, which often comes from UEs, UEs need to upload this data, and the upload overhead of this data is relatively large. Moreover, with a huge amount of training data concentrated in one node, the training efficiency will be limited by the storage space. In addition, collecting data at the UE side may violate user privacy. Storing a huge amount of data for centralized training may trigger major concerns among users about the leakage of private data. On the other hand, if users only use the data they own to train machine learning algorithms, the training effect is often limited by the limited amount of local data.

[0147] (6) Distributed training

[0148] Distributed training is an effective solution to address the above challenges. This type of technology allows the machine training process to be divided among multiple sub-nodes on the user side, achieving the scalability of learning algorithms. It allows a cloud or server acting as a central node to collect machine learning models trained by multiple sub-nodes, and the central node improves the overall machine learning training effect by integrating the models trained by sub-nodes. Since the training data always remains at the sub-nodes, distributed learning technology is expected to achieve the same performance as centralized training while protecting user data privacy and using the data and / or computing power of UEs.

[0149] To achieve this goal, the field of machine learning integrates computing and communication technologies and designs a simple distributed training framework called the parameter server. The parameter server mainly consists of two parts: a central node and child nodes. Among them, the central node is responsible for storing and / or updating parameters, while the child nodes are responsible for training. Briefly speaking, the basic idea of parameter server training is to introduce multiple child nodes for simultaneous training. The child nodes use the central node as a medium for parameter exchange between child nodes to synchronize the model parameters of child nodes. Each child node first receives the training data distributed by the central node and / or the broadcast model parameters, then uses the received data to execute a stochastic gradient algorithm to calculate the gradient, and uploads the gradient to the central node. The central node receives the gradients uploaded by the child nodes, aggregates the gradients to obtain new model parameters, and broadcasts them to the child nodes again.

[0150] However, the above training process will face the following two problems:

[0151] 1. Each time a child node calculates one or more gradients, it needs to communicate with the central node. Since machine learning algorithms, such as deep learning algorithms, require a large number of gradient updates to converge, the above distributed training framework has extremely high communication overhead.

[0152] 2. The data of the child nodes is distributed by the central node, that is, the central node can coordinate the training data distribution of the child nodes to ensure that the data distribution between the child nodes is the same. In machine learning, independent and identically distributed training data is very important for ensuring an unbiased estimate of the gradient of the stochastic gradient algorithm. However, in actual scenarios, the child nodes may be user terminals, and the data of the user terminals cannot be accessed by the central node, that is, the central node loses control of the data distribution of the child nodes. This will lead to large differences between child nodes and damage the overall performance.

[0153] To address the above problems, one approach is to adopt the federated learning algorithm. Federated learning allows child nodes to use local datasets to perform multiple steps of gradient updates and then feedback the updated models to the central node to reduce the communication overhead caused by the need for each child node to interact with the central node for each gradient update. Specifically, in the federated learning framework, K child nodes and a central server node cooperate to train a globally shared model. The server aggregates the models sent by the child nodes, and one or more child node models are trained for multiple rounds using local private data. Here, one round means traversing all available local data. Assuming that the local dataset contains N training samples and b samples are randomly selected each time for training to execute the stochastic gradient descent algorithm, then one round calculates times of gradients, where Denote floor function. Then, the central node performs model aggregation. After the aggregation is completed, the central node shares the aggregated model with the child nodes, and then repeats the above steps until the model converges. Since all local models are trained based on locally stored data, the data privacy of the child nodes can be protected. The entire process of this framework includes local training, communication between child nodes and the central node, and model aggregation. In the local training phase, the child nodes independently and parallelly train their own models. The models used by the child nodes for local training can be any type of machine learning model. Currently, since deep neural networks have been proven to have powerful capabilities in various fields, usually the local models of the child nodes are DNNs for optimal performance. In the communication phase, the data transmitted during each communication between the child nodes and the central node is only the model trained by the child nodes or the model aggregated by the central node. To improve efficiency, federated learning usually selects a part of the child nodes to perform model upload. The central node aggregates the received models within a set time window. For aggregation, model averaging is usually adopted in federated learning, that is, the central node performs weighted averaging on the child node models received. The above training, communication, and aggregation steps will be executed for multiple iterations until the model of the central node converges.

[0154] However, although federated learning can protect user privacy and reduce the communication overhead of traditional parameter server technology, due to the central node's inability to access user data, federated learning has the following disadvantages:

[0155] 1. When the data distributions of the child nodes are non-independent and identically distributed, the performance cannot be guaranteed. In practice, the data amounts between child nodes may be unbalanced. Due to different user preferences, the data on user devices may be non-independent and identically distributed. The distribution of a single child node's dataset cannot represent the overall data distribution. At this time, the gradients obtained by the child nodes using the stochastic gradient descent algorithm may be biased, which will ultimately lead to poor performance in distributed training.

[0156] 2. When encountering an unstable feedback link, the performance fluctuates greatly. When the user child nodes feedback gradients through an unstable wireless transmission, the number of child node gradients successfully received by the server has a certain randomness, which will bring fluctuations to the model parameter updates of the central server.

[0157] The above problems will all lead to a large difference between the models of the child nodes after multiple local gradient updates and the model of the central node, ultimately resulting in a decline in distributed training performance.

[0158] Another method is the Gossip learning method. This learning framework manages a group of sub-computation nodes, where one or more sub-nodes have a machine learning model and perform an iterative two-step process: local gradient update and inter-node aggregation update. Specifically, one or more sub-nodes perform multiple gradient updates on the local machine learning model in the gradient update step. Then, the sub-nodes share their model parameters with another randomly selected sub-node in the aggregation update step. These steps are repeated until all sub-nodes converge to a consensus model. Similar to federated learning, this technique allows sub-nodes to perform multi-step gradient calculations before communicating, so there is no need for frequent communication. At the same time, it can achieve distributed learning without a fixed central node, so the Gossip learning method has stronger scalability.

[0159] Similar to the federated learning algorithm, the Gossip learning method still faces the problem of non-independent and identically distributed data. Moreover, during the process of interacting with each other's model parameters between nodes, the receiving node can also be regarded as the central node, and the sending node can be regarded as the sub-computation node. Therefore, the Gossip learning framework will also encounter the instability of feedback. These problems will lead to significant differences between the models of the computation nodes, and ultimately make it difficult to converge to a consensus model with good performance.

[0160] As can be seen from the above, to ensure that user privacy is not violated, distributed training requires the central node not to access the data of the sub-nodes. Existing federated learning technologies and Gossip learning technologies alleviate the problem of frequent communication in distributed training systems by allowing sub-nodes to perform multiple local gradient updates before communicating. However, when faced with the problems of non-independent and identically distributed data of sub-node data and unstable wireless feedback, the differences between sub-nodes and between sub-nodes and the central node are relatively large, ultimately resulting in the performance of the distributed system not meeting service requirements. Therefore, it is necessary to improve the performance of distributed training.

[0161] To address the above problems, this application provides a distributed training method. The sub-nodes participating in the training of the central node model feedback the representational differences between the models to the central node, and the central node gives training instructions for the model based on this representational difference, enabling a high-performance machine learning model to be obtained during a limited number of communications between sub-nodes and the central node, improving the performance of distributed training.

[0162] The distributed training method provided by the embodiments of the present application will be described in detail below. It can be understood that in the present application, the central node and the sub-nodes are used as examples of the execution entities of the interaction schematic, but the present application does not limit the execution entities of the interaction schematic. For example, the central node in the method provided by the present application can also be a chip, a chip system, a circuit or a processor applied to the central node, and can also be a logical node, a logical module or software that can implement all or part of the central node functions; the sub-nodes in the method provided by the present application can also be a chip, a chip system, a circuit or a processor applied to the sub-nodes, and can also be a logical node, a logical module or software that can implement all or part of the sub-node functions.

[0163] As Figure 9 shown, it is a schematic flowchart of a distributed training method provided by an embodiment of the present application. Exemplarily, the method may include the following steps:

[0164] S901. The first sub-node sends the first information to the central node. Correspondingly, the central node receives the first information.

[0165] As Figure 1 shown, the distributed training system may include a central node and multiple sub-nodes. The first sub-node in this embodiment is any one of the multiple sub-nodes participating in the model training of the central node. The interaction operations between the central node and the first sub-node, as well as the operations performed inside the first sub-node itself described in this embodiment, can all be applied to other sub-nodes.

[0166] To reduce the communication frequency between the sub-nodes and the central node, the sub-nodes have the ability to perform multiple steps of gradient calculation and / or model parameter update locally. When the data distribution between the sub-nodes is independently and identically distributed, the gradients obtained by the sub-nodes using the stochastic gradient descent (SGD) method are unbiased. At this time, the distributed training system can exhibit good performance. However, in practice, since the training data of the sub-nodes may come from users in different geographical locations, different time periods, and different preferences, the assumption of independent and identical distribution often does not apply to the actual system. This results in the model differences between the sub-nodes and the model differences between the sub-nodes and the central node increasing continuously as the training progresses, leading to the deterioration of the overall performance of the distributed training.

[0167] During distributed training, the central node sends the initialized model information to each sub-node participating in the model training. Each sub-node performs model training based on local data and performs local model updates.

[0168] In this embodiment, each time the model of the first child node is updated locally, the first child node can calculate the first representation difference between the model of the central node and the model of the first child node, or calculate the first representation difference between the models of all child nodes participating in the model training of the central node, and send the first information to the central node. Wherein, the first information is used to indicate the first representation difference between the model of the central node and the model of the first child node, or the first representation difference between the models of all child nodes. Exemplarily, the first child node can send the first information to the central node through wireless transmission.

[0169] For example, if the first representation difference is the Euclidean norm (abbreviation: "L2 norm") distance (abbreviation: "L2 distance") or the cosine distance, then the first information includes the L2 distance between the model of the central node and the model of the first child node, or includes the cosine distance between the model of the central node and the model of the first child node, or includes the L2 distance between the models of each child node, or includes the cosine distance between the models of each child node. At the beginning of model training, the first child node receives the initialization model information of the central node, and after the global model is updated, the first child node receives the updated model information of the central node. The first child node can calculate the L2 distance or cosine distance between the model of the central node and the model of the first child node based on the received model information of the central node and its own model information. In another possible implementation, the first child node can receive the model information of each child node, and calculate the L2 distance or cosine distance between the model of the first child node and the models of other child nodes based on the received model information of other child nodes and its own model information. Wherein, the other child nodes can be one or more child nodes.

[0170] Among them, the L2 distance is defined as:

[0171] L2(w o ,w k )=||w k -w o ||2, where w o is the parameter of the central node model, w k is the model parameter of the first child node, ||·||2 is the L2 distance operator, for the vector x=(x1,x2,…,x n ),

[0172] The cosine distance is defined as:

[0173] where w o is the parameter of the central node model, w k is the model parameter of the first child node, represents the transpose of w k .

[0174] It can be understood that it can also be that the first child node sends the initialized model information of the received central node to a third-party device. The third-party device performs model training based on the local data of the first child node, updates the model, calculates the first representation difference between the model of the central node and the model of the first child node, or calculates the first representation difference between the model of other child nodes participating in the model training of the central node and the model of the first child node, and sends the first representation difference to the first child node. The first child node then sends the first representation difference to the central node. Among them, the other child nodes can be one or more child nodes. In a possible implementation, the other child nodes include all child nodes participating in the model training of the central node.

[0175] S902. The central node sends second information to the first child node according to the first representation difference. Correspondingly, the first child node receives the second information.

[0176] After receiving the first representation differences sent by all child nodes participating in the model training of the central node, the central node performs difference value analysis, generates training instruction information for each child node, and sends second information to each child node. Among them, the second information includes the training instruction information of the models of each child node.

[0177] If the first representation difference between the model of the central node and the model of the first child node is too large, or the first representation difference between the model of the first child node and the models of other child nodes is too large, it indicates that the model trained by the first child node deviates far from the model of the central node, or the model trained by the first child node deviates far from the models trained by other child nodes. The central node issues training instruction information, which is used to instruct the first child node to stop training the model of the first child node. Thus, it can timely prevent the model difference between the first child node and the central node (or other child nodes) from continuously increasing as the training progresses, resulting in the deterioration of the overall performance of distributed training.

[0178] Furthermore, the central node can also allocate the parameters of the corresponding loss function to the child nodes with large difference values according to the difference value analysis result, so that the child nodes can update their models according to the parameters, making the impact of the subsequent updates of the child nodes' models on the training results smaller.

[0179] If the first representation difference between the model of the central node and the model of the first child node is small, or the first representation difference between the model of the first child node and the models of other child nodes is small, it indicates that the model trained by the first child node does not deviate far from the model of the central node, or the models trained by the first child node and other child nodes do not deviate far from each other. The central node sends training instruction information, which is used to instruct the first child node to continue training the model of the first child node.

[0180] According to a distributed training method provided by an embodiment of the present application, child nodes participating in the training of the central node model feedback the representation differences between the models to the central node, and the central node gives training instructions based on the representation differences, so that a high-performance machine learning model can be obtained during a limited number of communications between the child nodes and the central node, improving the performance of distributed training.

[0181] As Figure 10 shown, it is a schematic flowchart of another distributed training method provided by an embodiment of the present application. Exemplarily, the method may include the following steps:

[0182] S1001. The central node broadcasts the third information. Correspondingly, the first child node receives the third information.

[0183] As Figure 1 shown, the distributed training system may include a central node and multiple child nodes. The first child node in this embodiment is any one of the multiple child nodes participating in the training of the central node model. The interaction operations between the central node and the first child node described in this embodiment, as well as the operations performed inside the first child node itself, can all be applied to other child nodes.

[0184] The central node usually sends the same third information to all child nodes participating in the training of the central node model in a broadcast manner.

[0185] Among them, the third information includes at least one of the following: model configuration information, public dataset. During initial training, the central node usually sends the same model configuration information and / or public dataset to one or more child nodes participating in the training of the central node model in a broadcast manner. The central node itself also stores the same model configuration information. Optionally, one or more child nodes may include all child nodes participating in the training of the central node model.

[0186] Among them, the model configuration information is used to indicate at least one of the following: the type of the models of multiple child nodes, the structural information of the models of multiple child nodes, the model parameters of the models of multiple child nodes, or the training configuration information of the models of multiple child nodes. Optionally, the models of multiple child nodes are all the same, that is, the model of the central node.

[0187] For example, the types of models of child nodes include DNN, CNN, random forest models, etc.

[0188] For example, the structural information of the model of a child node includes the number of hidden layers of the DNN, the neuron data and activation functions of each layer or some layers.

[0189] For example, the types and structural information of the models of multiple child nodes sent down can be reflected in the form of configuration text, or can be code scripts for compiling corresponding machine learning models.

[0190] Among them, the model parameters of the model of the child node are generated by the central node through a certain strategy, including but not limited to random generation, pre-training generation or obtaining from other third-party entities.

[0191] Among them, the training configuration information of the model of the child node includes the optimizer used by the child node to perform gradient update (such as SGD, root mean square propagation (RMSprop), adaptive moment estimation (Adam), etc.), the loss function of the task (such as using the cross-entropy loss function for classification tasks and the mean square error loss function for regression tasks), the regularization penalty term (such as the L2 penalty term), the initial learning rate, the gradient update batch size, and the maximum number of training epochs of the child node, and so on. And it is agreed that the loss function used for local training of the child node is the loss function of the task plus the representation difference penalty term. For example, if the task loss is the classification cross-entropy CrossEntropy(), then for a batch of training samples The loss function used for local training of child node k is: CrossEntropy(D)+A*e k (D), where x i is the input of the model; y i is the label of the model; N is the number of training samples; i is the i-th training sample; A is the representation difference coefficient; e k is the representation difference.

[0192] S1002. The first child node calculates the first representation difference between the model of the central node and the model of the first child node based on the local dataset or public dataset of the first child node.

[0193] Exemplarily, every time the model of the first child node is updated locally, it can trigger each child node of the first child node to calculate the first representation difference based on its own local dataset or public dataset. Among them, the local dataset can be pre-stored in each child node at the time of factory shipment. The local datasets of each child node are the same.

[0194] As described above, the common dataset is sent by the central node to each sub-node. The common datasets received by each sub-node are the same.

[0195] By enabling each sub-node to calculate the first representation difference based on the same local dataset or common dataset, the representation differences between the models of each sub-node and the central node, or the representation differences between the models of each sub-node, can be more accurately compared.

[0196] Among them, the first representation difference refers to the difference between the outputs or intermediate quantities calculated by the machine learning models of each sub-node and / or the central node using the same input. That is, the first representation difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first sub-node, and / or the first representation difference is the difference between the outputs or intermediate quantities of the models of multiple sub-nodes, where the output or intermediate quantity is obtained based on the same input. This same input is the local dataset or common dataset of the above-mentioned sub-node.

[0197] As Figure 11 shown, it is a schematic diagram for calculating a representation difference according to an embodiment of the present application. The machine learning model is a neural network, and its structure includes a feature extractor, a classifier or a regressor. For this model, the calculation of its representation difference follows the following steps: Place the central node model in the verification mode (turn off the randomness of the modules with randomness in the model, for example, turn off the Dropout function in the DNN), randomly select a batch of training samples from the local dataset or common dataset, send the samples into the central node model and obtain the features z o extracted by the model; Place the local model of the sub-node in the training mode (turn on the randomness of the modules with randomness in the model, for example, turn on the Dropout function in the DNN), send the samples into the sub-node model and obtain the features z k extracted by the model; Calculate a certain metric dis(z o and z k )(such as the L2 distance or the cosine distance) between the features z o and z k , and this metric is the representation difference e k . It can be understood that Figure 11 in the example, the difference between the intermediate quantities of the model of the central node and the model of the first sub-node is used as the representation difference, and the representation difference can also be the difference between the output of the model of the central node and the model of the first sub-node.

[0198] It can be understood that Figure 11That is, the L2 distance or cosine distance between the output of the last layer of the feature extractor of the central node and the output of the last layer of the feature extractor of the child node is used as the representation difference. Additionally, the L2 distance or cosine distance between the output of any layer of the feature extractor of the central node and the output of any layer of the feature extractor of the child node can also be calculated as the representation difference. Or, the L2 distance or cosine distance between the output of any layer of the classifier or regressor of the central node and the output of any layer of the classifier or regressor of the child node can be calculated as the representation difference. This application does not limit this.

[0199] Based on Figure 11 the process shown, the representation difference between the model of the central node and the model of the child node can be calculated.

[0200] In another scenario, the model of the central node can also be replaced with the model of another child node (reference node model), and then the representation differences between each child node model and the reference node model can be calculated.

[0201] S1003. The first child node sends the first information to the central node. Correspondingly, the central node receives the first information.

[0202] After calculating the first representation difference, the first child node sends the first information to the central node. Exemplarily, the first child node can send the first information to the central node through wireless transmission. Wherein, the first information is used to indicate the first representation difference between the model of the central node and the model of the first child node, or is used to indicate the first representation difference between the models of each child node.

[0203] For example, if the first representation difference is the L2 distance or cosine distance, the first information includes the L2 distance between the model of the central node and the model of the first child node, or includes the cosine distance between the model of the central node and the model of the first child node, or includes the L2 distance between the models of each child node, or includes the cosine distance between the models of each child node.

[0204] S1004. The central node aggregates the first representation differences of each child node and generates the first training indication information R1 of each child node.

[0205] After receiving the first representation differences sent by all child nodes participating in the training of the central node model, the central node performs difference value analysis and generates the first training indication information R1 of each child node.

[0206] The central node can set a first condition and a second condition to determine whether the first characterization differences of each child node meet the first condition or the second condition. If the first characterization difference of a child node meets the first condition, the first training instruction information is used to instruct the first child node to continue training the model of the first child node; if the first characterization difference of a child node meets the second condition, the first training instruction information is used to instruct the first child node to stop training the model of the first child node. Exemplarily, the first characterization difference meeting the first condition includes: the first characterization difference is less than or equal to a first threshold. The central node can compare the first characterization differences of each child node with the first threshold. If the first characterization difference of any child node is less than or equal to the first threshold, the first training instruction information is used to instruct the child node to continue training the model of the child node. The central node can compare the first characterization differences of each child node with a second threshold. If the first characterization difference of any child node is greater than or equal to the second threshold, the first training instruction information is used to instruct the child node to stop training the model of the child node. By selecting a unified condition, the central node can determine whether the first characterization differences of each child node meet the above conditions, which can improve the efficiency and accuracy of the difference value analysis. Optionally, the above first condition, for example, the first threshold, is pre-determined by the central node or pre-defined by the protocol; the above second condition, for example, the second threshold, is pre-determined by the central node or pre-defined by the protocol. Exemplarily, the second threshold and the first threshold can be the same or different.

[0207] Furthermore, the central node can also allocate corresponding training function parameters, such as a characterization difference coefficient A, to the child nodes with large difference values according to the difference value analysis results. It can be understood that the larger the first characterization difference, the larger the characterization difference coefficient. Because the larger the first characterization difference indicates that the model trained by the child node deviates further from the model of the central node, or among all the child nodes participating in the model training of the central node, the model trained by this child node deviates further from the models trained by other child nodes. Therefore, the central node allocates a characterization difference coefficient A, so that the child node can update its loss function according to the characterization difference coefficient A, making the impact of the subsequent update of the child node's model on the training result smaller.

[0208] S1005. The central node sends second information to the first child node. Correspondingly, the first child node receives the second information.

[0209] Regarding this second information, there can be the following several implementations:

[0210] One implementation is that the second information includes the first training instruction information R1 of the model of the first child node.

[0211] According to the above analysis result of the difference value, if the first representation difference of the first child node satisfies the first condition, for example, the first representation difference of the first child node is less than or equal to the above first threshold, then the first training indication information is used to indicate that the first child node continues the training of the model of the first child node.

[0212] Furthermore, the first training indication information can also be used to indicate the representation difference coefficient in the loss function of the model of the first child node.

[0213] According to the above analysis result of the difference value, if the first representation difference of the first child node satisfies the second condition, for example, the first representation difference of the first child node is greater than or equal to the above second threshold, then the first training indication information is used to indicate that the first child node stops the training of the model of the first child node.

[0214] For example, the first training indication information can include one or several bits, where one bit is used to indicate whether the first child node continues or stops the training of the model of the first child node. For example, if the value of this bit is "1", it is used to indicate that the first child node continues the training of the model of the first child node; if the value of this bit is "0", it is used to indicate that the first child node stops the training of the model of the first child node. Vice versa.

[0215] If the value of this bit is "1", that is, it is used to indicate that the first child node continues the training of the model of the first child node, the remaining bits of the first training indication information can be used to indicate the representation difference coefficient in the loss function of the model of the first child node.

[0216] Another implementation is that the second information includes the first training indication information R1 of the model of the first child node, and the first training indication information is used to indicate whether the first child node continues or stops the training of the model of the first child node.

[0217] And if the first representation difference of the first child node satisfies the first condition, the second information further includes the representation difference coefficient in the loss function of the model of the first child node.

[0218] S1006. The first child node updates the model of the first child node according to the second information.

[0219] After receiving the second information, if the first training indication information is used to indicate that the first child node continues the training of the model of the first child node and is also used to indicate the representation difference coefficient in the loss function of the model of the first child node; or the second information further includes the representation difference coefficient in the loss function of the model of the first child node, then the first child node updates the model parameters one or more steps according to the second information, including updating the loss function of the model by using the representation difference coefficient.

[0220] This step assumes that the first training instruction information is used to instruct the first child node to continue training the model of the first child node. If the first training instruction information is used to instruct the first child node to stop training the model of the first child node, it will be described in detail later.

[0221] S1007. The first child node calculates a second representation difference between the model of the central node and the model of the first child node based on the local dataset or the public dataset of the first child node.

[0222] After the first child node updates the local model according to the second information, it can trigger again to calculate the second representation difference between the model of the central node and the model of the first child node based on the local dataset or the public dataset of the first child node. The description of this calculation process can refer to step S1002 and will not be repeated here. The second representation difference calculated by the first child node can be the same as or different from the above-mentioned first representation difference.

[0223] S1008. The first child node sends the fourth information to the central node. Correspondingly, the central node receives the fourth information.

[0224] After the first child node calculates the second representation difference, it sends the fourth information to the central node, where the fourth information is used to indicate the second representation difference between the model of the central node and the model of the first child node. The specific implementation of this step can refer to the description of step S1003 and will not be repeated here.

[0225] S1009. The central node aggregates the second representation differences of each child node to generate second training instruction information R2 for each child node.

[0226] The specific implementation of this step can refer to the description of step S1004 and will not be repeated here. It can be understood that the second training instruction information R2 of each child node generated by the central node can be the same as or different from the first training instruction information R1 of the corresponding child node.

[0227] S1010. The central node sends the fifth information to the first child node. Correspondingly, the first child node receives the fifth information.

[0228] After the central node generates the second training instruction information R2 of each child node, it sends the fifth information to the first child node. The fifth information includes the second training instruction information R2 of the model of the first child node.

[0229] Regarding this fifth information, there can be the following several implementations:

[0230] One implementation is that the fifth information includes the second training instruction information R2 of the model of the first child node.

[0231] According to the above analysis result of the difference value, if the second representation difference of the first child node is greater than or equal to the above threshold, the second training instruction information is used to instruct the first child node to stop the training of the model of the first child node.

[0232] Another implementation is that the fifth information includes the second training instruction information R2 of the model of the first child node, and the second training instruction information is used to instruct the first child node to stop the training of the model of the first child node.

[0233] And if the second representation difference of the first child node is less than or equal to the above threshold, the fifth information further includes the representation difference coefficient in the loss function of the model of the first child node.

[0234] S1011. The first child node stops training the model.

[0235] If the first representation difference of the first child node is greater than or equal to the above threshold, the first training instruction information is used to instruct the first child node to stop the training of the model of the first child node. After receiving the above fifth information, the first child node stops training the model.

[0236] Alternatively, if the first child node reaches the maximum number of training rounds indicated by the above training configuration information, it stops training the model.

[0237] S1012. The first child node sends the sixth information to the central node. Correspondingly, the central node receives the sixth information.

[0238] After the first child node stops training the model, it can send the model information obtained by training to the central node. Exemplarily, the first child node can send the sixth information to the central node. Among them, the sixth information includes the model information obtained by training the first child node.

[0239] S1013. The central node updates the model of the central node according to the sixth information.

[0240] After receiving the sixth information, the central node can update the model of the central node according to the sixth information.

[0241] After a child node stops training the model, when the central node receives the model information obtained by training this child node, it can update the model of the central node based on the model information obtained by training this child node; or it can also wait until it receives the model information obtained by training all the child nodes participating in the training of the central node model, and then uniformly update the model of the central node.

[0242] The ways for the central node to update the model include but are not limited to weighted averaging of model parameters, etc.

[0243] It can be understood that the above steps S1002 - S1006 and steps S1007 - S1013 can be implemented independently or jointly. The above steps S1002 - S1006 can also be executed once or multiple times until the maximum number of training rounds of the first child node is reached, or the training of the model of the first child node is instructed to stop by the central node due to the calculated representation difference being greater than or equal to the threshold.

[0244] S1014. The central node distributes the model of the new central node. Correspondingly, the first node receives the model of the new central node distributed by the central node.

[0245] After all child nodes participating in the training of the central node model reach the maximum number of training rounds, or the training of the model of the child node is instructed to stop by the central node due to the calculated representation difference being greater than or equal to the threshold, the central node receives the model information obtained from the training of all child nodes participating in the training of the central node model, updates the model of the central node using the model information of these child nodes, and broadcasts the new model configuration information of the central node to all participating child nodes again to start a new round of training. The above steps S1001 - S1014 can be repeatedly executed during the new round of training.

[0246] Optionally, the central node is the third - party device that performs the actions related to the central node. For example, the above steps S1001, S1003 - S1005, S1008 - S1010, S1012 - S1014 are all executed by the third - party device.

[0247] Optionally, the first child node is the third - party device that performs the actions related to the first child node. For example, the above steps S1001 - S1003, S1005 - S1008, S1010 - S1012, S1014 are all executed by the third - party device.

[0248] Optionally, the central node is a network device. In this case, the network device can complete the model training. For example, the above steps S1001, S1003 - S1005, S1008 - S1010, S1012 - S1014 are all executed by the network device.

[0249] Optionally, the first child node is a terminal device. In this case, the terminal device can complete the model training. For example, the above steps S1001 - S1003, S1005 - S1008, S1010 - S1012, S1014 are all executed by the terminal device.

[0250] Optionally, the central node includes a network device and a third-party device. In one example, one or more of the above steps S1004, S1009, and S1013 may be executed by a third-party device such as an OTT or a cloud server, etc., and one or more of the above steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 may be executed by the network device. In addition, the network device and the third-party device can also communicate with each other to transmit the content transmitted in one or more of the above steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014.

[0251] Optionally, the first sub-node includes a terminal device and a third-party device. In one example, one or more of the above steps S1002, S1006, S1007, and S1011 may also be executed by a third-party device such as an OTT or a cloud server, etc., and one or more of the above steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 may be executed by the terminal device. In addition, the terminal device and the third-party device can also communicate with each other to transmit the content transmitted in one or more of the above steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014.

[0252] According to a distributed training method provided by an embodiment of the present application, the sub-nodes participating in the training of the central node model feedback the representation differences between the models to the central node, and the central node gives training instructions for the model according to the representation differences, so that a high-performance machine learning model can be obtained in the process of communicating between the sub-nodes and the central node a limited number of times, improving the performance of distributed training.

[0253] By selecting appropriate physical quantities for representing the differences between the models of the distributed nodes (the differences between the outputs or intermediate quantities of the models of the central node and the first sub-node, or the differences between the outputs or intermediate quantities of the models of each sub-node, such as the L2 distance or the cosine distance), and giving training instructions for the models of the sub-nodes according to the representation differences, the model training of the distributed training system can be better coordinated, and the overall performance can be improved.

[0254] However, in the prior art, due to the lack of monitoring of the sub-node status, the differences between the sub-nodes and between the sub-node and the central node are relatively large, resulting in performance degradation.

[0255] By restricting the model representation differences in the embodiments of the present application, the divergence of the sub-node update directions can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the deviation caused by packet loss can be reduced, and the final performance of the central node model can be improved.

[0256] When there are significant differences in the representations in the embodiments of this application, the central node instructs the child nodes not to update the model, which can reduce unnecessary computational overhead of the child nodes, or the central node feeds back a large representation difference coefficient to the child nodes to constrain the representation difference during the training process of the child nodes from being too large. Thus, the performance of the child nodes during joint training with the central node can be improved. If a child node loses a packet when feedbacking model parameters, the models of the child node and the central node still have a high degree of synchronization.

[0257] As Figure 12 shown, it is a system block diagram of a distributed training example in the embodiments of this application. The distributed training system includes a central node and K child nodes, where K is a positive integer greater than or equal to 1. The model of the central node is f o (., w o ), and the model of child node k is the model f k (., w k ). Among them, k ∈ {1, 2,..., K}. The central node broadcasts model configuration information and a common data set. After each child node receives the information broadcast by the central node, it initializes its local model. Each child node calculates the representation difference e k between the model of the central node and the model of this child node based on its local data set or the common data set. Each child node respectively feeds back the representation difference e k to the central node. The central node aggregates the representation differences of each child node and generates training instruction information for each child node. The central node respectively sends the training instruction information to each child node. Each child node performs local training according to the received training instruction information. During the training process, each child node calculates the loss using local data, adds the representation difference coefficient as a penalty term to the loss function, then calculates the gradient of the model parameters, and performs local training. If a certain child node meets the training stop condition, it stops training and feeds back the trained model information to the central node. The central node receives the model information from the child node, updates its own parameter w o according to the received model, and feeds back a new model to one or more child nodes and re-executes the foregoing representation difference calculation and reporting process. Exemplarily, for example, if the task loss is categorical cross-entropy CrossEntropy(), then for a batch of training samples The loss function used for local training of child node k is: CrossEntropy(D) + A * e k (D), where x i is the input of the model; y i is the label of the model; N is the number of training samples; i is the i-th training sample; A is the representation difference coefficient; e kThe larger the first representation difference is, the larger the representation difference coefficient is. Then, updating the loss function according to the representation difference coefficient A can reduce the impact of the subsequent sub-node model update on the training result.

[0258] In the above process, the child node can feedback the model information to the central node after completing the multi-step gradient update, which can alleviate the communication overhead between the child node and the central node. In addition, each child node uses a local data set or a public data set to calculate the representation difference. By making each child node calculate the first representation difference based on the same local data set or public data set, the representation difference between the model of each child node and the model of the central node, or the representation difference between the models of each child node can be more accurately compared. By aggregating the representation differences of each child node and sending training instruction information to each child node in a targeted manner, the central node can ensure that the representation differences between the models of the child nodes and between the models of the child nodes and the model of the central node will not be too large, thereby improving the performance of distributed training.

[0259] The above embodiment describes that the sub-nodes feed back the representation differences to the central node, and the central node aggregates the representation differences of each sub-node and issues training instruction information. The following embodiment will describe that the sub-nodes themselves calculate the representation differences and decide to continue or stop the training of the model:

[0260] like Figure 13 FIG. 1 is a flow chart of another distributed training method provided in an embodiment of the present application. Exemplarily, the method may include the following steps:

[0261] S1301. The central node broadcasts first information. Correspondingly, the first subnode receives the first information.

[0262] The first information includes at least one of the following: model configuration information, a public data set, a threshold update rule, and a characterization difference coefficient update rule.

[0263] For details about model configuration information and public datasets, please refer to Figure 10 Relevant description in step S1001 in the illustrated embodiment.

[0264] Among them, the update rule of the threshold can be pre - formulated by the central node. For example, the update rule is formulated as follows: the threshold decreases as the number of communications between the first child node and the central node increases, and the threshold of the first child node remains unchanged before the communication between the first child node and the central node. Here, when the central node sends a new model to the first child node once, it is regarded as one communication. It can be understood that as the child nodes train and the central node aggregates the models of the child nodes, the models of both the central node and the child nodes gradually converge. Therefore, as the number of communications between the first child node and the central node increases, the representational difference between the model of the central node and the model of the first child node will gradually decrease. Therefore, a threshold update rule can be formulated such that the threshold decreases as the number of communications between the first child node and the central node increases.

[0265] Optionally, the update rule of the first information or the threshold may further include an initial threshold.

[0266] Among them, the update rule of the representational difference coefficient is used to indicate that after the child node calculates the representational difference, an updated representational difference coefficient is obtained based on the initial representational difference coefficient and the calculated representational difference. The update rule of the representational difference coefficient is pre - formulated by the central node. For example, the update rule of the representational difference coefficient is: taking the linear or non - linear scaling of the previously updated representational difference coefficient as the update amount of the representational difference coefficient compared to the previous representational difference coefficient. It can be understood that the greater the representational difference, the greater the representational difference coefficient. Because the greater the representational difference, it indicates that the model trained by this child node deviates further from the model of the central node, or among all the child nodes participating in the model training of the central node, the model trained by this child node deviates further from the models trained by other child nodes. Therefore, the child node obtains an updated representational difference coefficient, enabling the child node to update its loss function according to this representational difference coefficient, making the impact of the subsequent update of the child node's model on the training result smaller.

[0267] Optionally, the update rule of the first information or the representational difference coefficient may further include an initial representational difference coefficient.

[0268] S1302. The first child node calculates the first representational difference between the model of the central node and the model of the first child node based on the local data set or the common data set of the first child node.

[0269] For the specific implementation of this step, reference can be made to Figure 10 the relevant description in step S1002 in the embodiments shown.

[0270] S1303. When the first child node determines that the first representational difference is less than the first threshold, continue the training of the model of the first child node.

[0271] The first threshold may be pre-stored in the first child node, or sent by the central node to the first child node in step S1301, for example, an initial threshold carried in the first information or the threshold update rule.

[0272] The first subnode compares the calculated first representation difference with the first threshold. If the first representation difference is greater than the first threshold, the first subnode stops the training of the model; if the first representation difference is less than the first threshold, the first subnode continues the training of the model; if the first representation difference is equal to the first threshold, the first subnode continues or stops the training of the model.

[0273] This step assumes that the first representation difference is less than or equal to the threshold, and the first child node continues the model training. The following will describe in detail that if the first representation difference is greater than the threshold, the first child node stops the model training.

[0274] S1304. The first child node calculates a gradient using the local data set and a loss function with a coefficient representing the difference, and updates the model parameters.

[0275] The first child node uses the local data set to update the model parameters for one or more steps, wherein the gradient is calculated using a loss function with a coefficient representing the difference. The meaning of the coefficient representing the difference can be found in the description above.

[0276] The characterization difference coefficient may be obtained according to the characterization difference coefficient update rule issued by the central node in step S1301.

[0277] S1305. The first subnode updates the threshold according to the threshold updating rule, and updates the characterization difference coefficient according to the characterization difference coefficient updating rule.

[0278] The first subnode updates the first threshold according to the threshold update rule received in step S1301 to obtain the second threshold, and updates the representation difference coefficient according to the representation difference coefficient update rule received in step S1301.

[0279] This step is optional, and the threshold update rule / characterization difference coefficient update rule may not stipulate that the threshold and / or characterization difference coefficient is updated after each local model update of the first child node.

[0280] S1306. The first sub-node calculates a second representation difference between the model of the central node and the model of the first sub-node based on the local data set or the public data set of the first sub-node.

[0281] After the first child node updates the local model according to the second information, it can trigger the calculation of the second representation difference between the model of the central node and the model of the first child node based on the local dataset or the common dataset of the first child node again. The description of this calculation process can refer to step S1302 and will not be elaborated here. The second representation difference calculated by the first child node can be the same as or different from the above first representation difference.

[0282] S1307. The first child node determines that the second representation difference is greater than the second threshold or reaches the maximum number of training rounds, and then stops the training of the model of the first child node.

[0283] Optionally, when the first child node determines that the second representation difference is equal to the second threshold, it stops the training of the model of the first child node.

[0284] S1308. The first child node sends the second information to the central node. Correspondingly, the central node receives the second information.

[0285] After the first child node stops the model training, it can send the trained model information to the central node. Exemplarily, the first child node can send the second information to the central node. Among them, the second information includes the model information trained by the first child node.

[0286] S1309. The central node updates the model of the central node according to the second information.

[0287] After receiving the second information, the central node can update the model of the central node according to the second information.

[0288] After a child node stops the model training, when the central node receives the model information trained by this child node, it can update the model of the central node based on the model information trained by this child node; or it can also wait until it receives the model information trained by all child nodes participating in the training of the central node model, and then uniformly update the model of the central node.

[0289] The ways for the central node to update the model include but are not limited to weighted averaging of model parameters, etc.

[0290] It can be understood that the above steps S1302 - S1305 and steps S1306 - S1307 can be implemented independently or jointly. The above steps S1302 - S1305 can also be executed once or multiple times until the maximum number of training rounds of the first child node is reached, or the training of the model of the first child node is stopped because the calculated representation difference is greater than or equal to the threshold.

[0291] S1310. The central node distributes the new model of the central node. Correspondingly, the first node receives the new model of the central node distributed by the central node.

[0292] After all child nodes participating in the training of the central node model reach the maximum number of training rounds, or when the calculated representation difference is greater than or equal to the threshold and the training of the child node model is stopped, the central node receives the model information obtained from the training of all child nodes participating in the training of the central node model. The central node uses the model information of these child nodes to update its own model, and then broadcasts the new model configuration information of the central node to the child nodes again to start a new round of training. The above steps S1301 - S1309 can be repeatedly executed during the new round of training.

[0293] Optionally, the central node is a third - party device that performs the actions related to the central node. For example, the above steps S1301, S1308 - S1310 are all executed by the third - party device.

[0294] Optionally, the first child node is a third - party device that performs the actions related to the first child node. For example, the above steps S1301 - S1308, S1310 are all executed by the third - party device.

[0295] Optionally, the central node is a network device. For example, the above steps S1301, S1308 - S1310 are all executed by the network device.

[0296] Optionally, the first child node is a terminal device. For example, the above steps S1301 - S1308, S1310 are all executed by the terminal device.

[0297] Optionally, the central node includes a network device and a third - party device. In one example, the above step S1309 can be executed by a third - party device such as OTT or a cloud server, etc., and one or more of the above steps S1301, S1308, S1310 can be executed by the network device. In addition, the network device and the third - party device can also communicate to transmit the content transmitted in one or more of the above steps S1301, S1308, S1310.

[0298] Optionally, the first child node includes a terminal device and a third - party device. In one example, one or more of the above steps S1302 - S1307 can also be executed by a third - party device such as OTT or a cloud server, etc., and one or more of the above steps S1301, S1308, S1310 can be executed by the terminal device. In addition, the terminal device and the third - party device can also communicate to transmit the content transmitted in one or more of the above steps S1301, S1308, S1310.

[0299] According to a distributed training method provided by an embodiment of the present application, the central node sends model configuration information, a public data set, an update rule for a threshold, and an update rule for a representation difference coefficient to a child node. The child node can calculate the representation difference between the model of the central node and its own model, compare the representation difference with the threshold, and determine whether to continue or stop the training of the model. If the training of the model continues, the model parameters are updated using a loss function with a representation difference coefficient. By constraining the model representation difference, the dispersion degree of the update direction of the child node can be effectively reduced, the impact of non-independent and identically distributed data on the performance of the distributed training system and the deviation caused by packet loss can be alleviated, the final performance of the central node model can be improved, and the performance of distributed training can be enhanced.

[0300] In the present application, "sending information to... (such as a child node)" or the relevant schematic in the drawings can be understood as the destination of the information being the child node. It may include directly or indirectly sending information to the child node. "Receiving information from... (such as a child node)" or "receiving information from... (such as a child node)", or the relevant schematic in the drawings can be understood as the source of the information being the child node, and it may include directly or indirectly receiving information from the child node. Necessary processing may be performed on the information between the source and the destination of the information transmission, such as format change, etc., but the destination can understand the valid information from the source. Similar expressions in the present application can be understood similarly, and will not be elaborated here.

[0301] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the interaction between each node. Correspondingly, the embodiment of the present application also provides a distributed training device for implementing the above various methods. The distributed training device can be the central node in the above method embodiment, or a component applicable to the central node; or, the distributed training device can be the child node in the above method embodiment, or a component applicable to the child node. It can be understood that, in order to implement the above functions, the distributed training device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0302] The embodiments of the present application can divide the distributed training device into functional modules according to the above method embodiments. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0303] Based on the same concept of the above distributed training method, the present application also provides the following distributed training device:

[0304] As Figure 14 shown, it is a schematic structural diagram of a distributed training device provided by an embodiment of the present application. The distributed training device 1400 includes a transceiver unit 1401 and a processing unit 1402; wherein:

[0305] When the distributed training device is used to implement the functions of the first sub-node in the above method embodiment, the transceiver unit 1401 is used to execute one or more of the operations of the first sub-node in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 shown in Figure 10 the embodiment, and the processing unit 1402 is used to execute one or more of the operations in steps S1002, S1006, S1007, and S1011 shown in Figure 10 the embodiment; or, the transceiver unit 1401 is used to execute one or more of the operations of the first sub-node in steps S1301, S1308, and S1310 shown in Figure 13 the embodiment, and the processing unit 1402 is used to execute one or more of the operations in steps S1302 to S1307 shown in Figure 13 the embodiment. Optionally, the distributed training device can be a terminal device, or a third-party device, such as an OTT or a cloud server, or can be a system composed of a terminal device and a third-party device.

[0306] When the distributed training device is used to implement the functions of the central node in the above method embodiment, the transceiver unit 1401 is used to execute one or more of the operations of the central node in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 shown in Figure 10 the embodiment, and the processing unit 1402 is used to execute one or more of the operations in steps S1004, S1009, and S1013 shown in Figure 10 the embodiment; or, the transceiver unit 1401 is used to execute one or more of the operations shown in Figure 13One or more of the operations of the central node in steps S1301, S1308, and S1310 of the illustrated embodiment, and the processing unit 1402 is used to execute one or more of the steps in S1309 of the illustrated embodiment. Optionally, the distributed training device may be a network device, or a third-party device such as an OTT or a cloud server, or may be a system composed of a network device and a third-party device. Figure 13 One or more of the steps in step S1309 of the illustrated embodiment. Optionally, the distributed training device may be a network device, or a third-party device such as an OTT or a cloud server, or may be a system composed of a network device and a third-party device.

[0307] For the specific implementation of the foregoing transceiver unit 1401 and processing unit 1402, reference may be made to the description in the foregoing method embodiments. In addition, it should be noted that the foregoing transceiver unit and / or processing unit may be implemented by a virtual module. For example, the processing unit may be implemented by a software functional unit or a virtual device, and the transceiver unit may be implemented by a software function or a virtual device. Alternatively, the processing unit or the transceiver unit may also be implemented by an entity circuit. For example, if the device is implemented by a chip / chip circuit, the transceiver unit may be an input / output circuit and / or a communication interface, performing input operations (corresponding to the foregoing receiving operations) and output operations (corresponding to the foregoing sending operations); the processing unit is a processing circuit, such as an integrated processor or a microprocessor or an integrated circuit.

[0308] The division of modules in this application is illustrative, merely a logical function division. In actual implementation, there may be other division methods. In addition, in each example of this application, each functional module may be integrated in a processor, may also exist separately physically, or two or more modules may be integrated in one module. The foregoing integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0309] As Figure 15 Shown in the figure is a schematic structural diagram of another distributed training device provided by an embodiment of this application. The distributed training device 1500 includes one or more processing circuits 1501 (one processing circuit is illustrated in the figure). Optionally, the distributed training device 1500 may further include a memory 1503 (shown in dotted lines in the figure). The memory 1503 is used to store instructions executed by the processing circuit 1501, or store input data required for the processing circuit 1501 to run the instructions, or store data generated after the processing circuit 1501 runs the instructions. Optionally, the distributed training device 1500 may further include an interface circuit 1502 (shown in dotted lines in the figure), and the processing circuit 1501 and the interface circuit 1502 are coupled to each other. It can be understood that the interface circuit 1502 may be a transceiver or an input / output interface.

[0310] Among them, the processing circuit may be a processor or a circuit in the processor for processing.

[0311] When the distributed training device is used to implement the functions of the first sub-node in the above method embodiments, the interface circuit 1502 is used to execute one or more of the operations of the first sub-node in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 of the embodiments shown in Figure 10 , and the processing circuit 1501 is used to execute one or more of the steps S1002, S1006, S1007, and S1011 of the embodiments shown in Figure 10 ; or, the interface circuit 1502 is used to execute one or more of the operations of the first sub-node in steps S1301, S1308, and S1310 of the embodiments shown in Figure 13 , and the processing circuit 1501 is used to execute one or more of the steps S1302 to S1307 of the embodiments shown in Figure 13 . Optionally, the distributed training device may be a terminal device, or a third-party device, such as an OTT or a cloud server, or may be a system composed of a terminal device and a third-party device.

[0312] When the distributed training device is used to implement the functions of the central node in the above method embodiments, the interface circuit 1502 is used to execute one or more of the operations of the central node in steps S1001, S1003, S1005, S1008, S1010, S1012, and S1014 of the embodiments shown in Figure 10 , and the processing circuit 1501 is used to execute one or more of the steps S1004, S1009, and S1013 of the embodiments shown in Figure 10 ; or, the interface circuit 1502 is used to execute one or more of the operations of the central node in steps S1301, S1308, and S1310 of the embodiments shown in Figure 13 , and the processing circuit 1501 is used to execute one or more of the steps S1309 of the embodiments shown in Figure 13 . Optionally, the distributed training device may be a network device, or a third-party device, such as an OTT or a cloud server, or may be a system composed of a network device and a third-party device.

[0313] When the above-mentioned distributed training device is a chip applied to the central node, the chip implements the functions of the central node in the above method embodiments. The chip receives information from other modules in the central node, and this information is sent by the child nodes to the central node; or, the chip sends information to other modules in the central node, and this information is sent by the central node to the child nodes. When the central node is a network device, the modules of the central node here can be the baseband chip of the central node, or CU, DU or other modules, or devices under the open radio access network (O-RAN) architecture, such as open CU, open DU and other devices. When the central node is a third-party device, the modules of the child nodes here can be the processing chips of the third-party device. Among them, the processing chip can be used to implement AI training.

[0314] When the above-mentioned distributed training device is a chip applied to the child node, the chip implements the functions of the child node in the above method embodiments. The chip receives information from other modules in the child node, and this information is sent by the central node to the child node; or, the chip sends information to other modules in the child node, and this information is sent by the child node to the central node. When the child node is a terminal device, the modules of the child node here can be the baseband chip of the child node, or the baseband chip and the processing chip. Among them, the processing chip can be used to implement AI training. When the child node is a third-party device, the modules of the child node here can be the processing chips of the third-party device. Among them, the processing chip can be used to implement AI training.

[0315] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed, the method in the above embodiments is implemented.

[0316] An embodiment of the present application also provides a computer program product containing instructions. When the instructions run on a computer, the computer is caused to execute the method in the above embodiments.

[0317] An embodiment of the present application also provides a distributed training system, including the above-mentioned distributed training device.

[0318] An embodiment of the present application also provides a circuit, which is coupled to a memory and is used to execute the method shown in the above embodiments. The circuit may include a chip circuit.

[0319] Optionally, an embodiment of the present application further provides a chip system, including: at least one processor and an interface, where the at least one processor is coupled to a memory through the interface. When the at least one processor runs a computer program or instruction in the memory, the chip system is caused to execute the method in any one of the above method embodiments. Optionally, the chip system may be composed of chips, or may include chips and other discrete devices. The embodiments of the present application do not make specific limitations thereto.

[0320] The memory in the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data. The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. For example, the memory may be a non-volatile memory, such as a digital versatile disc (DVD), a hard disk drive (HDD), or a solid-state drive (SSD), etc., or may also be a volatile memory, such as a random-access memory (RAM).

[0321] As used in the following description of the present application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0322] It should be understood that in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; wherein A and B can be singular or plural. Also, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be necessarily different. Meanwhile, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.

[0323] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer program instructions on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0324] Although the present application has been described in connection with various embodiments, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce a good effect.

[0325] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The magnitude of the serial numbers of the above processes does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic.

[0326] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0327] The components in the device embodiments of the present application can be combined, divided, and deleted according to actual needs. Those skilled in the art can combine or combine the different embodiments and the features of different embodiments described in this specification.

[0328] In the present application, on the premise of no logical contradiction, the examples can refer to each other. For example, the methods and / or terms between method embodiments can refer to each other, for example, the functions and / or terms between device embodiments can refer to each other, and for example, the functions and / or terms between device examples and method examples can refer to each other.

Claims

1. A distributed training method, characterized in that, The method includes: Receiving first information from a first child node, where the first information is used to indicate a first representation difference between the model of the central node and the model of the first child node, and the first child node is any one of a plurality of child nodes participating in the model training of the central node; Sending second information to the first child node, where the second information is obtained based on the first representation difference, and the second information includes training indication information for the model of the first child node.

2. The method according to claim 1, wherein The first representation difference satisfies a first condition, and the first training indication information is used to indicate the first child node to continue training the model of the first child node.

3. The method according to claim 2, wherein The training indication information is further used to indicate a representation difference coefficient in the loss function of the model of the first child node.

4. The method according to claim 2, wherein The second information further includes a representation difference coefficient in the loss function of the model of the first child node.

5. The method according to claim 3 or 4, characterized in that, The greater the first representation difference, the greater the representation difference coefficient.

6. The method according to claim 1, wherein The first representation difference satisfies a second condition, and the first training indication information is used to indicate the first child node to stop training the model of the first child node.

7. The method according to any one of claims 1-6, characterized in that, The first representation difference is obtained based on the local data set or the public data set of the first child node.

8. The method according to any one of claims 1 to 7, characterized in that The first representation difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, and / or the first representation difference is the difference between the output or intermediate quantity of the models of the plurality of child nodes, where the output or intermediate quantity is obtained based on the same input.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Broadcasting third information, where the third information includes at least one of the following: model configuration information, public data set; where the model configuration information is used to indicate at least one of the following: the type of the models of the plurality of child nodes, the structural information of the models of the plurality of child nodes, the model parameters of the models of the plurality of child nodes, or the training configuration information of the models of the plurality of child nodes.

10. A distributed training method, characterized in that, The method includes: Sending first information to a central node, where the first information is used to indicate a first representation difference between the model of the central node and the model of a first child node, and the first child node is any one of a plurality of child nodes participating in the model training of the central node; Receiving second information from the central node, where the second information is obtained based on the first representation difference, and the second information includes training indication information for the model of the first child node.

11. The method according to claim 10, wherein The first representation difference satisfies a first condition, and the training indication information is used to indicate the first child node to continue training the model of the first child node.

12. The method according to claim 11, wherein The method further includes: Updating the model of the first child node according to the second information.

13. The method according to claim 11 or 12, characterized in that, The training indication information is further used to indicate a representation difference coefficient in the loss function of the model of the first child node.

14. The method according to claim 11 or 12, characterized in that, The second information further includes a representation difference coefficient in the loss function of the model of the first child node.

15. The method according to claim 13 or 14, characterized in that, The greater the first representation difference, the greater the representation difference coefficient.

16. The method according to claim 10, wherein The first representation difference satisfies a second condition, and the training indication information is used to indicate the first child node to stop training the model of the first child node.

17. The method according to any one of claims 10 - 16, characterized in that, The first characterization difference is obtained based on the local data set or the common data set of the first child node.

18. The method according to any one of claims 10-17, characterized in that, The first characterization difference is the difference between the output or intermediate quantity of the model of the central node and the model of the first child node, and / or the first characterization difference is the difference between the output or intermediate quantity of the models of the multiple child nodes, where the output or intermediate quantity is obtained based on the same input.

19. The method according to any one of claims 10-18, characterized in that, The method further includes: Receiving third information, where the third information includes at least one of the following: model configuration information, common data set; where the model configuration information is used to indicate at least one of the following: the type of the models of the multiple child nodes, the structural information of the models of the multiple child nodes, the model parameters of the models of the multiple child nodes, or the training configuration information of the models of the multiple child nodes.

20. A distributed training device, characterized in that, It includes a module for executing the method according to any one of claims 1-9, or includes a module for executing the method according to any one of claims 10-19.

Citation Information

Cited By

  • Distributed training method and apparatus

    WO2025131014A1