Distributed training method, apparatus and system, and chip module and storage medium

By using the gated network model output of the child nodes to align the child nodes in federated learning, the problem of performance degradation of the central node aggregation model caused by the differences in child nodes' understanding is solved, and the generalization ability and training efficiency of the model are improved.

WO2025140055A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141150
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-20
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the federated learning scenario, each child node has different understandings of how to use the output of the gated network model and how to train the expert model, resulting in a degradation in the performance of the aggregation model of the central node.

Method used

The central node instructs the selection parameters so that all the child nodes participating in the training use the selection parameters on the output results of the gated network model during the training of the gated network model and the expert model, so that each child node can align the selection of the output results of the gated network model and improve the performance of the aggregation model of the central node.

Benefits of technology

By aligning the gated network model output results selection of the child nodes, the performance of the aggregation model of the central node is improved, and the generalization ability and training efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141150_03072025_PF_FP_ABST
    Figure CN2024141150_03072025_PF_FP_ABST
Patent Text Reader

Abstract

An artificial intelligence (AI) distributed training method, apparatus and system, and a chip module and a storage medium. A central node indicates a selection parameter, such that each sub-node participating in training selects an output result of a gating network model on the basis of the selection parameter and trains both the gating network model and expert models on the basis of the selected result. Therefore, the sub-nodes align the selection of output results of the gating network model, thereby improving the performance of an aggregation model at the central node.
Need to check novelty before this filing date? Find Prior Art

Description

Distributed training method, device, system, chip module and storage medium

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 26, 2023, with application number 202311819054.6 and invention name “Distributed training method, device, system, chip module and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to artificial intelligence (AI), and in particular to a distributed training method, device, system, chip module, and storage medium. Background Art

[0003] Federated learning (FL) is a machine learning paradigm based on distributed training. In a FL architecture, different child nodes train a common model based on local data, generating local models that effectively fit the characteristics of the local data. However, after these child node local models are aggregated by a central node, the generalization of the overall model improves, but the common model's ability to fit the local data characteristics of each child node decreases, requiring personalized enhancement.

[0004] Therefore, a mixture of experts (MOE) model is proposed. The global model issued by the central node includes multiple expert models. Based on the output of the gating network model, each child node selects the output of at least some expert models to train the global model, obtains gradient information, and reports this gradient information to the central node for model aggregation.

[0005] However, if each child node has a different understanding of how to use the output results of the gating network model and how to train the expert model, the performance of the central node's aggregation model will be degraded.

[0006] In view of this, for federated learning scenarios that include hybrid expert models, how to improve the performance of the aggregation model of the central node is an urgent problem to be solved. Summary of the Invention

[0007] The present application provides a distributed training method, device, system, chip module and storage medium to improve the performance of the aggregation model of the central node.

[0008] In a first aspect, a distributed training method is provided, which is applied to a distributed training system, wherein a global model of the distributed training system includes N expert models and a gated network model, where N is a positive integer, and the method includes: receiving first information, where the first information is used to indicate a selection parameter, where the selection parameter is used to select an output result of the gated network model; and sending second information, where the second information is used to indicate model parameter information of the gated network model and model parameter information of at least one expert model selected from the N expert models, where the selected at least one expert model is obtained based on the output result of the gated network model.

[0009] In this aspect, all child nodes participating in the training receive the selection parameters sent by the central node, and during the training process of the gated network model and the expert model, the selection parameters are applied to the output results of the gated network model, so that each child node can align the selection of the output results of the gated network model and improve the performance of the aggregation model of the central node.

[0010] Exemplarily, the output result of the gated network model is a value between [0, 1] after being processed by a set function.

[0011] In a possible implementation, the method further includes: inputting sample data into the N expert models and the gated network model to obtain N first output results of the N expert models and N second output results of the gated network model, wherein the N second output results correspond one-to-one to the N first output results; for each first output result of the N first output results and each second output result of the N second output results, for the second output result that satisfies the selection parameter, selecting the first output result corresponding to the second output result; and based on the selected at least one first output result and the true value information, obtaining model parameter information of the selected at least one expert model and the model parameter information of the gated network model.

[0012] In another possible implementation, the selection parameter is a first threshold.

[0013] In this implementation, for example, the first threshold value may be in the range of (0, 1). If the output of the gated network model is greater than the first threshold value, the output of the gated network model is retained; if the output of the gated network model is less than the first threshold value, the output of the gated network model is discarded. In particular, if the output of the gated network model is equal to the first threshold value, it can be agreed that the output of the gated network model is retained or discarded.

[0014] In another possible implementation, the first information is further used to indicate at least one of the following information: identifications of the N expert models, a competition mode of the N expert models, a training task, and types of input and output of the global model.

[0015] In this implementation, the competition mode can also be referred to as a usage mode, a cooperation mode, a collaborative mode, or information indicating whether to cooperate. The competition mode indicates whether the N expert models are collaborating or competing. Because the same training node will obtain different gating weights and expert weights when training a hybrid expert model using different competition modes, it is desirable that all participating child nodes use / train the expert model in the same manner. Therefore, the central node can indicate the competition mode of the N expert models to the child nodes.

[0016] In another possible implementation, the method further includes: sending third information, where the third information is used to indicate at least one of the following information: the memory space size of the child node, the computing power information of the child node, whether model training is supported, and the type of model supported for training.

[0017] In yet another possible implementation, the second information is further used to indicate an identifier of the at least one selected expert model.

[0018] In another possible implementation, the obtaining of model parameter information of the selected at least one expert model and model parameter information of the gated network model based on the selected at least one first output result and true value information includes any one of the following operations: obtaining model parameter information of the selected at least one expert model and model parameter information of the gated network model based on the average value and true value information of the selected at least one first output result; or weighting and averaging the selected at least one first output result based on at least one second output result corresponding to each of the selected at least one first output result to obtain a weighted average value of the selected at least one first output result, and obtaining model parameter information of the selected at least one expert model and model parameter information of the gated network model based on the weighted average value and true value information of the selected at least one first output result.

[0019] In yet another possible implementation, the method further includes: receiving fourth information, where the fourth information is used to indicate an updated selection parameter, and the updated selection parameter is obtained based on the second information.

[0020] In this implementation, if the central node updates the selection parameters, it can send fourth information to each child node, where the fourth information is used to indicate the updated selection parameters, so that the fourth information selects the output result of the gated network model based on the updated selection parameters during the next round of training.

[0021] In a second aspect, a distributed training method is provided, which is applied to a distributed training system, wherein the global model of the distributed training system includes N expert models and a gated network model, and the N are positive integers. The method includes: sending first information to multiple child nodes, wherein the first information is used to indicate a selection parameter, and the selection parameter is used to select an output result of the gated network model; and receiving multiple second information respectively, wherein each second information in the multiple second information is used to indicate model parameter information of the gated network model and model parameter information of at least one expert model selected from the N expert models, and the selected at least one expert model is obtained based on the output result of the gated network model.

[0022] In this aspect, the central node indicates the selection parameters so that all child nodes participating in the training will apply the selection parameters to the output results of the gated network model during the training process of the gated network model and the expert model, thereby enabling each child node to align the selection of the output results of the gated network model and improve the performance of the aggregation model of the central node.

[0023] Exemplarily, the output result of the gated network model is a value between [0, 1] after being processed by a set function.

[0024] In a possible implementation, the selection parameter is a first threshold.

[0025] In another possible implementation, the first information is further used to indicate at least one of the following information: identifications of the N expert models, a competition mode of the N expert models, a training task, and types of input and output of the global model.

[0026] In another possible implementation, the method further includes: receiving third information, where the third information is used to indicate at least one of the following information: the memory space size of the child node, the computing power information of the child node, whether model training is supported, and the type of model supported for training.

[0027] In yet another possible implementation, the second information is further used to indicate an identifier of the at least one selected expert model.

[0028] In yet another possible implementation, the method further includes: updating the global model based on the plurality of second information.

[0029] In yet another possible implementation, the method further includes: updating the selection parameter based on the plurality of second information; and sending fourth information, where the fourth information is used to indicate the updated selection parameter.

[0030] In a third aspect, a distributed training device is provided for implementing the distributed training method in the above-mentioned first aspect or any one of the implementations of the first aspect. The device can be a sub-node / third-party device, or a module applied to a sub-node / third-party device (such as a processor, chip, or chip system, etc.), or a logical node, logical module or software that can implement all or part of the functions of the sub-node / third-party device. In one implementation, the distributed training device may include a sending unit, a receiving unit, and may also include a processing unit. The sending unit and the receiving unit may be independent or combined together (which may be referred to as a "transceiver unit").

[0031] In a fourth aspect, a distributed training device is provided for implementing the distributed training method in the second aspect or any one of the implementations of the second aspect. The device can be a central node, or a module applied to the central node (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module or software that can implement all or part of the functions of the central node. In one implementation, the distributed training device may include a sending unit, a receiving unit, and may also include a processing unit. The sending unit and the receiving unit may be independent or combined together (which may be referred to as a "transceiver unit").

[0032] In a possible implementation, the distributed training device in the third to fourth aspects includes a unit for respectively executing the method in any one of the first to second aspects or any one of their implementations.

[0033] In another possible implementation, the distributed training device in the third to fourth aspects above includes a processing circuit coupled to a memory; the processing circuit is configured to enable the device to perform the corresponding functions in the above-mentioned distributed training method. The memory is used to couple with the processing circuit, which stores the necessary programs (instructions) and / or data for the device. Optionally, the distributed training device may further include a communication interface for enabling communication between the device and other network elements. Optionally, the memory may be located inside the distributed training device or outside the distributed training device. Exemplarily, the processing circuit may be a processor or a circuit in a processor for processing.

[0034] When the distributed training device in the third and fourth aspects is a chip, the sending unit may be an output unit, such as an output circuit or a communication interface; the receiving unit may be an input unit, such as an input circuit or a communication interface. When the distributed training device is a terminal device, the sending unit may be a transmitter or a transmitter; and the receiving unit may be a receiver or a receiver.

[0035] In a fifth aspect, a computer-readable storage medium is provided, in which a computer program or instruction is stored. When the computer program or instruction is executed, the methods described in the above aspects are implemented.

[0036] In a sixth aspect, a computer program product comprising instructions is provided, which, when executed on a distributed training device, causes the distributed training device to execute the methods described in the above aspects.

[0037] In a seventh aspect, a distributed training system is provided, which includes the distributed training device described in the third aspect and the distributed training device described in the fourth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] FIG1 is a schematic diagram of the architecture of a distributed training system provided in an embodiment of the present application;

[0039] FIG2 is a simplified schematic diagram of a wireless communication system provided by an embodiment of the present application;

[0040] FIG3 is a schematic diagram of the architecture of another distributed training system provided by the present application;

[0041] 4A to 4D are schematic diagrams of a network architecture provided in an embodiment of the present application;

[0042] FIG5 is a schematic diagram of a neuron structure;

[0043] Figure 6 is a schematic diagram of a neural network;

[0044] Figure 7 is a schematic diagram of an AI application framework;

[0045] FIG8 is a schematic diagram of the architecture of another communication system provided in an embodiment of the present application;

[0046] 9A to 9E are schematic structural diagrams of a global model provided in an embodiment of the present application;

[0047] FIG10 is a flow chart of a distributed training method according to an embodiment of the present application;

[0048] FIG11 is a schematic diagram of optimal beam training according to an embodiment of the present application;

[0049] FIG12 is a schematic structural diagram of a distributed training device provided in an embodiment of the present application;

[0050] FIG13 is a schematic structural diagram of another distributed training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0052] The embodiments of the present application can be applied to a distributed training system as shown in Figure 1, which includes a central node and multiple child nodes. Exemplarily, the distributed training system can be a federated learning system or a gossip learning system. Model parameter information, etc. can be transmitted between the central node and each child node. The distributed training system can also include a third-party device (not shown in the figure), which can serve as a model training entity for the central node / child node.

[0053] The machine learning models trained by this distributed training system can be for non-wireless communication services, such as image recognition, natural language processing, etc., or for wireless communication services, such as beam selection based on environmental information.

[0054] The technology provided by this application can be applied to various communication systems. For example, the communication system can be a fourth generation (4G) th generation, 4G) communication systems (such as long term evolution (LTE) systems), fifth generation (5 th generation (5G) communication systems, worldwide interoperability for microwave access (WiMAX), wireless local area network (WLAN) systems, satellite communication systems, integrated systems of multiple systems, or future communication systems such as the sixth generation (6 th generation, 6G) communication system, etc. Among them, the 5G communication system can also be called a new radio (NR) system.

[0055] A network element in a communication system can send a signal to another network element or receive a signal from another network element. The signal may include information, signaling, or data, etc. The network element can also be replaced by an entity, a network entity, a device, a terminal device, a communication module, a node, a communication node, etc. The present application uses the network element as an example for description. For example, the communication system may include at least one terminal device and at least one network device. The network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the network device. In addition, it can be understood that if the communication system includes multiple terminal devices, the multiple terminal devices can also send signals to each other, that is, the signal sending network element and the signal receiving network element can both be terminal devices.

[0056] Refer to Figure 2, which is a simplified schematic diagram of a wireless communication system provided in an embodiment of the present application. As shown in Figure 2, the wireless communication system includes a wireless access network 100. The wireless access network 100 can be a next-generation (e.g., 6G or higher) wireless access network, or a traditional (e.g., 5G, 4G) wireless access network. One or more terminal devices (120a-120j, collectively referred to as 120) can be connected to each other, or connected to one or more network devices (110a, 110b, collectively referred to as 110) in the wireless access network 100. Optionally, Figure 2 is only a schematic diagram, and the wireless communication system may also include other devices, such as core network devices, wireless relay devices and / or wireless backhaul devices, which are not shown in Figure 2.

[0057] Optionally, in actual applications, the wireless communication system may include multiple network devices (also called access network devices) and multiple terminal devices at the same time. A network device can serve one or more terminal devices at the same time. A terminal device can also access one or more network devices at the same time. The embodiments of the present application do not limit the number of terminal devices and network devices included in the wireless communication system.

[0058] The network device may be an entity on the network side for transmitting or receiving signals. The network device may be an access device for a terminal device to access the wireless communication system in a wireless manner, such as a base station. The base station can broadly cover various names as follows, or be replaced with the following names, such as: radio access network (RAN) node, NodeB, evolved NodeB (eNB), next generation NodeB (gNB), network equipment in open radio access network (O-RAN), relay station, access point, transmission point (TRP), transmitting point (TP), master eNB (MeNB), secondary eNB (SeNB), multi-standard radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, building baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), centralized unit (CU), distributed unit (DRU), etc. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network. The term "network device" refers to a wireless communication device or a wireless network that is configured to communicate with the user via the cellular network.The network device can support networks with the same or different access technologies. The embodiments of the present application do not limit the specific technology and specific device form used by the network device.

[0059] Network devices can be fixed or mobile. For example, base stations 110a and 110b are stationary and are responsible for wireless transmission and reception in one or more cells from terminal device 120. The helicopter or drone 120i shown in Figure 2 can be configured to act as a mobile base station, and one or more cells can move according to the location of the mobile base station 120i. In other examples, the helicopter or drone (120i) can be configured to act as a terminal device communicating with base station 110b.

[0060] In this application, the communication device used to implement the above-mentioned network access function can be a network device, or a network device with partial network access functions, or a device capable of supporting the implementation of the network access function, such as a chip system, a hardware circuit, a software module, or a hardware circuit plus a software module. The device can be installed in the network device or used in combination with the network device. In the method of this application, the communication device used to implement the network device function is described as an example of a network device.

[0061] A terminal device may be an entity on the user side for receiving or transmitting signals, such as a mobile phone. The terminal device may be used to connect people, objects, and machines. The terminal device may communicate with one or more core networks through a network device. The terminal device includes a handheld device with wireless connection capabilities, other processing devices connected to a wireless modem, or a vehicle-mounted device. The terminal device may be a portable, pocket-sized, handheld, computer-built-in, or vehicle-mounted mobile device. The terminal device 120 may be widely used in various scenarios, such as cellular communication, D2D, V2X, point-to-point (P2P), machine-to-machine (M2M), machine type communication (MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, telemedicine, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.Some examples of the terminal device 120 include: user equipment (UE) of the 3GPP standard, fixed equipment, mobile equipment, handheld equipment, wearable equipment, cellular phones, smart phones, session initiated protocol (SIP) phones, laptops, personal computers, smart books, vehicles, satellites, global positioning system (GPS) equipment, target tracking equipment, drones, helicopters, aircraft, ships, remote control equipment, smart home equipment, industrial equipment, personal communication service (PCS) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), wireless network cameras, tablet computers, handheld computers, mobile internet devices (MIDs), wearable devices such as smart watches, VR devices, AR devices, wireless terminals in industrial control, terminals in vehicle networking systems, wireless terminals in self-driving cars, wireless terminals in smart grids, wireless terminals in transportation safety, and smart cities. The terminal device 120 may be a wireless terminal in a city such as a smart gas pump, a terminal device on a high-speed rail, and a wireless terminal in a smart home, such as a smart speaker, a smart coffee machine, a smart printer, etc. The terminal device 120 may be a wireless device in the above various scenarios or a device for being set in a wireless device, for example, a communication module, a modem or a chip in the above device. The terminal device may also be referred to as a terminal, a terminal device, a UE, a mobile station (MS), a mobile terminal (MT), etc. The terminal device may also be a terminal device in a future wireless communication system. The terminal device may be used in a dedicated network device or a general device. The embodiments of the present application do not limit the specific technology and specific device form adopted by the terminal device.

[0062] Alternatively, a terminal device can function as a base station. For example, a UE can act as a dispatching entity, providing sidelink signals between UEs in V2X, D2D, or P2P scenarios. As shown in Figure 2, a cell phone 120a and a car 120b communicate with each other using sidelink signals. Cell phone 120a and smart home device 120e communicate without relaying the communication signal through base station 110b.

[0063] In this application, the communication device used to implement the functions of the terminal device can be a terminal device, or a terminal device with some of the functions of the above terminal devices, or a device that can support the implementation of the functions of the above terminal devices, such as a chip system, which can be installed in the terminal device or used in combination with the terminal device. In this application, the chip system can be composed of chips, or it can include chips and other discrete devices. In the technical solution provided in this application, the communication device is described as a terminal device or UE as an example.

[0064] Optionally, a wireless communication system is typically composed of cells, with base stations providing cell management and communication services to multiple mobile stations (MS) in the cell. The base station includes a baseband unit (BBU) and a remote radio unit (RRU). The BBU and RRU can be placed in different locations, for example: the RRU is remote and placed in an area with high traffic volume, while the BBU is placed in a central computer room. The BBU and RRU can also be placed in the same computer room. The BBU and RRU can also be different components under the same rack. Optionally, a cell can correspond to a carrier or component carrier.

[0065] In some deployments, the network device referred to in the embodiments of the present application may include a CU, a DU, or both a CU and a DU, or a control plane CU node (centralized unit-control plane (CU-CP)), a user plane CU node (centralized unit-user plane (CU-UP)), and a DU node. For example, the network device may include a gNB-CU-CP, a gNB-CU-UP, and a gNB-DU.

[0066] In some deployments, multiple RAN nodes collaborate to assist terminals in achieving wireless access, with different RAN nodes implementing portions of the base station's functionality. For example, a RAN node can be a CU, DU, CU-CP, CU-UP, or RU. The CU and DU can be separate or included in the same network element, such as the BBU. The RU can be included in a radio frequency device or radio unit, such as an RRU, AAU, or RRH.

[0067] The RAN node may support one or more types of fronthaul interfaces, and different fronthaul interfaces correspond to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more baseband functions, and the RU is configured to implement one or more radio frequency functions. If the fronthaul interface between the DU and the RU is another type of interface, relative to the CPRI, some of the downlink and / or uplink baseband functions, such as precoding, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) for downlink, are moved from the DU to the RU for implementation; and for uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / cyclic prefix removal, are moved from the DU to the RU for implementation. In one possible implementation, the interface may be an enhanced common public radio interface (eCPRI). In the eCPRI architecture, the division between the DU and RU is different, corresponding to different types (category, Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, and F.

[0068] Taking eCPRI Cat A as an example, for downlink transmission, based on layer mapping, the DU is configured to implement layer mapping and one or more functions preceding it (i.e., one or more of coding, rate matching, scrambling, modulation, and layer mapping). Other functions after layer mapping (e.g., resource element (RE) mapping, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) addition) are moved to the RU for implementation. For uplink transmission, based on RE demapping, the DU is configured to implement demapping and one or more functions preceding it (i.e., one or more of decoding, rate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, and RE demapping). Other functions after demapping (e.g., one or more of digital BF or FFT / CP removal) are moved to the RU for implementation. It is understandable that for the functional description of DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol, which will not be described in detail here.

[0069] In one possible design, the processing unit for implementing baseband functions in the BBU is called a baseband high layer (BBH) unit, and the processing unit for implementing baseband functions in the RRU / AAU / RRH is called a baseband low layer (BBL) unit.

[0070] In different systems, CU (or CU-CP and CU-UP), DU or RU may also have different names, but those skilled in the art can understand their meanings. For example, in an open radio access network (ORAN) system, CU may also be referred to as O-CU (open CU), DU may also be referred to as O-DU, CU-CP may also be referred to as O-CU-CP, CU-UP may also be referred to as O-CU-UP, and RU may also be referred to as O-RU. Any of the CU (or CU-CP, CU-UP), DU and RU in this application may be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.

[0071] In the embodiments of the present application, the device for implementing the functions of the network device can be a network device; it can also be a device that can support the network device to implement the functions, such as a chip system, a hardware circuit, a software module, or a hardware circuit and a software module. The device can be installed in the network device or used in conjunction with the network device. In the embodiments of the present application, only the device for implementing the functions of the network device is used as an example to illustrate, and does not constitute a limitation on the solutions of the embodiments of the present application.

[0072] It is understandable that the present application can be applied between network devices and terminal devices.

[0073] Protocol layer structure between network devices and terminal devices:

[0074] The communication between the network device and the terminal device follows a certain protocol layer structure. The protocol layer structure may include a control plane protocol layer structure and a user plane protocol layer structure. For example, the control plane protocol layer structure may include the functions of the radio resource control (RRC) layer, the packet data convergence protocol (PDCP) layer, the radio link control (RLC) layer, the medium access control (MAC) layer, and the physical layer. For example, the user plane protocol layer structure may include the functions of the PDCP layer, the RLC layer, the MAC layer, and the physical layer. In one possible implementation, a service data adaptation protocol (SDAP) layer may also be included above the PDCP layer.

[0075] Optionally, the protocol layer structure between the network device and the terminal device may further include an artificial intelligence (AI) layer for transmitting data related to AI functions.

[0076] Taking data transmission between network devices and terminal devices as an example, data transmission needs to pass through the user plane protocol layers, such as the SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer. The SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer can also be collectively referred to as the access layer. Data transmission is divided into sending or receiving based on the direction of transmission, and each of these layers is further divided into a sending part and a receiving part. Taking downlink data transmission as an example, after the PDCP layer obtains data from the upper layer, it transmits the data to the RLC layer and MAC layer. The MAC layer then generates a transport block, which is then wirelessly transmitted through the physical layer. Data is encapsulated accordingly in each layer. For example, data received by a layer from the layer above it is considered a service data unit (SDU) of that layer. After encapsulation by that layer, it becomes a protocol data unit (PDU) and is then passed to the next layer.

[0077] For example, a terminal device may also include an application layer and a non-access layer. The application layer can be used to provide services to applications installed in the terminal device. For example, downlink data received by the terminal device can be sequentially transmitted from the physical layer to the application layer, which then provides it to the application. For another example, the application layer can obtain data generated by the application and sequentially transmit the data to the physical layer for transmission to other communication devices. The non-access layer can be used to forward user data, such as forwarding uplink data received from the application layer to the SDAP layer, or forwarding downlink data received from the SDAP layer to the application layer.

[0078] It should be understood that the number and type of each device in the communication system shown in Figure 2 are for illustration only, and the present application is not limited to this. In actual applications, the communication system may also include more terminal devices, more network devices, and other network elements, such as core network devices, and / or network elements for implementing artificial intelligence functions.

[0079] It is understandable that all or part of the functions implemented by one or more of the terminal devices, network devices, core network devices, or network elements for implementing artificial intelligence functions can be virtualized, that is, implemented by one or more of the proprietary processors or general-purpose processors and the corresponding software modules. Among them, since the terminal devices and network devices involve interfaces for air interface transmission, the transceiver functions of the interfaces can be implemented by hardware. Core network devices, such as operation administration and maintenance (OAM) network elements, can be virtualized. Optionally, one or more functions of the virtualized terminal devices, network devices, core network devices, or network elements for implementing artificial intelligence functions can be implemented by cloud devices, such as cloud devices in over the top (OTT) systems.

[0080] In an embodiment of the present application, when the central node is a network device and the child node is a terminal device (such as a UE), the network device and UE1 to UE5 can form a distributed AI training system as shown in Figure 3. In this communication system, UE1 to UE5 can send data to the network device, and the network device needs to receive uplink data sent by UE1 to UE5. The uplink data can be the characterization difference or model parameter calculated by the child node, or it can be the feedback amount containing its status information. At the same time, the network device can send configuration information to UE1-UE5. The configuration information can be the model parameter data used by the central node to synchronize each child node, or it can be control data that indicates the training method of the child node. The data between the network equipment and the UE can be carried on physical channels, such as the physical downlink control channel (PDCCH), the physical downlink shared channel (PDSCH), the physical uplink shared channel (PUSCH) or the physical uplink control channel (PUCCH); for example, the physical sidelink control channel (PSCCH) and the physical sidelink shared channel (PSSCH).

[0081] In order to support AI technology in wireless networks, AI nodes may also be introduced into the network.

[0082] Optionally, the AI ​​node can be deployed in one or more of the following locations in the communication system: network equipment, terminal equipment, or core network equipment. Alternatively, the AI ​​node can be deployed separately, for example, in a location other than any of the above devices, such as a host or cloud server in an over-the-top (OTT) system. The AI ​​node can communicate with other devices in the communication system, such as one or more of the following: network equipment, terminal equipment, or core network elements.

[0083] It is understood that this application does not limit the number of AI nodes. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on function, such as different AI nodes are responsible for different functions.

[0084] It can also be understood that AI nodes can be independent devices, or they can be integrated into the same device to implement different functions, or they can be network elements in hardware devices, or they can be software functions running on dedicated hardware, or they can be virtualized functions instantiated on a platform (for example, a cloud platform). This application does not limit the specific form of the above-mentioned AI nodes.

[0085] An AI node can be an AI network element or an AI module.

[0086] One or more AI modules are provided in one or more of these network element nodes, such as core network equipment, access network nodes (RAN nodes), terminals or OAM devices. The access network node can be a separate RAN node, or it can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be provided with one or more AI modules. Optionally, the CU can also be split into CU-CP and CU-UP. One or more AI models are provided in the CU-CP and / or CU-UP.

[0087] The AI ​​module is used to implement the corresponding AI function. The AI ​​modules deployed in different network elements can be the same or different. The model of the AI ​​module can implement different functions according to different parameter configurations. The model of the AI ​​module can be configured based on one or more of the following parameters: structural parameters (such as the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of the neuron, the activation function of the neuron, or at least one of the bias in the activation function), input parameters (such as the type of input parameters and / or the dimension of the input parameters), or output parameters (such as the type of output parameters and / or the dimension of the output parameters). Among them, the bias in the activation function can also be called the bias of the neural network.

[0088] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or on the same node or device.

[0089] The communication system includes a RAN intelligent controller. For example, the RIC can be the above-mentioned AI module, which is used to implement AI-related functions. The RIC includes a near-real-time RIC (near-real time RIC, near-RT RIC) and a non-real-time RIC (non-real time RIC, non-RT RIC). Among them, the non-real-time RIC mainly processes non-real-time information, such as data that is not sensitive to latency, and the latency of this data can be in the order of seconds. The real-time RIC mainly processes near-real-time information, such as data that is relatively sensitive to latency, and the latency of this data is in the order of tens of milliseconds.

[0090] Near real-time RIC is used for model training and reasoning. For example, it is used to train an AI model and use the AI ​​model for reasoning. Near real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or reasoning data. Optionally, near real-time RIC can deliver the reasoning results to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and DU, and / or between the DU and RU. For example, the near real-time RIC delivers the reasoning results to the DU, and the DU sends it to the RU.

[0091] Non-real-time RIC is also used for model training and reasoning. For example, it is used to train AI models and use the models for reasoning. Non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU and / or RU) and / or terminals. This information can be used as training data or reasoning data, and the reasoning results can be submitted to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and DU, and / or between the DU and RU. For example, the non-real-time RIC submits the reasoning results to the DU, and the DU sends it to the RU.

[0092] The near-real-time RIC and non-real-time RIC can also be set up as separate network elements. Optionally, the near-real-time RIC and non-real-time RIC can also be part of other devices. For example, the near-real-time RIC is set up in a RAN node (e.g., a CU or DU), while the non-real-time RIC is set up in an OAM, a cloud server, a core network device, or other network devices.

[0093] For example, the configuration of near real-time RIC and non-real-time RIC in the network architecture may be as shown in FIG4A to FIG4D :

[0094] As shown in (a) of FIG4A , in a first possible implementation, the network device includes a near real-time RIC module for performing model learning and / or reasoning.

[0095] As shown in (b) of FIG4A , in a second possible implementation, in a communication system, a non-real-time RIC may be included outside the network device. Optionally, the non-real-time RIC may be located in the OAM or in a core network device.

[0096] As shown in (c) of Figure 4A, in a third possible implementation, the network device includes a near real-time RIC and a non-real-time RIC is also included outside the network device. Optionally, the non-real-time RIC can be located in the OAM or the core network device.

[0097] Compared to (c) in Figure 4A, the CU is separated into CU-CP and CU-UP in Figure 4B. The settings of near real-time RIC and non-real-time RIC are the same as those in (c) in Figure 4A.

[0098] As shown in Figure 4C, optionally, the network device includes one or more AI entities, and the function of the AI ​​entity is similar to the above-mentioned near real-time RIC. Optionally, the OAM includes one or more AI entities, and the function of the AI ​​entity is similar to the above-mentioned non-real-time RIC. Optionally, the core network device includes one or more AI entities, and the function of the AI ​​entity is similar to the above-mentioned non-real-time RIC. When both the OAM and the core network device include AI entities, the models trained by their respective AI entities are different, and / or the models used for reasoning are different. In the present application, the difference in models may include at least one of the following differences: structural parameters of the model (such as the number of layers of the model, and / or weights, etc.), input parameters of the model, or output parameters of the model.

[0099] Relative to Figure 4C, the network device in Figure 4D is separated into CU and DU. Optionally, the CU may include an AI entity, and the function of the AI ​​entity is similar to the above-mentioned near real-time RIC. Optionally, the DU may include an AI entity, and the function of the AI ​​entity is similar to the above-mentioned near real-time RIC. When both the CU and the DU include AI entities, the models trained by their respective AI entities are different, and / or the models used for reasoning are different. Optionally, the CU in Figure 4D can be further split into CU-CP and CU-UP. Optionally, one or more AI models can be deployed in the CU-CP. And / or, one or more AI models can be deployed in the CU-UP. Optionally, in Figure 4C or Figure 4D, the OAM of the network device and the OAM of the core network device can be deployed separately and independently.

[0100] For ease of understanding, the following first introduces the AI ​​technology involved in this application. It should be understood that this introduction does not limit this application.

[0101] (1) AI Model

[0102] AI refers to the intelligence exhibited by machines created by humans. Generally, AI refers to the technology that replicates human intelligence through ordinary computer programs. AI can be defined as machines or computers that mimic humans and possess cognitive functions associated with human thinking, such as learning and problem-solving. AI is able to learn from past experiences, make rational decisions, and respond quickly. The goal of AI is to understand intelligence by building computer programs capable of symbolic reasoning or deduction.

[0103] Machine learning (ML) is a path to artificial intelligence (AI), specifically using machine learning to solve AI problems. Machine learning theory primarily involves the design and analysis of algorithms that enable computers to automatically "learn." Machine learning algorithms automatically analyze data to identify patterns and use these patterns to make predictions about unknown data. Because learning algorithms involve extensive statistical theory, machine learning is particularly closely linked to inferential statistics, also known as statistical learning theory.

[0104] Machine learning can be divided into supervised learning, unsupervised learning, and reinforcement learning.

[0105] Supervised learning uses a machine learning algorithm to learn the mapping relationship between sample values ​​and sample labels based on collected sample values ​​and sample labels. This learned mapping relationship is then expressed using a machine learning model. The process of training a machine learning model is the process of learning this mapping relationship. For example, in signal detection, a noisy received signal is a sample, and the true constellation point corresponding to this signal is the label. Through training, machine learning aims to learn the mapping relationship between samples and labels, essentially enabling the machine learning model to become a signal detector. During training, the model parameters are optimized by calculating the error between the model's predicted values ​​and the true labels. Once the mapping relationship is learned, it can be used to predict the label of each new sample. The mapping relationship learned by supervised learning can include linear and nonlinear mappings. Learning tasks can be categorized into classification and regression tasks based on the type of label.

[0106] Unsupervised learning relies solely on collected sample values, using algorithms to discover inherent patterns within them. One type of unsupervised learning algorithm uses the samples themselves as supervisory signals, meaning the model learns the mapping from one sample to another. This is called self-supervised learning. During training, the model parameters are optimized by calculating the error between the model's predictions and the samples themselves. Self-supervised learning can be used in signal compression and decompression recovery applications. Common algorithms include autoencoders and generative adversarial networks.

[0107] Reinforcement learning, unlike supervised learning, is a type of algorithm that learns problem-solving strategies through interaction with the environment. Unlike supervised and unsupervised learning, reinforcement learning problems lack explicit label data for "correct" actions. Instead, the algorithm must interact with the environment to obtain reward signals from the environment, and then adjust its decision-making actions to maximize the reward signal value. For example, in downlink power control, the reinforcement learning model adjusts the downlink transmit power of each user based on the overall system throughput fed back by the wireless network, hoping to achieve higher system throughput. The goal of reinforcement learning is also to learn the mapping between environmental states and optimal decision-making actions. However, because the labels for "correct actions" are not available in advance, network optimization cannot be achieved by calculating the error between actions and "correct actions." Reinforcement learning training is achieved through iterative interaction with the environment.

[0108] An AI model is an algorithm or computer program that implements AI functions. It is the concrete implementation of AI technology functions. An AI model represents the mapping relationship between the model's input and output. AI models can be neural networks, linear regression models, decision tree models, support vector machines (SVMs), Bayesian networks, Q-learning models, or other machine learning models.

[0109] (2) Deep neural network (DNN)

[0110] Deep neural networks are a specific implementation of AI or machine learning technology. According to the universal approximation theorem, neural networks can theoretically approximate any continuous function, enabling them to learn arbitrary mappings. Traditional communication systems require extensive expert knowledge to design communication modules. However, deep learning communication systems based on DNNs can automatically discover implicit patterns in massive data sets and establish mapping relationships between data, achieving performance superior to traditional modeling methods.

[0111] The idea of ​​DNN is derived from the neuronal structure of the brain. For example, each neuron performs a weighted sum operation on its input values ​​and outputs the result through an activation function. Figure 5 shows a schematic diagram of the neuron structure. Assume that the input of the neuron is x = [x0, x1, ..., xn ], and the weights corresponding to each input are w=[w0,w1,…,w n ], where w i As x i The weight of x i Weighted. The bias of the weighted sum of the input values ​​according to the weight is, for example, b. The activation function can take many forms. Assuming that the activation function of a neuron is: y = f(z) = max(0,z), then the output of the neuron is:

[0112] For another example, if the activation function of a neuron is: y = f(z) = z, then the output of the neuron is:

[0113] Among them, b, w i 、x i It can be a decimal, an integer (such as 0, a positive integer or a negative integer), or a complex number. The activation functions of different neurons in a neural network can be the same or different.

[0114] A neural network generally includes multiple layers, each of which may include one or more neurons. By increasing the depth and / or width of a neural network, its expressive power can be improved, providing more powerful information extraction and abstract modeling capabilities for complex systems. The depth of a neural network can refer to the number of layers it comprises, while the number of neurons in each layer can be referred to as the width of that layer. In one implementation, a neural network includes an input layer and an output layer. The input layer processes the input information received by the neural network through neurons, passes the processing results to the output layer, and the output layer obtains the output of the neural network. In another implementation, the neural network includes an input layer, a hidden layer, and an output layer. The neural network schematic diagram in Figure 6 is provided. The input layer processes the input information received by the neural network through neurons, passes the processing results to an intermediate hidden layer, which performs calculations on the received processing results to obtain a calculation result. The hidden layer then passes the calculation results to the output layer or an adjacent hidden layer, and the output layer ultimately obtains the output of the neural network. A neural network can include one hidden layer or multiple hidden layers connected in sequence, without limitation.

[0115] Depending on how the network is constructed, DNNs can include feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). Figure 6 shows an FNN network, which is characterized by complete connectivity between neurons in adjacent layers. This typically requires a large amount of storage space and results in high computational complexity.

[0116] CNNs are neural networks specifically designed to process data with a grid-like structure. For example, time series data and image data can both be considered grid-like. CNNs don't use all input information at once for computation. Instead, they use a fixed-size window to intercept a portion of the information for convolution operations, significantly reducing the computational complexity of model parameters. Furthermore, depending on the type of information intercepted by the window (e.g., people and objects in an image represent different types of information), different convolution kernels can be used for each window, enabling CNNs to better extract features from the input data.

[0117] RNNs are a type of DNN that utilizes feedback time series information. Their input consists of a new input value at the current moment and their own output value at the previous moment. RNNs are suitable for capturing temporally correlated sequence features and are particularly well-suited for applications such as speech recognition and channel coding.

[0118] The above-mentioned FNN, CNN, and RNN are common neural network structures, which are all constructed based on neurons. As mentioned above, each neuron performs a weighted sum operation on its input values, and the weighted summation result generates an output through a nonlinear function. We call the weights of the weighted summation operation of neurons in the neural network and the nonlinear function the parameters of the neural network. Taking the neuron with max{0,x} as the nonlinear function as an example, The parameters of the neuron to be operated are weights w=[w0,…,w n ], the weighted sum bias is b, and the nonlinear function max{0,x}. The parameters of all neurons in a neural network constitute the parameters of the neural network.

[0119] (3) Training dataset and inference data

[0120] The training dataset is used to train the AI ​​model. The training dataset may include the input of the AI ​​model, or the input and target output of the AI ​​model. The training dataset includes one or more training data. The training data may be a training sample input to the AI ​​model or the target output of the AI ​​model. The target output may also be referred to as a label or a label sample. The training dataset is an important part of machine learning. Model training is essentially learning certain features from the training data so that the output of the AI ​​model is as close as possible to the target output, such as the difference between the output of the AI ​​model and the target output is as small as possible. The composition and selection of the training dataset can, to a certain extent, determine the performance of the trained AI model.

[0121] In addition, during the training process of an AI model (such as a neural network), a loss function can be defined. The loss function describes the gap or difference between the output value of the AI ​​model and the target output value. This application does not limit the specific form of the loss function. The training process of the AI ​​model is a process of adjusting the model parameters of the AI ​​model so that the value of the loss function is less than the threshold, or the value of the loss function meets the target requirements. For example, the AI ​​model is a neural network, and adjusting the model parameters of the neural network includes adjusting at least one of the following parameters: the number of layers, width, weights of neurons, or parameters in the activation function of neurons.

[0122] Inference data can be used as input to a trained AI model for inference. During the model inference process, the inference data is input into the AI ​​model, and the corresponding output is the inference result.

[0123] (4) AI model design

[0124] The design of an AI model primarily includes a data collection phase (e.g., collecting training data and / or inference data), a model training phase, and a model inference phase. It may further include an inference result application phase. See Figure 7, which illustrates an AI application framework. In the aforementioned data collection phase, a data source is used to provide training data sets and inference data. In the model training phase, an AI model is obtained by analyzing or training the training data provided by the data source. The AI ​​model represents the mapping relationship between the model's input and output. Learning an AI model through model training nodes is equivalent to learning the mapping relationship between the model's input and output using training data. In the model inference phase, the AI ​​model trained in the model training phase is used to perform inference based on the inference data provided by the data source to obtain an inference result. This phase can also be understood as: inputting inference data into the AI ​​model, obtaining an output through the AI ​​model, and the output being the inference result. The inference result may indicate: configuration parameters used (executed) by the execution object, and / or the operation performed by the execution object. Inference results are published during the application phase. For example, the inference results can be centrally planned by an execution entity (actor). For example, the execution entity can send the inference results to one or more execution targets (e.g., core network equipment, network equipment, or terminal devices) for execution. Furthermore, the execution entity can provide feedback on model performance to the data source to facilitate subsequent model update and training.

[0125] It is understood that a communication system may include network elements with artificial intelligence capabilities. The aforementioned AI model design-related steps can be performed by one or more network elements with AI capabilities. In one possible design, AI capabilities (such as AI modules or AI entities) can be configured within existing network elements in the communication system to implement AI-related operations, such as AI model training and / or inference. For example, these existing network elements can be network devices (such as gNBs), terminal devices, core network devices, or network management systems. Network management systems can categorize network management tasks into three categories based on the actual needs of the operator's network operations: operations, administration, and maintenance. Network management systems are also referred to as Operational and Administered Management (OAM) network elements, or OAM for short. Operations primarily perform routine network and service analysis, forecasting, planning, and configuration; maintenance primarily involves daily operational activities such as testing and fault management of the network and its services. Network management systems can monitor network operating status, optimize network connectivity and performance, improve network stability, and reduce network maintenance costs. Alternatively, in another possible design, independent network elements can be introduced into the communication system to perform AI-related operations, such as training AI models. The independent network element can be called an AI network element or an AI node, etc., and this application does not limit this name. The AI ​​network element can be directly connected to the network equipment in the communication system, or it can be indirectly connected to the network equipment through a third-party device. Among them, the third-party device can be a core network element such as an authentication management function (AMF) network element, a user plane function (UPF) network element, an OAM, a cloud server or other network elements, without limitation. For example, referring to Figure 8, the communication system includes a network device 810, terminal devices 820 and 830, and an AI network element 840 is also introduced in the communication system.

[0126] In this application, a model can be inferred to obtain a single parameter or multiple parameters. The training process of different models can be deployed on different devices or nodes, or on the same device or node. The inference process of different models can be deployed on different devices or nodes, or on the same device or node.

[0127] Among them, the model parameters may include one or more of the following structural parameters of the model (such as the number of layers and / or weights of the model, etc.), the input parameters of the model (such as input dimension, number of input ports), or the output parameters of the model (such as output dimension, number of output ports). It can be understood that the input dimension may refer to the size of an input data. For example, when the input data is a sequence, the input dimension corresponding to the sequence may indicate the length of the sequence. The number of input ports may refer to the number of input data. Similarly, the output dimension may refer to the size of an output data. For example, when the output data is a sequence, the output dimension corresponding to the sequence may indicate the length of the sequence. The number of output ports may refer to the number of output data.

[0128] (5) Centralized training

[0129] Over the past decade, the number of smart devices, such as mobile phones and wearables, has continued to increase. It is foreseeable that billions of IoT devices will be deployed across communication networks in the near future, enabling the automation and intelligence of social operations. The performance of current intelligent services based on advanced machine learning is likely to benefit from the explosive growth of data on these devices, as well as the computing power available in these UEs. Most machine learning techniques, such as deep neural network-based learning algorithms, require centralized training using all available data. However, centralized training requires the collection of sufficient data, often originating from UEs. This data must be uploaded by the UEs, which incurs significant upload overhead. Furthermore, the concentration of massive amounts of training data in a single node limits training efficiency due to storage space constraints. Furthermore, collecting UE-side data may infringe on user privacy. Storing massive amounts of data for centralized training can raise significant concerns about privacy leaks. On the other hand, if users train machine learning algorithms solely using their own data, training effectiveness is often limited by the limited amount of local data.

[0130] (6) Distributed training

[0131] Distributed training is an effective solution to these challenges. This technology allows the machine learning process to be divided among multiple sub-nodes on the user side, achieving scalability of learning algorithms. It allows a cloud or server, acting as a central node, to collect machine learning models trained by multiple sub-nodes. The central node then integrates these models to improve the overall machine learning training results. Because the training data is always stored on the sub-nodes, distributed learning technology has the potential to achieve the same performance as centralized training while utilizing the user end's data and / or computing power, while protecting user data privacy.

[0132] Federated learning is a distributed machine learning paradigm. Its original purpose was to effectively enable multiple organizations to utilize data and conduct machine learning modeling while ensuring user privacy and data security. Within the federated learning framework, nodes communicate not the data itself but rather intermediate results from training, such as model parameters or gradients. As a distributed machine learning paradigm, federated learning can effectively address data silos, enabling participants to jointly model without sharing data. This technically breaks down data silos and enables AI collaboration. Based on the distribution of data sources among participating parties, federated learning can be categorized into three types: horizontal federated learning, vertical federated learning, and federated transfer learning.

[0133] Horizontal federated learning refers to the situation where two datasets have a lot of overlap in user features but little in user overlap. We split the dataset horizontally (i.e., the user dimension) and extract the data with the same user features but different user identities for training. Vertical federated learning refers to the situation where two datasets have a lot of overlap in users but little in user features. We split the dataset vertically (i.e., the feature dimension) and extract the data with the same users but different user identities for training. Federated transfer learning refers to the situation where two datasets have little overlap in both users and user features. We do not split the data, but instead use transfer learning to overcome insufficient data or labels.

[0134] Taking the high-frequency beam management problem as an example, when the network device side performs beam scanning of codebook-based synchronization signals and physical broadcast channel (PBCH) blocks (i.e., synchronization signal blocks (SSBs)) or channel state information-reference signals (CSI-RSs), the channels between users and network devices at different locations are different. The user measures the received SSB or CSI-RS beam, such as measuring the physical layer reference signal received power (L1-reference signal receiving power, L1-RSRP), and feeds back the beam identity (ID) corresponding to the maximum RSRP value. AI / ML can be used to train a model that uses, for example, a plurality of received SSB / CSI-RS signals (part or all of them), or the strength (RSRP) (part or all of) of a plurality of received SSB / CSI-RS signals, or the estimated channel as input to infer the optimal beam ID and feed it back to the network device. Each user can collect their own receive beam / channel information and the corresponding optimal beam ID as samples (i.e., local samples) for training the aforementioned AL / ML model. However, the number of samples each user can collect is limited, and the performance of the model trained solely on local data is limited. Specifically, due to the user's location, the optimal beam ID may only be a subset of the SSB / CSI-RS codebook. If the user sends local data to the server, the server aggregates the data of all users for model training. While this improves model performance, it also carries the risk of leaking user privacy information, such as inferring the user's current location through the channel. To address this issue, federated learning can be used. The central node distributes a global model to each participating user. Each user trains the global model using local data to obtain a local model. The local model parameter information, such as gradients and weights, is then sent (encrypted) to the server. The server performs model aggregation (MA) to update the global model and then sends the global model to each user. The user then continues to update the local model and sends it to the central node. This process repeats multiple times until convergence.

[0135] In a federated learning architecture, different child nodes train a common model based on local data, generating local models that better fit the characteristics of the local data. However, after these local models of the child nodes are aggregated by the central node, the generalization of the entire model improves, but the common model's ability to fit the local data characteristics of each child node decreases, requiring personalized enhancement.

[0136] Therefore, a hybrid expert model is proposed. The global model issued by the central node includes multiple expert models. Based on the output of the gating network model, each child node selects the output of at least some expert models to train the global model, obtains gradient information, and reports this gradient information to the central node for model aggregation.

[0137] Hybrid expert models combine multiple networks, or expert models, to achieve better model performance. Large models based on hybrid expert models can effectively improve their training efficiency and the number of model parameters.

[0138] As shown in Figures 9A to 9E, various global model structures are given.

[0139] Based on whether the global model includes a general layer, whether the expert layer includes a gated network model, and whether there is a clear definition of the expert layer, the following classifications can be made:

[0140] (1) The expert layer includes a gated network model and a clear definition of the expert layer:

[0141] In Figure 9A, the global model includes a general layer and an expert layer. The general layer is a network layer that is common to all input data and can be a feature extraction network such as a convolutional neural network. The expert layer also includes a gated network model and N expert models. N is a positive integer. The N expert models can be based on different types of neural networks, such as CNN, RNN, Transformer, MLP, etc.; or based on the same type of neural network, but with different parameter configurations, such as the number of layers, depth, and specific configurations including kernel size.

[0142] The gating network model is trained, and its output is used to select the output of the expert model. Choosing an appropriate and complementary gating output mechanism can combine and balance the expert choices. Different expert models here can refer to different network structures, for example, expert model 1 can be a convolutional network, expert model 2 can be a fully connected network, and so on. However, it is important to ensure that the outputs of the various expert networks can be combined. Selecting a subset of expert models for prediction based on the output of the gating network model reduces computational effort and allows the selection of the most appropriate expert model for different inputs.

[0143] In this architecture, the output of the general layer serves as the common input to the expert model and the gating network model.

[0144] In Figure 9B, the global model includes only the expert layer and does not include the general layer. The expert layer includes a gated network model and N expert models. The meanings of the expert model and the gated network model can be found in the above description.

[0145] In this architecture, the input of the local model serves as the common input of the expert model and the gating network model.

[0146] (2) The expert layer does not include the gated network model, and there is a clear definition of the expert layer:

[0147] In Figure 9C , the expert layer does not include a gated network model, but only N expert models. The gated network model is independent of the expert layer. The meanings of expert models and gated network models can be found in the above description. The expert layer and the gated network model share the same input, and the result of the universal preprocessing serves as the common input for both the expert layer and the gated network model. This universal preprocessing result may or may not be based on the output of the universal layer model.

[0148] In Figure 9D , the expert layer does not include a gated network model; it only includes N expert models. The gated network model is independent of the expert layer. The meanings of expert models and gated network models can be found in the above description. Furthermore, the inputs of the expert layer and the gated network model are different. The results of the first type of preprocessing on the input local data serve as the input for expert layers 1 through N; the results of the second type of preprocessing on the input local data serve as the input for the gated network model.

[0149] (3) No clear definition of the expert layer:

[0150] In Figure 9E , there is no clear definition of the expert layer. The global model includes N networks and a gated network model. The functions and roles of these N networks can be referred to above for the expert model. The meaning of the gated network model can be referred to above. The inputs of Networks 1 through N are different from those of the gated network model. The result of the first type of preprocessing on the input local data serves as the input for Network 1; the result of the second type of preprocessing on the input local data serves as the input for Network 2; and so on; the result of the N+1th type of preprocessing on the input serves as the input for the gated network model.

[0151] However, if each child node has a different understanding of how to use the output results of the gating network model and how to train the expert model, the performance of the central node's aggregation model will be degraded.

[0152] To address the above problems, the present application provides a distributed training solution. The central node indicates the selection parameters so that all sub-nodes participating in the training will apply the selection parameters to the output results of the gated network model during the training of the gated network model and the expert model. This allows each sub-node to align the selection of the output results of the gated network model, thereby improving the performance of the aggregation model of the central node.

[0153] The distributed training method provided by the embodiment of the present application is described in detail below. It will be understood that the present application uses the central node and the sub-node as an example to illustrate the execution subject of the interactive diagram, but the present application does not limit the execution subject of the interactive diagram. For example, the central node in the method provided by the present application may also be a chip, chip system, circuit or processor applied to the central node, or a logical node, logic module or software that can realize all or part of the central node; the sub-node in the method provided by the present application may also be a chip, chip system, circuit or processor applied to the sub-node, or a logical node, logic module or software that can realize all or part of the sub-node functions.

[0154] As shown in Figure 10, a flow chart of a distributed training method provided in an embodiment of the present application is provided. The method is applied to a distributed training system, which includes a central node and multiple child nodes. The global model of the distributed training system includes N expert models and a gated network model, where N is a positive integer. Exemplarily, the method may include the following steps:

[0155] S1001. The child node sends third information to the central node. Correspondingly, the central node receives the third information.

[0156] The child node is any child node participating in the distributed training. This embodiment is described by taking the interaction process of a child node and a central node as an example. The interaction process of other child nodes and central nodes can refer to the interaction process of this embodiment.

[0157] When initially building the distributed training system, child nodes can report their capabilities to the central node. For example, the child nodes send third information to the central node, where the third information indicates at least one of the following: the child node's memory size, the child node's computing power, whether it supports model training, and the types of models supported for training.

[0158] When training begins, the central node sends the central node's initialization model to the child node, allowing the child node to train the initialization model. Therefore, before training begins, the child node can report the size of its memory space to the central node, so that the central node knows whether the child node's memory space is large enough to store the central node's initialization model and subsequent training data. The memory space size of the child node refers to the amount of memory space available for the child node to store the AI / ML model.

[0159] Before training begins, child nodes can also report their computing power information to the central node, allowing the central node to know whether the child node has sufficient computing power and can provide timely feedback on the trained model information. The computing power information of the child node refers to the computing power required to run the AI / ML model.

[0160] Before training begins, the child node can also report to the central node whether it supports model training, so that the central node can determine whether the child node can participate in distributed training and send the global model to it.

[0161] Before training begins, the child node can also report the types of models it supports to the central node so that the central node can determine whether the child node can participate in distributed training. For example, the types of models supported by the child node include CNN, RNN, fully connected, random forest models, etc.

[0162] In addition, the third information may also include hardware information of the subnode, including but not limited to the antenna configuration of the subnode (number of antennas, polarization direction, etc.), number of RF channels, sensor type (position sensor / global positioning system (GPS), motion sensor, etc.) and parameters.

[0163] It is understandable that since it is federated learning, the child nodes use local data for training, so the child nodes do not need to report a series of information related to the actual collected data or involving privacy, such as the amount of data that can be processed.

[0164] It is understandable that the central node may also obtain the above information of the child nodes in advance. Therefore, this step is optional and is represented by a dotted line in the figure.

[0165] S1002: The central node sends first information to the child node, and the child node receives the first information accordingly.

[0166] After receiving the third information reported by each child node, the central node selects a child node to participate in this round of distributed training based on the third information of each child node.

[0167] As shown in Figures 9A-9E , the global model includes N expert models (or Network 1 through Network N as shown in Figure 9E ). These N expert models can be based on different types of neural networks, such as CNN, RNN, Transformer, and MLP. They can also be based on the same type of neural network but with different parameter configurations, such as the number of layers, depth, and specific configurations including kernel size. Therefore, the output results of the N expert models may differ.

[0168] Based on experience and simulation results, the more sparsely the output results of the expert model are selected, the more data feature types the expert network can adapt to. The extreme case is that a network only learns one type of feature data.

[0169] However, each child node participating in distributed training needs to align its selection of the output results of the N expert models, otherwise the performance of the global model aggregated by the central node will deteriorate. In Figures 9A-9E, the global model includes a gated network model. When the child nodes train the expert models, they also train the gated network model. The gated network model can include different types of neural networks, or the same type of neural network but with different parameter configurations, such as the number of layers, depth, and specific configurations including kernel size. Therefore, the gated network model has N outputs, and the N outputs of the gated network model are respectively connected to the outputs of the N expert models. The output results of the gated network model can be used to select the output results of the expert models, and the output results of the gated network model correspond one-to-one with the output results of the expert models to which it is connected. Therefore, to align the selection results of the expert models by each child node, the central node sends first information to the child nodes. This first information is used to indicate the selection parameters used to select the output results of the gated network model. The child nodes can determine whether to retain or discard the output results of the gated network model based on the selection parameters. If the output result of the gated network model is retained according to the selection parameters, the output result of the corresponding expert model is also retained; if the output result of the gated network model is discarded according to the selection parameters, the output result of the corresponding expert model is also discarded.

[0170] Exemplarily, the output result of the gated network model is a value between [0, 1] after being processed by the set function.

[0171] Exemplarily, the selection parameter is a first threshold. For example, the value range of the first threshold can be (0, 1). If the output result of the gated network model is greater than the first threshold, the output result of the gated network model is retained; if the output result of the gated network model is less than the first threshold, the output result of the gated network model is discarded. In particular, if the output result of the gated network model is equal to the first threshold, it can be agreed that the output result of the gated network model is retained or discarded.

[0172] Furthermore, in addition to being used to indicate selection parameters, the first information can also be used to indicate at least one of the following information: the identification of the N expert models, the competition mode of the N expert models, the training task, and the type of input and output of the global model.

[0173] The first information is used to indicate the identifiers of the N expert models, so that the model parameter information of the expert models reported by the subsequent child nodes is based on the output results of which expert models.

[0174] For N expert models, the central node can further instruct the child nodes on how to train each expert model. Generally, the central node will instruct the child nodes on the purpose of training the network. For example, in channel recovery, the goal is to make the recovered channel (as the output of the central node) as close to the true value as possible (that is, to minimize the normalized mean squared error (NMSE)). However, under the architecture of the hybrid expert model, even based on the same task (such as making the recovered channel close to the true value), there may still be at least two methods. The main difference between these two methods is the competitive mode of the N expert models, where the competitive mode can also be called the usage mode, cooperation mode, collaborative mode, and indication information of whether to cooperate, etc.:

[0175] Method 1: N expert models collaborate: loss = || A - ∑ i p i o i || 2 .

[0176] Method 2: N expert models compete with each other: loss = ∑ i p i ||Ao i || 2 .

[0177] Among them, in the above two formulas, A represents the true value (which can represent a single sample), p i represents the weight of the i-th expert model (determined based on the output of the gating network model), o irepresents the output of the i-th expert model. As can be seen, in the collaborative approach, multiple expert models collaborate to approximate the true value; in the competitive approach, each expert model has its own loss, enabling independent judgment without relying on the results of other expert models. Even if the same training node trains the hybrid expert model using different competitive approaches, the resulting gating weights and expert weights will differ. Therefore, it is desirable that all participating subnodes use / train the expert models in the same manner.

[0178] Therefore, the central node may carry the competition mode of the N experts in the first information, where the competition mode indicates whether the N expert models cooperate or compete with each other.

[0179] The central node also indicates the training task of the distributed training. The training task can be understood as the function of the global model, that is, what the global model can be used to train. For example, the training task can be beam prediction, channel recovery, channel prediction, etc.

[0180] In the beam management scenario, the input of the global model can be, for example, channel quality, RSRP of the beam, etc.; the output of the global model can be, for example, the optimal beam index, etc.

[0181] In addition, the central node can also send model parameter information of the global model to the child nodes.

[0182] S1003. The subnode inputs sample data to N expert models and the gated network model to obtain N first output results of the N expert models and N second output results of the gated network model, where the N second output results correspond one-to-one to the N first output results.

[0183] When training a child node, sample data needs to be input into N expert models and the gating network model. Therefore, before executing this step, the child node collects different types of sample data in different application scenarios.

[0184] For example, in a beam management scenario, the child nodes in a federated learning architecture can be network devices, while the central node can be an independent federated learning management node; or the child nodes can be terminal devices, while the central node can be a network device acting as a central node. Assuming that the global model to be trained is an AI / ML model that takes estimated channel measurements or the received signal itself as input and outputs the optimal beam index, then during the data collection phase, the child nodes are responsible for collecting the channel measurements or received signals used as model inputs and the labels used for training the model, namely the optimal beam index. All possible beams (codebook-based synchronization signal blocks (SSBs) or channel state information-reference signal (CSI-RS) beams) can be sent one by one to the terminal device through the network device. The terminal device selects the beam direction index with the best performance (the best can refer to the beam with the largest physical layer-reference signal received power (L1-RSRP) or signal-to-noise ratio (SNR) measurement value among all SSB / CSI-RS beams) as the label.

[0185] Before model training, the central node can also configure downlink resources for the child nodes to transmit the central node's initial global model information. The downlink resources can be control channel resources, such as PDCCH resources, or data channel resources, such as PDSCH resources. Exemplarily, the downlink resources include parameters such as the frequency domain resource block number, starting position, subband number, subband bandwidth, frequency hopping parameters, and modulation and coding scheme (MCS).

[0186] The global model can be distributed from the central node via broadcast or multicast. For example, in a single-cell federated learning architecture where the central node is a network device and the child nodes are terminal devices, the global model can be distributed via broadcast. Due to the nature of broadcast, child nodes not participating in the federated learning can also receive the broadcast information. In a multi-cell federated learning architecture where a network device with federated learning management functions serves as the central node and other network devices serve as child nodes, the central node can also distribute the global model to each child node via broadcast. Similarly, other child nodes not participating in the federated learning can also receive the broadcast information. Alternatively, multicast can be used for child nodes participating in the federated learning. Child nodes associated with the same central node are grouped together, sharing the same group number and configured with the same downlink resources. In multicast mode, child nodes not participating in the federated learning will not receive the multicast information.

[0187] The central node can also configure uplink resources for subnodes to report local models, which are used by subnodes to report models, gradients, and weights. Another federated learning management node can also configure uplink resources for the central node and subnodes for reporting local model information and necessary signaling. Similar to downlink resource configuration, uplink resources can be control channel resources, such as PUCCH resources, or data channel resources, such as PUSCH resources.

[0188] After collecting a certain amount of sample data, the child node inputs the sample data into N expert models and the gated network model, obtaining N first output results of the N expert models and N second output results of the gated network model. Because the N outputs of the gated network model are connected to the outputs of the N expert models, the N second output results correspond one-to-one with the N first output results.

[0189] S1004. The child node selects, for each first output result among the N first output results and each second output result among the N second output results, a first output result corresponding to the second output result that satisfies the selection parameter.

[0190] For each second output result of the gating network model, the child node determines whether to retain or discard the second output result based on a selection parameter. For example, if the selection parameter is a first threshold, the child node determines whether the second output result is greater than or equal to the first threshold. If so, the second output result is retained; otherwise, the second output result is discarded.

[0191] After the at least one retained second output result is determined, at least one first output result corresponding to the at least one retained second output result is also retained or selected.

[0192] S1005. The subnode obtains model parameter information of the selected at least one expert model and model parameter information of the gated network model based on the selected at least one first output result and the true value information.

[0193] After the subnode selects at least one first output result, the subnode obtains model parameter information of the selected at least one expert model and model parameter information of the gated network model based on the selected at least one first output result and the true value information.

[0194] The model parameter information of the expert model includes the weight, gradient, gradient change, etc. of the expert model. The model parameter information of the gated network model includes the weight, gradient, gradient change, etc. of the gated network model.

[0195] Exemplarily, the child node obtains model parameter information of at least one selected expert model and model parameter information of the gated network model based on the selected at least one first output result and the true value information. There are two possible implementation methods:

[0196] In one possible implementation, the child node obtains model parameter information for at least one selected expert model and model parameter information for the gated network model based on the average and true value information of at least one selected first output result. For example, if the first threshold is 0.5, the child node may retain the second output results of the gated network model whose value is greater than 0.5 and select at least one first output result of the expert model corresponding to the retained at least one second output result. The child node then multiplies each of the at least one selected first output results by 0.5, sums and averages these products, and obtains the average value of the at least one selected first output result.

[0197] Another possible implementation is that the child node weights and averages the at least one selected first output result based on the at least one second output result corresponding to each of the at least one selected first output results to obtain a weighted average of the at least one selected first output result, and obtains model parameter information of the at least one selected expert model and the model parameter information of the gated network model based on the weighted average and true value information of the at least one selected first output result. For example, if the first threshold is 0.5, the child node may retain the second output results of the gated network model whose value is greater than 0.5 and select at least one first output result of the expert model corresponding to the retained at least one second output result. The child node multiplies each of the at least one selected first output results by its corresponding second output result (for example, some second output results are 0.6, and some second output results are 0.8), sums and averages these products, and obtains the weighted average of the at least one selected first output result.

[0198] S1006. The child node sends the second information to the central node. Correspondingly, the central node receives the second information.

[0199] After the child node obtains the model parameter information of the selected at least one expert model and the model parameter information of the gated network model, it sends second information to the central node, wherein the second information is used to indicate the model parameter information of the gated network model and the model parameter information of at least one expert model selected from the N expert models.

[0200] Furthermore, the second information is also used to indicate the identifier of the selected at least one expert model. The child node carries the indication information of the identifier of the selected at least one expert model in the second information, so that the central node can perform model aggregation based on the indication information.

[0201] S1007. The central node updates the global model and the selection parameters based on the plurality of second information received from the plurality of child nodes.

[0202] The central node receives multiple pieces of second information from multiple child nodes respectively, and can update the global model based on the multiple pieces of second information.

[0203] In some scenarios, the central node can further update selection parameters based on the updated global model. For example, during initial training, the central node may configure a higher first threshold, which can be updated based on information such as training feedback from child nodes. This first threshold can also be derived from neural network learning.

[0204] S1008. The central node sends the fourth information, and the child node receives the fourth information accordingly.

[0205] If the central node updates the selection parameters, it can send fourth information to each child node, where the fourth information is used to indicate the updated selection parameters, so that the fourth information selects the output result of the gated network model based on the updated selection parameters during the next round of training.

[0206] The distributed training system may execute steps S1002 to S1008 multiple times until the global model converges.

[0207] In one example, the global model consists of a general layer, Net_com, three expert models, Net1, Net2, and Net3, and a fully connected network, Net_FC. A central node (e.g., a base station) sends the global model's model parameters, including Net_com, Net1, Net2, and Net_FC, to each child node. The corresponding neural network gradients and weights are represented by W_com, W_1, W_2, and W_3, and W_FC.

[0208] Assume that only two child nodes participate in this round of federated learning: UE1 and UE2. Their respective training processes are the same. Local sample data is input to the model and the following training is performed:

[0209] i) The outputs after the general layer and Net_1 / 2 / 3 are W_1_out, W_2_out, and W_3_out respectively. The output of the gated Net_FC is W_FC (a vector consisting of three [0,1] elements). After fusion, the final output is: Out = W_FC(1)*W_1_out + W_FC(2)*W_2_out + W_FC(3)*W_3_out;

[0210] ii) Calculate the loss based on, for example, the loss function formula loss = ||Out–A||^2, where A is the true value / label;

[0211] iii) After training, the following weights are obtained: UE1: W_com(UE1), W_1(UE1) / W_2(UE1) / W_3(UE1) and W_FC(UE1); UE2: W_com(UE2), W_1(UE2) / W_2(UE2) / W_3(UE2) and W_FC(UE2).

[0212] The child nodes (UE1, UE2) feed back the above weight value (which may also be a gradient or a gradient change) to the central node (base station).

[0213] After the central node obtains the model parameter information from the two child nodes, it corresponds to the same global model and begins to fuse the two sets of parameters. Here, we take the averaging method as an example to obtain the following:

[0214] ⅰ) Fusion results of general layer

[0215] ii) Fusion results corresponding to the three expert models

[0216] iii) Fusion results of the gated network model

[0217] After completing this round of federated learning, the central node sends new model parameter information to the child nodes for the next round of training.

[0218] As shown in Figure 11, it is a schematic diagram of the optimal beam training of an example embodiment of the present application. Taking the high-frequency beam management problem as an example, when the network device side performs beam scanning of the synchronization signal block or CSI-RS based on the codebook, the channels between users (UE1~UEN) at different locations (which may also include training nodes with similar data characteristics to these UEs) and the network device are different. The network device sends a unified global model to users at different locations. Users at different locations measure the received SSB or CSI-RS beam, such as measuring L1-RSRP, and feedback the beam identifier corresponding to the maximum RSRP value. AI / ML can be used to train a model, using, for example, a plurality of received SSB / CSI-RS signals (part or all of them), or the strength (RSRP) of a plurality of received SSB / CSI-RS signals (part or all of them) or the estimated channel as input to infer the optimal beam ID and feed it back to the network device.

[0219] The global model includes a gated network model and N expert models, and may also include other layers (i.e., the general layers described above, which may or may not be trained). The other layers are connected to the gated network model and the N expert models. The gated network model has N outputs, each of which is connected to the output of an expert model. During each user's training process, each user inputs the measured RSRP into other layers, and the trained outputs of the other layers serve as common inputs for the gated network model and the N expert models. Each user trains the gated network model and the N expert models based on these inputs. Before training, the network device sends selection parameters to each user. Each user determines whether to retain the N second output results of the gated network model based on the selection parameters and, based on at least one retained second output result, obtains the first output result of at least one expert model corresponding to the at least one retained second output result. Then, based on the selected at least one first output result and the true value information, each user obtains model parameter information for the selected at least one expert model and the gated network model, and sends the model parameter information for the selected at least one expert model and the gated network model to the network device. Furthermore, each user obtains an optimal beam identifier based on the selected at least one first output result and the true value information.

[0220] Optionally, the central node is a third-party device that performs the aforementioned central node-related actions. For example, the above steps S1001-S1002 and S1006-S1008 are all performed by a third-party device.

[0221] Optionally, the child node is a third-party device that performs the aforementioned actions related to the child node. For example, the above steps S1001-S1008 are all performed by a third-party device.

[0222] Optionally, the central node is a network device. In this case, the network device can complete the training of the model. For example, the above steps S1001-S1002 and S1006-S1008 are all performed by the network device.

[0223] Optionally, the child node is a terminal device. In this case, the terminal device can complete the training of the model. For example, the above steps S1001-S1008 are all performed by the terminal device.

[0224] Optionally, the central node includes a network device and a third-party device. In one example, step S1007 can be performed by a third-party device, such as an OTT or cloud server, and one or more of steps S1001-S1002, S1006, and S1008 can be performed by the network device. Furthermore, the network device and the third-party device can communicate with each other to transmit the content transmitted in one or more of steps S1001-S1002, S1006-S1008.

[0225] Optionally, the sub-nodes include a terminal device and a third-party device. In one example, one or more of steps S1003-S1005 may be performed by a third-party device, such as an OTT or cloud server, and one or more of steps S1001-S1002, S1006, and S1008 may be performed by the terminal device. Furthermore, the terminal device and the third-party device may communicate with each other to transmit the content transmitted in one or more of steps S1001-S1008.

[0226] According to a distributed training method provided in an embodiment of the present application, the central node indicates a selection parameter so that all child nodes participating in the training will apply the selection parameter to the output result of the gated network model during the training process of the gated network model and the expert model, thereby enabling each child node to align the selection of the output result of the gated network model and improve the performance of the aggregation model of the central node.

[0227] In this application, "sending information to... (e.g., a child node)" or the related illustrations in the accompanying drawings can be understood as the destination end of the information being a child node. This can include sending information to a child node directly or indirectly. "Receiving information from... (e.g., a child node)" or "receiving information from... (e.g., a child node)", or the related illustrations in the accompanying drawings can be understood as the source end of the information being a child node, which can include receiving information from a child node directly or indirectly. The information may be processed as necessary between the source end and the destination end of the information transmission, such as format changes, etc., but the destination end can understand the valid information from the source end. Similar expressions in this application can be understood similarly and will not be repeated here.

[0228] The above mainly introduces the solutions provided by the embodiments of the present application from the perspective of the interaction between various nodes. Accordingly, the embodiments of the present application also provide a distributed training device, which is used to implement the various methods described above. The distributed training device can be the central node in the above method embodiments, or a component that can be used for the central node; or, the distributed training device can be a sub-node in the above method embodiments, or a component that can be used for the sub-node. It is understandable that in order to implement the above functions, the distributed training device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0229] The embodiments of the present application can divide the functional modules of the distributed training device according to the above-mentioned method embodiments. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing unit. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical functional division. In actual implementation, other division methods can be used.

[0230] Based on the same concept of the above-mentioned distributed training method, this application also provides the following distributed training device:

[0231] FIG12 is a schematic diagram of the structure of a distributed training device provided in an embodiment of the present application. The distributed training device 1200 includes a transceiver unit 1201 and a processing unit 1202.

[0232] When the distributed training device is used to implement the functions of the subnode in the above method embodiment, the transceiver unit 1201 is used to perform one or more of the operations of the subnode in steps S1001, S1002, S1006, and S1008 of the embodiment shown in Figure 10, and the processing unit 1202 is used to perform one or more of steps S1003-S1005 of the embodiment shown in Figure 10. Optionally, the distributed training device can be a terminal device, or a third-party device such as an OTT or cloud server, or can be a system composed of a terminal device and a third-party device.

[0233] When the distributed training device is used to implement the functions of the central node in the above method embodiment, the transceiver unit 1201 is used to perform one or more of the operations of the central node in steps S1001, S1002, S1006, and S1008 of the embodiment shown in Figure 10, and the processing unit 1202 is used to perform step S1007 of the embodiment shown in Figure 10. Optionally, the distributed training device can be a network device, or a third-party device such as an OTT or cloud server, or can be a system composed of a network device and a third-party device.

[0234] For the specific implementation of the above-mentioned transceiver unit 1201 and the processing unit 1202, reference may be made to the description in the above-mentioned method embodiment.

[0235] In addition, it should be noted that the aforementioned transceiver unit and / or processing unit can be implemented through virtual modules. For example, the processing unit can be implemented through a software function unit or a virtual device, and the transceiver unit can be implemented through a software function or a virtual device. Alternatively, the processing unit or transceiver unit can also be implemented through physical circuits. For example, if the device is implemented using a chip / chip circuit, the transceiver unit can be an input / output circuit and / or a communication interface to perform input operations (corresponding to the aforementioned receiving operations) and output operations (corresponding to the aforementioned sending operations); the processing unit is a processing circuit, such as an integrated processor or microprocessor or integrated circuit.

[0236] The division of modules in this application is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the examples of this application may be integrated into a single processor, exist physically as separate modules, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in either hardware or software functional modules.

[0237] As shown in Figure 13, it is a structural diagram of another distributed training device provided in an embodiment of the present application. The distributed training device 1300 includes one or more processing circuits 1301 (one processing circuit is illustrated in the figure). Optionally, the distributed training device 1300 may also include a memory 1303 (indicated by a dotted line in the figure). The memory 1303 is used to store instructions executed by the processing circuit 1301, or to store input data required for the processing circuit 1301 to run instructions, or to store data generated after the processing circuit 1301 runs instructions. Optionally, the distributed training device 1300 may also include an interface circuit 1302 (indicated by a dotted line in the figure), and the processing circuit 1301 and the interface circuit 1302 are coupled to each other. It will be understood that the interface circuit 1302 can be a transceiver or an input-output interface.

[0238] The processing circuit may be a processor or a circuit in a processor used for processing.

[0239] When the distributed training device is used to implement the functions of the sub-node in the above-mentioned method embodiment, the interface circuit 1302 is used to execute one or more of the operations of the sub-node in steps S1001, S1002, S1006 and S1008 of the embodiment shown in Figure 10, and the processing circuit 1301 is used to execute one or more of steps S1003-S1005 of the embodiment shown in Figure 10.

[0240] When the distributed training device is used to implement the function of the central node in the above method embodiment, the interface circuit 1302 is used to execute one or more of the operations of the central node in steps S1001, S1002, S1006 and S1008 of the embodiment shown in Figure 10, and the processing circuit 1301 is used to execute step S1007 of the embodiment shown in Figure 10.

[0241] When the above-mentioned distributed training device is a chip applied to the central node, the chip implements the function of the central node in the above-mentioned method embodiment. The chip receives information from other modules in the central node, and the information is sent by the child node to the central node; or, the chip sends information to other modules in the central node, and the information is sent by the central node to the child node. When the central node is a network device, the module of the central node here can be the baseband chip of the central node, or it can be a CU, DU or other module, or it can be a device under the open radio access network (O-RAN) architecture, such as an open CU, open DU and other devices. When the central node is a third-party device, the module of the child node here can be a processing chip of the third-party device. Among them, the processing chip can be used to implement AI training.

[0242] When the above-mentioned distributed training device is a chip applied to a sub-node, the chip implements the functions of the sub-node in the above-mentioned method embodiment. The chip receives information from other modules in the sub-node, and the information is sent by the central node to the sub-node; or, the chip sends information to other modules in the sub-node, and the information is sent by the sub-node to the central node. When the sub-node is a terminal device, the module of the sub-node here can be the baseband chip of the sub-node, or, the baseband chip and the processing chip. Among them, the processing chip can be used to implement AI training. When the sub-node is a third-party device, the module of the sub-node here can be the processing chip of the third-party device. Among them, the processing chip can be used to implement AI training.

[0243] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed, the method in the above embodiment is implemented.

[0244] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method in the above embodiment.

[0245] An embodiment of the present application also provides a distributed training system, including the above-mentioned distributed training device.

[0246] The present application also provides a circuit, which is coupled to a memory and is used to execute the method shown in the above embodiment. The circuit may include a chip circuit.

[0247] Optionally, an embodiment of the present application further provides a chip system, comprising: at least one processor and an interface, wherein the at least one processor is coupled to a memory via the interface, and when the at least one processor executes a computer program or instruction in the memory, the chip system executes the method in any of the above method embodiments. Optionally, the chip system may be composed of a chip, or may include a chip and other discrete devices, which is not specifically limited in the embodiments of the present application.

[0248] The memory in the present application may also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data. A memory is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. For example, the memory may be a non-volatile memory, such as a digital versatile disc (DVD), a hard disk drive (HDD), or a solid-state drive (SSD), or a volatile memory, such as a random-access memory (RAM).

[0249] As used in the following description of this application, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0250] It should be understood that in the description of this application, unless otherwise specified, " / " indicates that the objects associated with each other are in an "or" relationship. For example, A / B can mean A or B; where A and B can be singular or plural. Also, in the description of this application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural. In addition, to facilitate the clear description of the technical solutions of the embodiments of this application, in the embodiments of this application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean different. At the same time, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0251] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0252] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art can understand and implement other changes to the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0253] It is understood that the various numbers used in the embodiments of this application are merely for ease of description and are not intended to limit the scope of the embodiments of this application. The order of the sequence numbers of the above-mentioned processes does not necessarily imply a specific order of execution; the order of execution of the processes should be determined by their functions and inherent logic.

[0254] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0255] The components in the device of the embodiment of the present application can be merged, divided, or deleted according to actual needs. Those skilled in the art can combine or combine the different embodiments and features of the different embodiments described in this specification.

[0256] In this application, under the premise of no logical contradiction, the examples can reference each other, for example, the methods and / or terms between method embodiments can reference each other, for example, the functions and / or terms between device embodiments can reference each other, for example, the functions and / or terms between device examples and method examples can reference each other.

Claims

1. A distributed training method, characterized in that, The method is applied to a distributed training system. The global model of the distributed training system includes N expert models and a gating network model, where N is a positive integer. The method includes: Receiving first information, where the first information is used to indicate a selection parameter, and the selection parameter is used to select the output result of the gating network model; Sending second information, where the second information is used to indicate the model parameter information of the gating network model and the model parameter information of at least one selected expert model among the N expert models, and the at least one selected expert model is obtained based on the output result of the gating network model.

2. The method according to claim 1, wherein The method further includes: Inputting sample data into the N expert models and the gating network model to obtain N first output results of the N expert models and N second output results of the gating network model, where the N second output results respectively correspond one-to-one to the N first output results; For each of the N first output results and each of the N second output results, for the second output results that meet the selection parameter, selecting the first output result corresponding to the second output result; Based on at least one selected first output result and truth value information, obtaining the model parameter information of at least one selected expert model and the model parameter information of the gating network model.

3. The method according to claim 1 or 2, characterized in that, The selection parameter is a first threshold.

4. The method according to any one of claims 1 to 3, characterized in that, The first information is further used to indicate at least one of the following information: the identifiers of the N expert models, the competition and cooperation mode of the N expert models, the training task, the types of the input and output of the global model.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Sending third information, where the third information is used to indicate at least one of the following information: the memory space size of the sub-node, the computing power information of the sub-node, whether it supports model training, and the types of models supported for training.

6. The method according to any one of claims 1-5, characterized in that, The second information is further used to indicate the identifiers of at least one selected expert model.

7. The method according to any one of claims 1-4, characterized in that, The obtaining the model parameter information of at least one selected expert model and the model parameter information of the gating network model based on at least one selected first output result and truth value information includes any one of the following operations: Based on the average value of at least one selected first output result and truth value information, obtaining the model parameter information of at least one selected expert model and the model parameter information of the gating network model; Or Based on at least one second output result corresponding to at least one selected first output result, performing weighted sum and average on at least one selected first output result to obtain the weighted average value of at least one selected first output result, and based on the weighted average value of at least one selected first output result and truth value information, obtaining the model parameter information of at least one selected expert model and the model parameter information of the gating network model.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Receiving fourth information, where the fourth information is used to indicate an updated selection parameter, and the updated selection parameter is obtained based on the second information.

9. A distributed training method, characterized in that, The method is applied to a distributed training system. The global model of the distributed training system includes N expert models and a gating network model, where N is a positive integer. The method includes: Sending first information to multiple child nodes, where the first information is used to indicate selection parameters for selecting the output result of the gating network model; Receiving multiple second information respectively, where each of the multiple second information is used to indicate the model parameter information of the gating network model and the model parameter information of at least one selected expert model among the N expert models, and the at least one selected expert model is obtained based on the output result of the gating network model.

10. The method according to claim 9, wherein, The selection parameter is a first threshold.

11. The method according to claim 9 or 10, characterized in that The first information is further used to indicate at least one of the following information: the identifiers of the N expert models, the competition and cooperation method of the N expert models, the training task, and the types of the input and output of the global model.

12. The method according to any one of claims 9-11, characterized in that, The method further includes: Receiving third information, where the third information is used to indicate at least one of the following information: the memory space size of the child node, the computing power information of the child node, whether it supports model training, and the types of models that support training.

13. The method according to any one of claims 9-12, characterized in that, The second information is further used to indicate the identifiers of the at least one selected expert model.

14. The method according to any one of claims 9 - 13, characterized in that, The method further includes: Updating the global model based on the multiple second information.

15. The method according to any one of claims 9-14, characterized in that, The method further includes: Updating the selection parameter based on the multiple second information; Sending fourth information, where the fourth information is used to indicate the updated selection parameter.

16. A distributed training device, characterized in that, Comprising a unit for executing the method according to any one of claims 1-8, or comprising a unit for executing the method according to any one of claims 9-15.

17. A distributed training system, characterized in that, The system includes a first distributed training device and a second distributed training device. The first distributed training device is used to implement the method according to any one of claims 1-8, and the second distributed training device is used to implement the method according to any one of claims 9-15.

18. A distributed training device, characterized in that, Comprising a processing circuit and an interface circuit. The interface circuit is used to receive signals from other distributed training devices outside the distributed training device and transmit them to the processing circuit, or send signals from the processing circuit to other distributed training devices outside the distributed training device. The processing circuit uses a logic circuit or executes code instructions to implement the method according to any one of claims 1-8, or implement the method according to any one of claims 9-15.

19. The distributed training device according to claim 18, wherein The distributed training device is a chip.

20. A chip module, characterized in that, Comprising a transceiver component and a chip, where the chip is used to execute the method according to any one of claims 1-8, or execute the method according to any one of claims 9-15.

21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-8, or implements the method according to any one of claims 9-15.

Citation Information

Patent Citations

  • Personalized federated learning method based on hybrid expert model

    CN112560991A

  • Deep learning model training method, target object detection method and device

    CN115906921A

  • Privacy preserving collaborative learning with domain adaptation

    US20210073677A1