Model training method and device
Patent Information
- Application Number
- CN202380081589.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-12
- Publication Date
- 2025-07-04
AI Technical Summary
In the federated learning algorithm, a large amount of model/gradient data needs to be exchanged between wireless network nodes and the central server, resulting in huge data transmission pressure. And due to node heterogeneity, nodes with poor performance will affect the overall training progress. How to reduce data Transmitting pressure and increasing training speed and efficiency become urgent issues to be solved.
By implementing intelligent flow training of models in a node set, each node updates the model and sends the updated model to the next-hop node, which is not a central server, reducing data transmission pressure and dynamically adapting to node heterogeneity. , improve training speed and efficiency.
It achieves efficient model training in wireless networks, reduces data transmission pressure, reduces transmission overhead and management and control complexity, dynamically adapts to node heterogeneity, improves training speed and efficiency, and protects data privacy.
Smart Images

Figure CN120266138A_ABST
Abstract
Description
Model training method and device Technical Field
[0001] The embodiments of the present application relate to the field of communication technology, and in particular to a model training method and device. Background Art
[0002] With the continuous development of communication technology, attempts have been made to combine artificial intelligence (AI) technology with communication networks to realize model training and reasoning through communication networks.
[0003] For example, a federated learning algorithm can be used to train the model, that is, the model can be distributed to each network node through a central server, and each network node performs model training and updates, and uploads the updated model / gradient data to the central server for aggregation without uploading the original data to protect data privacy.
[0004] However, federated learning algorithms require the exchange of large amounts of model and gradient data between network nodes and central servers. As the scale of these models and gradient data grows, wireless network data transmission faces significant pressure. Furthermore, network nodes at all levels in wireless networks are highly heterogeneous, with significant variations in computing power, memory, and transmission bandwidth. Poorly performing nodes can hinder overall training progress.
[0005] Therefore, how to use various network nodes to train the model to reduce data transmission pressure and improve training speed and efficiency has become a technical problem that needs to be solved urgently.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a model training method and device, which can reduce data transmission pressure and improve training speed and efficiency when using various network nodes to train the model.
[0008] In a first aspect, an embodiment of the present application provides a model training method, which may include: a first node updates an acquired first model to obtain an updated first model, and sends the updated first model to a next-hop node; wherein the first node is any node in a node set, and the node set is used to train the first model; the updated first model converges on the first node; and the next-hop node is a node in the node set.
[0009] Based on the first aspect, when training the first model, a certain node in the node set can be used to train the first model, and the updated first model can be sent to another node in the node set to realize the intelligent flow training of the first model between the nodes in the node set, without being limited to a single node to train the first model, so that each node can obtain the updated training results of the first model by other nodes. At the same time, since the next hop node of each node is a node in the node set, not a central server, the data transmission pressure can be reduced, the transmission overhead can be reduced, and the complexity of management and control can be reduced. Since each node sends the updated first model to the next hop node, rather than the local original data, data privacy can be protected. The embodiment of the present application can also dynamically adapt to the heterogeneity of each node to improve the training speed and efficiency.
[0010] In one possible design, the first node updates the first model to obtain an updated first model, including: the first node determines activation parameters based on the first model, updates the activation parameters, and obtains the updated first model; wherein the activation parameters are part or all of the parameters of the first model.
[0011] Based on this possible design, when the first node updates the first model, it can selectively update some parameters of the first model and freeze the remaining parameters (i.e., not update the remaining parameters). Alternatively, the first node can also update all parameters of the first model without restriction.
[0012] In one possible design, the first node determines the activation parameters based on the first model, including: the first node determines the activation parameters based on one or more of the following: data characteristics of the first node, computing power of the first node, and update status of parameters of the first model.
[0013] In one possible design, the activation parameter is a parameter in the first model whose correlation with the data of the first node is greater than or equal to a preset threshold; or, the activation parameter is a parameter in the first model that has not been updated; or, the activation parameter is any one or more parameters in the first model.
[0014] Based on the above two possible designs, the first node can determine the parameters with strong correlation as activation parameters according to the data characteristics of the local original data and the training objectives. Alternatively, the first node can also determine the parameters that have not been updated in the first model as activation parameters, so that the first model can be completely traversed as soon as possible, while reducing the impact on the update results of other nodes (such as the nodes traversed by the first model before the first node). Alternatively, the first node can also use randomness to randomly select one or more parameters in the first model as activation parameters to break the problem of excessive weight on a certain factor in a fixed mode, and it is relatively simple to implement and does not require the collection of additional information. Provide multiple feasible solutions for the first node to determine the activation parameters.
[0015] In one possible design, the first node determines the next-hop node based on the node information of each node in the node set; wherein the node information includes one or more of the following: first indication information, data characteristics, computing power information, and channel status information; the first indication information is used to indicate whether the node has been traversed.
[0016] In one possible design, the next-hop node is a node in the node set that has not been traversed; or, the next-hop node is the node in the node set that has the strongest correlation with the data of the first node; or, the next-hop node is the node in the node set that is closest to the first node; or, the next-hop node is the node in the node set that has the highest connection power with the first node; or, the next-hop node is the node in the node set that has the strongest computing power; or, the next-hop node is any node in the node set.
[0017] Based on the above two possible designs, multiple feasible solutions are provided for the first node to determine the next hop node. At the same time, the first node determines the next hop node by itself, adopting a completely self-organizing approach, which can reduce the complexity of management and control.
[0018] In one possible design, the first node sends the updated first model to the next-hop node, including: if the first condition is not met, the first node sends the updated first model to the next-hop node; wherein the first condition is that the number of times the first node is traversed is greater than or equal to a preset round, or the first condition is that the model prediction accuracy of the first model is greater than or equal to a preset accuracy.
[0019] Based on this possible design, when the first condition is not met, the first node can send the updated first model to the next-hop node to continue training the first model. If the first condition is met, the model training process can be exited to complete the training of the first model.
[0020] In one possible design, each node in the node set is used to update the first model in each round corresponding to a preset round.
[0021] Based on this possible design, in each round of traversal, each node in the node set may update the first model and send the updated first model to the next hop node until the node set is completely traversed.
[0022] In one possible design, the first node sends the updated first model to the next-hop node, including: the first node sends the updated first model to multiple next-hop nodes.
[0023] Based on this possible design, when the first node sends the updated first model to the next-hop node, it can send the updated first model to multiple next-hop nodes to obtain multiple final training results of the first model, increase the degree of parallelism, and at the same time, enable each node to obtain the transfer of part of the knowledge, achieving better results than independent training.
[0024] In the second aspect, an embodiment of the present application provides a communication device, which can be applied to the first node in the above-mentioned first aspect or the possible design of the first aspect to implement the function performed by the above-mentioned first node. The communication device can be the first node, or it can be a chip or system on chip for implementing the function of the first node. The communication device can implement the function performed by the above-mentioned first node by executing the corresponding software through hardware. The hardware or software includes one or more modules corresponding to the above-mentioned functions. For example, a transceiver module and a processing module. The transceiver module is used to obtain the first model; the processing module is used to update the first model to obtain the updated first model; the transceiver module is also used to send the updated first model to the next hop node; wherein the updated first model converges on the first node; the first node is any node in the node set, the node set is used to train the first model, and the next hop node is a node in the node set.
[0025] In one possible design, the processing module is specifically used to: determine activation parameters based on the first model, update the activation parameters, and obtain an updated first model; wherein the activation parameters are part or all of the parameters of the first model.
[0026] In one possible design, the processing module is specifically used to determine the activation parameters based on one or more of the following: data characteristics of the first node, computing power of the first node, and update status of parameters of the first model.
[0027] In one possible design, the activation parameter is a parameter in the first model whose correlation with the data of the first node is greater than or equal to a preset threshold; or, the activation parameter is a parameter in the first model that has not been updated; or, the activation parameter is any one or more parameters in the first model.
[0028] In one possible design, the processing module is also used to determine the next hop node based on the node information of each node in the node set; wherein the node information includes one or more of the following: first indication information, data characteristics, computing power information, channel status information; the first indication information is used to indicate whether the node has been traversed.
[0029] In one possible design, the next-hop node is a node in the node set that has not been traversed; or, the next-hop node is the node in the node set that has the strongest correlation with the data of the first node; or, the next-hop node is the node in the node set that is closest to the first node; or, the next-hop node is the node in the node set that has the highest connection power with the first node; or, the next-hop node is the node in the node set that has the strongest computing power; or, the next-hop node is any node in the node set.
[0030] In one possible design, the transceiver module is specifically used to send the updated first model to the next hop node if the first condition is not met; wherein the first condition is that the number of times the first node is traversed is greater than or equal to a preset round, or the first condition is that the model prediction accuracy of the first model is greater than or equal to a preset accuracy.
[0031] In one possible design, each node in the node set is used to update the first model in each round corresponding to a preset round.
[0032] In one possible design, the transceiver module is also used to send the updated first model to multiple next-hop nodes.
[0033] It should be noted that the specific implementation method of the communication device in the second aspect can refer to the behavioral function of the first node in the model training method provided by the first aspect or any possible design of the first aspect.
[0034] In a third aspect, an embodiment of the present application provides a communication device, which includes one or more processors; one or more processors are used to run computer programs or instructions, and when the one or more processors execute the computer instructions or instructions, the communication device executes the model training method described in the first aspect.
[0035] In one possible design, the communication device further includes one or more memories, the one or more memories being coupled to one or more processors, and the one or more memories being used to store the above-mentioned computer programs or instructions. In one possible implementation, the memory is located outside the communication device. In another possible implementation, the memory is located within the communication device. In an embodiment of the present application, the processor and the memory may also be integrated into one device, that is, the processor and the memory may also be integrated together. In one possible implementation, the communication device further includes a transceiver, and the transceiver is used to receive information and / or send information.
[0036] In one possible design, the communication device further includes one or more communication interfaces, the one or more communication interfaces are coupled to one or more processors, and the one or more communication interfaces are used to communicate with other modules outside the communication device.
[0037] In a fourth aspect, an embodiment of the present application provides a communication device, which includes an input and output interface and a logic circuit; the input and output interface is used to input and / or output information; the logic circuit is used to execute the model training method described in the first aspect, and process and / or generate information based on the information.
[0038] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions or programs. When the computer instructions or programs are run on a computer, the model training method described in the first aspect is executed.
[0039] In a sixth aspect, an embodiment of the present application provides a computer program product comprising computer instructions, which, when run on a computer, enables the model training method described in the first aspect to be executed.
[0040] In a seventh aspect, an embodiment of the present application provides a computer program, which, when run on a computer, enables the model training method described in the first aspect to be executed.
[0041] Among them, the technical effects brought about by any design method in the third to seventh aspects can refer to the technical effects brought about by the above-mentioned first aspect.
[0042] In an eighth aspect, an embodiment of the present application provides a communication system, which may include: a first node and a next-hop node of the first node; the first node is used to obtain a first model, update the first model, and obtain an updated first model; wherein the first node is any node in a node set, and the nodes in the node set are used to train the first model; the updated first model converges on the first node; the first node is also used to send the updated first model to the next-hop node; wherein the next-hop node is a node in the node set; the next-hop node of the first node is used to receive the updated first model from the first node.
[0043] In one possible design, the first node is specifically used to: determine activation parameters based on the first model; wherein the activation parameters are part or all of the parameters of the first model; and update the activation parameters to obtain an updated first model.
[0044] In one possible design, the first node is specifically used to determine activation parameters based on one or more of the following: data characteristics of the first node, computing power of the first node, and update status of parameters of the first model.
[0045] In one possible design, the activation parameter is a parameter in the first model whose correlation with the data of the first node is greater than or equal to a preset threshold; or, the activation parameter is a parameter in the first model that has not been updated; or, the activation parameter is any one or more parameters in the first model.
[0046] In one possible design, the first node is also used to determine the next hop node based on the node information of each node in the node set; wherein the node information includes one or more of the following: first indication information, data characteristics, computing power information, and channel status information; the first indication information is used to indicate whether the node has been traversed.
[0047] In one possible design, the next-hop node is a node in the node set that has not been traversed; or, the next-hop node is the node in the node set that has the strongest correlation with the data of the first node; or, the next-hop node is the node in the node set that is closest to the first node; or, the next-hop node is the node in the node set that has the highest connection power with the first node; or, the next-hop node is the node in the node set that has the strongest computing power; or, the next-hop node is any node in the node set.
[0048] In one possible design, the first node is specifically used to: send the updated first model to the next hop node if the first condition is not met; wherein the first condition is that the number of times the first node is traversed is greater than or equal to a preset round, or the first condition is that the model prediction accuracy of the first model is greater than or equal to a preset accuracy.
[0049] In one possible design, each node in the node set is used to update the first model in each round corresponding to a preset round.
[0050] In one possible design, the first node is specifically used to send the updated first model to multiple next-hop nodes.
[0051] Among them, the technical effects brought about by any design method in the eighth aspect can refer to the technical effects brought about by the above-mentioned first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] FIG1a is a schematic diagram of a communication system provided in an embodiment of the present application;
[0053] FIG1b is a schematic diagram of a communication system provided in an embodiment of the present application;
[0054] FIG2 is a schematic diagram of the composition of a communication device provided in an embodiment of the present application;
[0055] FIG3 is a schematic diagram of a model training method provided in an embodiment of the present application;
[0056] FIG4 is a flow chart of a model training method provided in an embodiment of the present application;
[0057] FIG5 is a flow chart of a model training method provided in an embodiment of the present application;
[0058] FIG6 is a schematic diagram of task parallelism provided in an embodiment of the present application;
[0059] FIG7 is a schematic diagram of a model training method provided in an embodiment of the present application;
[0060] FIG8 is a schematic diagram of a communication device provided in an embodiment of the present application;
[0061] FIG9 is a structural diagram of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] With the continuous development of communication technology, attempts have been made to combine big data artificial intelligence (AI) technology with communication networks (such as federated learning between network data analytics functions (NWDAF)) to achieve model training and reasoning through communication networks.
[0063] For example, taking the federated learning algorithm as an example, the model can be distributed to each network node (or described as data node, node, etc.) through the central server, and each network node will train and update the model, and upload the updated model / gradient data to the central server for aggregation without uploading the original data to protect data privacy.
[0064] However, federated learning algorithms require the exchange of large amounts of model / gradient data between network nodes and central servers. As the scale of model / gradient data grows (for example, the VGG16 model is 552MB in size, the Vision Transformer is 1337MB+, and the number of parameters has even exceeded one trillion), wireless network data transmission faces tremendous pressure.
[0065] In addition, unlike Internet technology (IT) networks, wireless networks have strong heterogeneity between network nodes at all levels, with large differences in computing power, memory, and transmission bandwidth. This leads to a serious straggler problem when applying federated learning algorithms, that is, network nodes with poor performance execute tasks slowly, affecting the overall federation progress.
[0066] Therefore, how to use various network nodes to train the model to reduce data transmission pressure and improve training speed and efficiency has become a technical problem that needs to be solved urgently.
[0067] In order to solve the above technical problems, an embodiment of the present application provides a model training method, in which the first node can update the acquired first model to obtain an updated first model, and send the updated first model to the next hop node; wherein the first node is any node in the node set, and the node set is used to train the first model; the updated first model converges on the first node; and the next hop node is a node in the node set.
[0068] The model training method provided in the embodiment of the present application is a distributed learning method that is more suitable for use in wireless networks. When training the first model, a node in the node set can be used to train the first model, and the updated first model can be sent to another node in the node set to realize intelligent flow training of the first model between the nodes in the node set, without being limited to a single node to train the first model, so that each node can obtain the updated training results of the first model by other nodes. At the same time, since the next hop node of each node is a node in the node set, not a central server, it can reduce the pressure on data transmission, reduce transmission overhead, and reduce the complexity of management and control. Since each node sends the updated first model to the next hop node, rather than the local original data, data privacy can be protected. The embodiment of the present application can also dynamically adapt to the heterogeneity of each node to improve training speed and efficiency.
[0069] The following describes in detail the implementation of the embodiments of the present application in conjunction with the accompanying drawings.
[0070] The model training method provided in the embodiment of the present application can be used for any communication system, which can be a third generation partnership project (3GPP) communication system, for example, a long term evolution (LTE) system, or a fifth generation (5G) mobile communication system, a new radio (NR) communication system, a 5G-NR communication system, a new radio vehicle to everything (NR V2X) system, and can also be applied to a system in which LTE and 5G are hybrid networks, or a non-terrestrial network (NTN) system, a device-to-device (D2D) communication system, a machine-to-machine (M2M) communication system, an Internet of Things (IoT), an IT system, and other next-generation communication systems, such as future communication systems such as 6G, and can also be non-3GPP communication systems without limitation.
[0071] The embodiments of the present application can also be applied to one or more of the following business scenarios: enhanced mobile broadband (eMBB), ultra-reliable and low latency communication (URLLC), machine-type communication (MTC), massive machine-type communication (mMTC), narrowband Internet of Things (NB-IoT), customer premise equipment (CPE), augmented reality / virtual reality (AR / VR), V2X, etc., without limitation.
[0072] eMBB refers to mobile broadband services that further enhance user experience and other performance, and is the application scenario most relevant to daily life. The most intuitive experience brought by 5G in this regard is the significant increase in network speed, with peak speeds reaching 10Gbps even for watching 4K HD video. For example, eMBB can refer to high-traffic mobile broadband services such as 3D and ultra-HD video.
[0073] URLLC features high reliability, low latency, and extremely high availability. It can be used in a variety of scenarios and applications, including industrial applications and control, traffic safety and control, remote manufacturing, remote training, and remote surgery. URLLC has great potential in autonomous driving services and is also crucial for the security industry. For example, URLLC can be used for services such as autonomous driving and industrial automation that require low-latency, highly reliable connections.
[0074] MTC, also known as M2M, has the advantages of low cost and enhanced coverage.
[0075] NB-IoT boasts wide coverage, multiple connections, low speed, low cost, low power consumption, and an optimized architecture. For example, it offers massive connections, lower power consumption, and lower chip costs. Applications include smart water meters, smart parking, smart pet tracking, smart bicycles, smart smoke detectors, smart toilets, and smart vending machines.
[0076] CPE is a mobile signal access device that receives mobile signals and forwards them as wireless-fidelity (Wi-Fi) signals. It also converts high-speed 4G or 5G signals into Wi-Fi signals, supporting a large number of mobile terminals accessing the internet simultaneously. CPE is widely used for wireless network access in rural areas, towns, hospitals, workplaces, factories, and residential areas, saving the cost of laying wired networks.
[0077] V2X is a key technology for future intelligent transportation systems. It enables communication between vehicles, between vehicles and base stations, and between base stations. This provides real-time traffic information, including road conditions, pedestrian information, and more. This improves driving safety, reduces congestion, increases traffic efficiency, and provides in-vehicle entertainment.
[0078] The following describes the communication system provided in an embodiment of the present application using Figure 1a as an example.
[0079] Figure 1a is a schematic diagram of a communication system provided in an embodiment of the present application. As shown in Figure 1a, the communication system may include multiple nodes (or described as network nodes, data nodes, devices, communication devices, equipment, communication equipment, etc.).
[0080] The nodes in FIG1a may be devices capable of training and updating the model.
[0081] For example, taking the AI model as an example, each node in Figure 1a can be a device with AI computing capabilities.
[0082] For example, each node may include an AI module, and each node may implement AI model reasoning through the AI module.
[0083] Optionally, each node in FIG1a may be any of the following: terminal equipment, access network equipment, core network equipment, server, etc., without limitation.
[0084] The terminal device can be a device with wireless transceiver capabilities or a chip or chip system that can be installed in the device. It can allow users to access the network and is used to provide voice and / or data connectivity to users. The terminal device can be vehicle-mounted, portable, or handheld. The terminal device and the user can be completely independent. All user-related information can be stored in a smart card (SIM card), which can be used on the terminal device. The terminal device can complete direct air interface interaction with the access network device. The terminal device can send and / or receive signals.
[0085] The terminal device may also be referred to as user equipment (UE), subscriber unit (subscriber unit), terminal, mobile station (MS), or mobile terminal (MT). Specifically, the terminal device may be a cellular phone, a smart phone, a wireless data card, a mobile phone, a personal digital assistant (PDA), a tablet computer or a computer with wireless transceiver function, a wireless modem, a handheld device (handset), or a laptop computer. The terminal device may also be a VR terminal, an AR terminal, a wireless terminal in industrial control, a wireless terminal in unmanned driving, a wireless terminal in telemedicine, a wireless terminal in smart grids, a wireless terminal in smart cities, a wireless terminal in smart homes, an MTC terminal, an in-vehicle terminal, a vehicle with vehicle-to-vehicle (V2V) communication capability, an intelligent connected vehicle, a drone with UAV to UAV (U2U) communication capability, and the like, without limitation.
[0086] Access network equipment can be any device deployed in the access network that can wirelessly communicate with terminal devices. It is responsible for all functions related to the air interface: wireless physical control functions, resource scheduling, wireless access control, wireless link maintenance, wireless resource management, mobility management, and other functions. The wireless link maintenance function is to maintain the wireless link with the terminal device, and is responsible for protocol conversion for wireless link data and IP data quality monitoring. Wireless resource management functions include establishing and releasing wireless links, scheduling and allocating wireless resources, etc. Mobility management functions include configuring terminal devices for measurement, evaluating terminal device wireless link quality, and making decisions about terminal device handovers between cells.
[0087] Exemplarily, the access network device may be an access network (AN) / radio access network (RAN) device, which is composed of multiple AN / RAN nodes. The AN / RAN node may be: an access point (AP), a base station (nodeB, NB), a macro base station, a micro base station (or described as a small station), a relay station, an enhanced nodeB (eNB), a next generation eNB (ng-eNB), a next generation nodeB (gNB), a transmission reception point (TRP), a transmission point (TP), a transmission measurement function (TMF), a wearable device, an in-vehicle device, or some other access node, etc., without limitation.
[0088] The access network device may also be a centralized unit (CU) / distributed unit (DU) architecture, in which case the access network device may include two network elements, CU and DU; the access network device may also be a control plane-user plane (CP-UP) architecture, in which case the access network device may include three network elements, namely the CU control plane (CU-CP), the CU user plane (CU-UP), and the DU, without limitation. The access network device may also include a remote unit (RU). In different systems, CU (or CU-CP and CU-UP), DU, or RU may also have different names, but those skilled in the art can understand their meanings. For example, in an open radio access network (open RAN, ORAN) system, CU may also be referred to as O-CU (open CU), DU may also be referred to as O-DU, CU-CP may also be referred to as O-CU-CP, CU-UP may also be referred to as O-CU-UP, and RU may also be referred to as O-RU. For convenience of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples for description. Any of the CU (or CU-CP, CU-UP), DU and RU in this application can be implemented through a software module, a hardware module, or a combination of a software module and a hardware module.
[0089] Optionally, when the access network device is a CU (or O-CU, CU-CP, CU-UP) or DU (or O-DU) or RU (or O-RU), the CU or DU or RU can perform all the sending and receiving operations performed by the node (such as the first node, other nodes, etc.) in the embodiments shown in Figures 3 to 7 below, and / or other processes for supporting the technology described herein; the CU or DU or RU can also be used to perform all operations other than the sending and receiving operations performed by the node in the embodiments shown in Figures 3 to 7 below, and / or other processes for supporting the technology described herein, without limitation.
[0090] Optionally, when the access network device includes CU, DU and RU, the DU or RU may perform all the transceiver operations performed by the node (such as the first node, or other nodes) in the embodiments shown in Figures 3 to 7 below, and / or other processes for supporting the technology described herein; the CU or DU may perform all the operations except the transceiver operations performed by the node in the embodiments shown in Figures 3 to 7 below, and / or other processes for supporting the technology described herein, without limitation.
[0091] Among them, the core network equipment is mainly responsible for providing user connections, user management, and business carrying, and provides an interface to the external network as a bearer network.
[0092] Exemplarily, the core network device may include a mobility management network element, a session management network element, a user plane network element, a session management network element, and other network elements, without limitation.
[0093] The server may be deployed in a data network, which may be an operator network that provides data transmission services to users, such as an operator network that provides Internet protocol multimedia services (IMS) to users, etc., without limitation.
[0094] Optionally, as shown in Figure 1b, the nodes in the communication system can be connected through an interface (e.g., NG, Xn) or an air interface. These nodes, such as core network equipment, access network equipment, terminal equipment, or one or more devices in operation administration and maintenance (OAM) are provided with one or more AI modules (for clarity, only one is shown in Figure 1b). The access network equipment can be a separate RAN node, or it can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be provided with one or more AI modules. Optionally, the CU can also be split into CU-CP and CU-UP. One or more AI modules are provided in the CU-CP and / or CU-UP.
[0095] The AI module is used to implement the corresponding AI function. The AI modules deployed in different nodes can be the same or different. The model of the AI module can implement different functions according to different parameter configurations. The model of the AI module can be configured based on one or more of the following parameters: structural parameters (such as the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of the neuron, the activation function of the neuron, or at least one of the bias in the activation function), input parameters (such as the type of input parameters and / or the dimension of the input parameters), or output parameters (such as the type of output parameters and / or the dimension of the output parameters). Among them, the bias in the activation function can also be called the bias of the neural network.
[0096] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or on the same node or device.
[0097] It should be noted that each node in the embodiments of the present application can be one or more chips, or a system on chip (SOC), etc. Figure 1a is only an exemplary figure, and the number of devices included is not limited. In addition, in addition to the devices shown in Figure 1a, the communication system may also include other devices. The names of the various devices and the names of the various links in Figure 1a are not limited. In addition to the names shown in Figure 1a, the various devices and links can also be named other names without limitation.
[0098] In a specific implementation, each node shown in Figure 1a may adopt the structure shown in Figure 2, or include the components shown in Figure 2. Figure 2 is a schematic diagram of the structure of a communication device 200 provided in an embodiment of the present application. The communication device 200 may be a node or a chip or system-on-chip in a node. As shown in Figure 2, the communication device 200 includes a processor 201, a transceiver 202, and a communication circuit 203.
[0099] Furthermore, the communication device 200 may further include a memory 204 , wherein the processor 201 , the memory 204 and the transceiver 202 may be connected via a communication line 203 .
[0100] The processor 201 is a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 201 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.
[0101] Transceiver 202 is used to communicate with other devices or other communication networks. Such other communication networks may be Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. Transceiver 202 may be a module, circuit, transceiver, or any device capable of communication.
[0102] The communication line 203 is used to transmit information between the components included in the communication device 200.
[0103] The memory 204 is used to store instructions, where the instructions may be computer programs.
[0104] The memory 204 may be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions, or a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, etc., without limitation.
[0105] It should be noted that the memory 204 can exist independently of the processor 201 or can be integrated with the processor 201. The memory 204 can be used to store instructions, program code, or some data. The memory 204 can be located within the communication device 200 or outside the communication device 200, without limitation. The processor 201 is used to execute the instructions stored in the memory 204 to implement the model training method provided in the following embodiments of this application.
[0106] In an example, the processor 201 may include one or more CPUs, such as CPU0 and CPU1 in FIG. 2 .
[0107] As an optional implementation, the communication device 200 includes multiple processors. For example, in addition to the processor 201 in FIG. 2 , it may also include a processor 207 .
[0108] As an optional implementation, the communication apparatus 200 further includes an output device 205 and an input device 206. For example, the input device 206 is a keyboard, a mouse, a microphone, a joystick, or the like, and the output device 205 is a display screen, a speaker, or the like.
[0109] It should be noted that the communication device 200 may be a desktop computer, a portable computer, a network server, a mobile phone, a tablet computer, a wireless terminal, an embedded device, a chip system, or a device having a structure similar to that shown in FIG2 . Furthermore, the structure shown in FIG2 does not limit the communication device. In addition to the components shown in FIG2 , the communication device may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0110] In the embodiment of the present application, the chip system can be composed of chips, or can include chips and other discrete devices.
[0111] In addition, the actions and terms involved in the various embodiments of this application can refer to each other without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are only examples, and other names can also be used in specific implementations without limitation.
[0112] In conjunction with the communication system shown in FIG. 1 a , taking the training of the first model as an example, a node set for training the first model can be determined from various nodes in the communication system.
[0113] The first model may be any model to be trained, the node set may include multiple nodes for training the first model, and the node set may also be called a collaboration set.
[0114] Optionally, the node set used to train the first model is determined based on data type, business type, computing power, etc.
[0115] For example, taking the first model as a model trained based on data of business A, multiple nodes executing business A can be determined as a node set used to train the first model.
[0116] The size of the node set determines the algorithm performance, which is related to the computing power of each node in the node set, data distribution characteristics, etc. For example, the stronger the computing power of each node, the stronger the algorithm performance.
[0117] Optionally, when each node in the node set trains the first model, as shown in Figure 3, the first model can be routed sequentially between the nodes in the node set, each node updates part or all of the parameters of the first model based on local original data, and sends the updated first model to the next hop node.
[0118] Exemplarily, referring to Figure 4 below, taking the first node as an example, the process of each node updating the first model and sending the updated first model to the next hop node is described in detail, wherein the first node can be any node in the node set, that is, any node in the node set can train the first model with reference to the method described in Figure 4 below, and send the updated first model to the next hop node.
[0119] FIG4 is a flow chart of a model training method provided in an embodiment of the present application. As shown in FIG4 , the method may include:
[0120] Step 401: The first node obtains a first model.
[0121] The first node is any node in the node set, and the node set is used to train the first model.
[0122] Optionally, if the first node is the first node to train the first model, the first node may obtain the first model from operations management and maintenance (OAM), or the first model may be pre-set in the first node. If the first node is not the first node to train the first model, the first node may obtain the first model from the previous hop node of the first node. That is, the first model obtained by the first node may be the first model updated by the previous hop node.
[0123] Step 402: The first node updates the first model to obtain an updated first model.
[0124] The first node updating the first model can also be described as the first node training the first model, and the updated first model can also be described as the trained first model.
[0125] Among them, the updated first model converges on the first node, or it can also be described as the updated first model is in a convergent state on the first node, or the updated first model is a converged model on the first node, or the first node trains the first model to a converged state, etc., without limitation.
[0126] Optionally, the local original data of the first node is used as a data set, and the data set is randomly divided into a training set, a validation set and a test set. The training set is used to update (or described as training) the first model, the validation set is used to verify the first model, the first model is continuously adjusted according to the verification results, and the final first model is evaluated with the test set to obtain an updated first model that has reached convergence.
[0127] For example, when the accuracy of the test set reaches a preset accuracy, it can be considered that the first model has reached convergence.
[0128] For example, taking the preset accuracy of 95% as an example, when the first node updates the first model, if the accuracy of the test set reaches 95%, the first node can consider that the updated first model has reached a convergence state, thereby stopping the update of the first model and outputting the updated first model.
[0129] Optionally, when the first node updates the first model, it determines activation parameters, updates the activation parameters, and obtains an updated first model.
[0130] The activation parameters may be part or all of the parameters of the first model, and the activation parameters may also be described as weight parameters.
[0131] Specifically, the first node may selectively update some parameters of the first model and freeze the remaining parameters (or describe as not updating the remaining parameters), or the first node may also update all parameters of the first model without restriction.
[0132] Exemplarily, the first node may determine the activation parameters based on one or more of the following: data characteristics of the first node, computing power of the first node, and update status of parameters of the first model.
[0133] In a first possible implementation, the activation parameter may be a parameter of the first model whose correlation with the data of the first node is greater than or equal to a preset threshold.
[0134] Among them, when updating the first model, not all parameters will be updated. The first node can determine the key parameters (or describe it as determining the parameters related to its own data) based on the data characteristics of the local original data and the training objectives, and determine the key parameters as activation parameters.
[0135] For example, the first node may determine a parameter in the first model whose correlation with the data of the first node is greater than or equal to a preset threshold as an activation parameter.
[0136] Optionally, the correlation between the parameters of the first model and the data of the first node is determined based on the information entropy.
[0137] Exemplarily, the information entropy may be Fisher information entropy.
[0138] In a second possible implementation, the activation parameters are parameters that have not been updated in the first model.
[0139] The first node can determine the active parameters based on the update status of the parameters of the first model. Specifically, if one or more parameters of the first model have already been updated by other nodes, the first node can freeze those parameters and determine the unupdated parameters as the active parameters. This allows the first model to be fully traversed as quickly as possible while minimizing the impact on the update results of other nodes (e.g., nodes traversed by the first model before the first node).
[0140] In a third possible implementation, the activation parameter is any one or more parameters in the first model.
[0141] Among them, the first node can also use randomness to randomly select one or more parameters in the first model as activation parameters to break the problem of excessive weight on a certain factor in a fixed mode. It is relatively simple to implement and does not require the collection of additional information (such as data characteristics, model update status, etc.).
[0142] In a fourth possible implementation, the first node determines the activation parameter based on the computing power of the first node.
[0143] When the first node has strong computing power, it can select more parameters as activation parameters. When the first node has weak computing power, it can select fewer parameters as activation parameters. This allows for flexible adaptation based on the first node's own computing resources (or dynamic adaptation to the heterogeneity of each node), thereby avoiding the straggler problem and improving training speed and efficiency.
[0144] The computing power may refer to the computing processing capability of the first node. The stronger the computing processing capability, the stronger the computing power.
[0145] For example, the computing power can be measured according to the number of CPUs included in the first node. For example, the larger the number of CPUs included, the stronger the computing power, and the smaller the number of CPUs included, the weaker the computing power.
[0146] Optionally, the first node may determine the activation parameter according to one or more of the first to fourth possible implementation manners described above, without limitation.
[0147] For example, the first node may select, from the parameters of the first model that have not been updated, parameters whose correlation with the data of the first node is greater than or equal to a preset threshold value and determine them as activation parameters.
[0148] Optionally, when the first node updates the activation parameter, a gradient may be calculated for the activation parameter to update the activation parameter.
[0149] Optionally, the first node inputs the data sample (ie, local original data) into the first model for forward propagation to obtain a loss function.
[0150] The loss function may also be called a target loss function, an objective function, etc., and is used to evaluate the degree to which the predicted value of the first model differs from the true value. The better the loss function, the better the performance of the first model.
[0151] Optionally, the first node updates the activation parameters according to an anti-catastrophic forgetting algorithm to prevent the first node's update of the first model from overwriting the update result of the previous node on the first model, causing catastrophic forgetting.
[0152] For example, when updating activation parameters, some regularization terms may be added to avoid excessively large updates to parameters that are heavily associated with old tasks (such as previous node updates to the first model).
[0153] For example, when updating activation parameters, an elastic weight consolidation (EWC) algorithm can be used to avoid catastrophic forgetting.
[0154] Optionally, when the first node updates the first model, the first model may be updated according to the following formula (1):
[0155] M n+1 =f(M n ,D n+1 ); Formula (1)
[0156] Where M represents the first model, D represents the local original data, and f represents the update function, which can be an update function such as continuous learning, distillation, or aggregation without restriction.
[0157] Step 403: The first node sends the updated first model to the next-hop node.
[0158] The next hop node may be a node in the node set.
[0159] In one possible design, the first node determines the next hop node based on a pre-planned unified path.
[0160] After the node set is determined, the training path of the first model can be planned uniformly based on the node information of each node in the node set.
[0161] For example, the training paths may be uniformly planned by nodes in a node set, or by control devices in a network, or by developers, without limitation.
[0162] In another possible design, the first node determines the next-hop node on its own, using a completely self-organizing approach to reduce management and control complexity.
[0163] Exemplarily, the first node may determine the next hop node according to the node information of each node in the node set.
[0164] Among them, the node information may include one or more of the following: first indication information, data characteristics, computing power information, and channel status information; the first indication information is used to indicate whether the node is traversed.
[0165] In a first possible implementation, the next hop node is a node in the node set that has not been traversed.
[0166] The first node may determine whether each node has been traversed according to the first indication information, and select a node that has not been traversed as the next hop node, so that the first model can be completely traversed in the node set as quickly as possible.
[0167] In a second possible implementation, the next hop node is a node in the node set that has the strongest correlation with the data of the first node.
[0168] Among them, since the first model is trained sequentially, the difference in data features between the previous and next nodes will affect the convergence effect of the first model. Improper selection of the next-hop node may cause oscillation in the model convergence direction. Therefore, when selecting the next-hop node, the correlation between the data features of each node and the data features of the first node can be fully measured, and the node with the strongest correlation can be selected as the next-hop node.
[0169] For example, the following formula (2) and formula (3) can be used to calculate the distance between data distributions using KL divergence to characterize the correlation between data:
[0170] q=Min q D(p||q); Formula (2)
[0171] Where p represents the sample distribution of the first node, and q represents the sample distribution of the other nodes. A larger KL divergence indicates a greater difference between the two; a smaller KL divergence indicates a smaller difference between the two.
[0172] In a third possible implementation, the next hop node is a node in the node set that is closest to the first node.
[0173] In consideration of data transmission delay and energy consumption, the next hop node may be determined as the node in the node set that is closest to the first node, so as to reduce data transmission delay and power consumption.
[0174] The distance can be the communication transmission distance between nodes, which can be determined based on the transmission delay. For example, the shorter the transmission delay, the shorter the communication transmission distance, indicating a closer proximity to the first node. In other words, the node closest to the first node can also be described as the node with the shortest transmission delay to the first node.
[0175] Exemplarily, taking the first node as a base station as an example, the next hop node of the first node may be a neighboring station.
[0176] In a fourth possible implementation, the next hop node is a node in the node set that has the greatest connection power with the first node.
[0177] Among them, considering the data transmission delay and energy consumption, the next hop node can be determined as the node with the largest connection power to the first node in the node set according to the channel state information, so as to reduce the data transmission delay and power consumption.
[0178] Exemplarily, the first node may determine, based on the channel state information, the node with the best channel quality as the node in the node set with the greatest connection power to the first node.
[0179] In the fifth possible implementation, the next hop node is the node with the strongest computing power in the node set.
[0180] Among them, each node in the node set can update some parameters of the first model according to its own computing power. Therefore, the first node can select the node with the strongest computing power as the next hop node to train more parameters faster.
[0181] In a sixth possible implementation, the next hop node is any node in the node set.
[0182] Among them, the first node can also use randomness to randomly select a node from the node set as the next hop node to break the problem of excessive weight on a certain factor in a fixed mode. It is relatively simple to implement and does not require the collection of additional information.
[0183] Optionally, the first node may determine the next hop node based on one or more of the first to sixth possible implementations. That is, the first node may comprehensively consider the multiple factors mentioned in the first to sixth possible implementations and perform optimization under multiple factor parameters.
[0184] Exemplarily, the first node can adopt mathematical modeling to mathematically analyze the impact of various factors (such as whether it has been traversed, correlation, node distance, node computing power, connection power, and one or more of the randomly selected nodes) on the final model training, so as to perform optimization solution with strong interpretability.
[0185] In another example, the first node can also adopt AI modeling, taking multiple factors (such as whether it has been traversed, correlation, node distance, node computing power, connection power, and one or more of the randomly selected nodes) as input features, using deep learning or reinforcement learning for modeling, and learning a routing solution, which is relatively simple to implement.
[0186] Based on the method shown in Figure 4 above, when training the first model, a certain node in the node set can be used to train the first model, and the updated first model can be sent to another node in the node set, so as to realize intelligent flow training of the first model between the nodes in the node set, without being limited to training the first model on a single node, so that each node can obtain the updated training results of the first model by other nodes.
[0187] Furthermore, because the next hop for each node is a node in the node set, not a central server, frequent uploads and downloads of model / gradient data in federated learning algorithms are avoided, significantly reducing communication overhead and pressure. Decentralized training eliminates the performance bottlenecks and security risks of central servers. If any node in the node set experiences a problem, it can be replaced at any time, ensuring uninterrupted training. It can also dynamically adapt to the heterogeneity of individual nodes, improving training speed and efficiency. This reduces management and control complexity.
[0188] In addition, a sequential routing approach is adopted to conduct model training among distributed nodes. Each node sends the updated first model to the next hop node instead of the local original data, which can protect data privacy.
[0189] Based on the above description, optionally, when training the first model, each node in the node set may perform one or more rounds of traversal on the first model.
[0190] Optionally, in each round of traversal, each node in the node set participates in updating the first model, that is, each node in the node set is used to update the first model in each round of traversal.
[0191] Exemplarily, as shown in FIG5 , in each round of traversal, each node in the node set may execute the method shown in FIG4 above, update the first model, and send the updated first model to the next hop node until the node set is completely traversed.
[0192] Optionally, in each round of traversal, the training path of the first model between nodes may be different.
[0193] Optionally, when the first condition is not met, each node can send the updated first model to the next-hop node to continue training the first model. If the first condition is met, the model training process can be exited to complete the training of the first model.
[0194] Exemplarily, the first condition may be that the number of times the node is traversed is greater than or equal to a preset round.
[0195] Among them, when each node does not meet the first condition, it sends the updated first model to the next hop node to continue training the first model. It can be replaced by the description that when the number of times the node is traversed is less than or equal to the preset rounds, each node sends the updated first model to the next hop node to continue training the first model.
[0196] Taking the first node as an example, the first condition may be that the number of times the first node is traversed is greater than or equal to a preset round.
[0197] In another example, the first condition may be that the model prediction accuracy of the first model is greater than or equal to a preset accuracy.
[0198] Optionally, each node in the node set can train multiple first models simultaneously to achieve multi-task parallelism.
[0199] Among them, since more than one task may be running on each node in the network, compared with single-task parallelism, by using each node in the node set to train multiple first models at the same time, multi-task parallelism can be achieved to improve training speed and efficiency.
[0200] Single-task parallelism refers to each node executing the same task in the same time period and executing different tasks in different time periods. Multi-task parallelism refers to each node executing different tasks in the same time period.
[0201] For example, as shown in FIG6 , for single-task parallelism, nodes 1, 2, and 3 can execute task 1 in a first time period, task 2 in a second time period, and task 3 in a third time period. For multi-task parallelism, node 1 can execute task 1 (e.g., updating first model 1) in the first time period and send node 1's execution result of task 1 to node 2. Node 2 then executes task 1 in the second time period and sends node 2's execution result of task 1 to node 3. Node 3 then executes task 1 in the third time period. Node 2 can also execute task 2 (e.g., updating first model 2) in the first time period and send node 2's execution result of task 2 to node 3. Node 3 then executes task 2 in the second time period and sends node 3's execution result of task 2 to node 1. Node 1 then executes task 2 in the third time period. Node 3 can also execute task 3 (e.g., updating first model 3) in the first time period and send node 3's execution result of task 3 to node 1. Node 1 then executes task 3 in the second time period and sends node 1's execution result of task 3 to node 2. Node 2 then executes task 3 in the third time period.
[0202] Optionally, when each node sends the updated first model to the next-hop node, it can send the updated first model to multiple next-hop nodes to obtain multiple final training results of the first model, increase the degree of parallelism, and at the same time, enable each node to obtain the transfer of part of the knowledge, achieving better results than independent training.
[0203] The migration of partial knowledge may mean that each node can learn the update of the first model by the previous node based on the updated first model sent by the previous hop node.
[0204] Exemplarily, as shown in FIG7 , each node may send the updated first model to two next-hop nodes to obtain 8 training results.
[0205] Optionally, for multiple training results of the first model, the training result with the highest prediction accuracy can be selected as the final training result of the first model, or the multiple training results can be used as different training results of the first model on different training paths, or the multiple training results can be aggregated to obtain the final training result of the first model, without restriction.
[0206] It should be noted that the various methods provided in the embodiments of the present application can be implemented individually or in combination without limitation.
[0207] It is understood that in the embodiments of the present application, the execution subject may perform some or all of the steps in the embodiments of the present application. These steps or operations are merely examples, and the embodiments of the present application may also perform other operations or variations of various operations. In addition, the various steps may be performed in a different order than those presented in the embodiments of the present application, and it is possible that not all operations in the embodiments of the present application need to be performed.
[0208] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of interaction between devices. It is understandable that, in order to realize the above functions, each device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0209] The embodiments of the present application can divide the functional modules of each device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical function division. In actual implementation, there may be other division methods.
[0210] In the case of dividing each functional module according to each function, Figure 8 shows a communication device 80, which can execute the actions performed by the first node in the method shown in Figures 3 to 7 above. All relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module. The technical effects that can be obtained can be referred to the above method embodiment and will not be repeated here.
[0211] The communication device 80 may include a transceiver module 801 and a processing module 802. Exemplarily, the communication device 80 may be a communication device, or a chip used in a communication device, or other combined device, component, etc. having the functions of the above-mentioned communication device. When the communication device 80 is a communication device, the transceiver module 801 may be a transceiver, which may include an antenna and a radio frequency circuit, etc.; the processing module 802 may be a processor (or a processing circuit), such as a baseband processor, which may include one or more CPUs. When the communication device 80 is a component having the functions of the above-mentioned communication device, the transceiver module 801 may be a radio frequency unit; the processing module 802 may be a processor (or a processing circuit), such as a baseband processor. When the communication device 80 is a chip system, the transceiver module 801 may be the input and output interface of the chip (such as a baseband chip); the processing module 802 may be the processor (or a processing circuit) of the chip system, which may include one or more central processing units. It should be understood that the transceiver module 801 in the embodiment of the present application can be implemented by a transceiver or a transceiver-related circuit component; the processing module 802 can be implemented by a processor or a processor-related circuit component (or, referred to as a processing circuit).
[0212] For example, the transceiver module 801 can be used to perform all transceiver operations performed by the communication device in the embodiments shown in Figures 3 to 7, and / or to support other processes of the technology described herein; the processing module 802 can be used to perform all operations other than transceiver operations performed by the communication device in the embodiments shown in Figures 3 to 7, and / or to support other processes of the technology described herein.
[0213] As another possible implementation, the transceiver module 801 in FIG8 can be replaced by a transceiver that integrates the functionality of the transceiver module 801; the processing module 802 can be replaced by a processor that integrates the functionality of the processing module 802. Furthermore, the communication device 80 shown in FIG8 can also include a memory.
[0214] Alternatively, when the processing module 802 is replaced by a processor and the transceiver module 801 is replaced by a transceiver, the communication device 80 involved in the embodiment of the present application may also be the communication device 90 shown in Figure 9, wherein the processor may be the logic circuit 901 and the transceiver may be the interface circuit 902. Furthermore, the communication device 90 shown in Figure 9 may also include a memory 903.
[0215] The embodiments of the present application also provide a computer program product, which, when executed by a computer, can implement the functions of any of the above method embodiments.
[0216] The embodiments of the present application also provide a computer program, which, when executed by a computer, can implement the functions of any of the above method embodiments.
[0217] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above-mentioned method embodiments can be completed by a computer program to instruct the relevant hardware, and the program can be stored in the above-mentioned computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. The computer-readable storage medium can be an internal storage unit of the terminal (including the data sending end and / or the data receiving end) of any of the above-mentioned embodiments, such as the hard disk or memory of the terminal. The above-mentioned computer-readable storage medium can also be an external storage device of the above-mentioned terminal, such as a plug-in hard disk equipped on the above-mentioned terminal, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. Further, the above-mentioned computer-readable storage medium can also include both the internal storage unit of the above-mentioned terminal and an external storage device. The above-mentioned computer-readable storage medium is used to store the above-mentioned computer program and other programs and data required by the above-mentioned terminal. The above-mentioned computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0218] It should be noted that the terms "first" and "second" in the specification, claims and drawings of this application are used to distinguish different objects, rather than to describe a specific order. "First" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "multiple" means two or more.
[0219] Furthermore, the terms "include," "comprise," and "have," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0220] It should be understood that in this application, "at least one (item)" refers to one or more. "Multiple" refers to two or more. "At least two (items)" refers to two or three and more than three. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple. “When” and “if” both mean that corresponding measures will be taken under certain objective circumstances. They do not limit the time, nor do they require any judgment action when they are implemented, nor do they mean that there are other limitations.
[0221] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.
[0222] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0223] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0224] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0225] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0226] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
Claims
1. A model training method, characterized in that: include: The first node acquires a first model; wherein the first node is any node in a node set, and the node set is used to train the first model; The first node updates the first model to obtain an updated first model; wherein the updated first model converges on the first node; The first node sends the updated first model to a next-hop node; wherein the next-hop node is a node in the node set.
2. The method according to claim 1, characterized in that The first node updates the first model to obtain the updated first model, including: The first node determines activation parameters according to the first model; wherein the activation parameters are part or all of the parameters of the first model; The first node updates the activation parameter to obtain the updated first model.
3. The method according to claim 2, characterized in that The first node determines the activation parameter according to the first model, including: The first node determines the activation parameter according to one or more of the following: data characteristics of the first node, computing power of the first node, and an update status of parameters of the first model.
4. The method according to claim 2 or 3, characterized in that: The activation parameter is a parameter of the first model whose correlation with the data of the first node is greater than or equal to a preset threshold; or The activation parameters are parameters in the first model that have not been updated; or The activation parameter is any one or more parameters in the first model.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: The first node determines the next hop node based on the node information of each node in the node set; wherein the node information includes one or more of the following: first indication information, data characteristics, computing power information, channel status information; the first indication information is used to indicate whether the node has been traversed.
6. The method according to any one of claims 1 to 5, characterized in that: The next hop node is a node in the node set that has not been traversed; or The next hop node is a node in the node set that has the strongest correlation with the data of the first node; or The next hop node is a node in the node set that is closest to the first node; or The next hop node is a node in the node set that has the largest connection power with the first node; or The next hop node is the node with the strongest computing power in the node set; or The next hop node is any node in the node set.
7. The method according to any one of claims 1 to 6, characterized in that: The first node sending the updated first model to the next hop node includes: If the first condition is not met, the first node sends the updated first model to the next hop node; wherein the first condition is that the number of times the first node is traversed is greater than or equal to a preset round, or the first The condition is that the model prediction accuracy of the first model is greater than or equal to the preset accuracy.
8. The method according to claim 7, characterized in that Each node in the node set is used to update the first model in each round corresponding to the preset round.
9. The method according to any one of claims 1 to 8, characterized in that: The first node sending the updated first model to the next hop node includes: The first node sends the updated first model to the multiple next-hop nodes.
10. A communication device, characterized in that: include: A transceiver module, used for acquiring a first model; A processing module, used for updating the first model to obtain an updated first model; wherein the updated first model converges on a first node; the first node is any node in a node set, and the node set is used to train the first model; The transceiver module is further used to send the updated first model to a next-hop node; wherein the next-hop node is a node in the node set.
11. The device according to claim 10, characterized in that The processing module is specifically used for: Determine activation parameters according to the first model; wherein the activation parameters are part or all of the parameters of the first model; The activation parameters are updated to obtain the updated first model.
12. The device according to claim 11, characterized in that The processing module is specifically used to determine the activation parameters based on one or more of the following: data characteristics of the first node, computing power of the first node, and update status of parameters of the first model.
13. The device according to claim 11 or 12, characterized in that The activation parameter is a parameter of the first model whose correlation with the data of the first node is greater than or equal to a preset threshold; or The activation parameters are parameters in the first model that have not been updated; or The activation parameter is any one or more parameters in the first model.
14. The device according to any one of claims 10 to 13, characterized in that: The processing module is also used to determine the next hop node based on the node information of each node in the node set; wherein the node information includes one or more of the following: first indication information, data characteristics, computing power information, channel status information; the first indication information is used to indicate whether the node has been traversed.
15. The device according to any one of claims 10 to 14, characterized in that: The next hop node is a node in the node set that has not been traversed; or The next hop node is a node in the node set that has the strongest correlation with the data of the first node; or The next hop node is a node in the node set that is closest to the first node; or The next hop node is a node in the node set that has the largest connection power with the first node; or The next hop node is the node with the strongest computing power in the node set; or The next hop node is any node in the node set.
16. The device according to any one of claims 10 to 15, characterized in that: The transceiver module is specifically used for: If the first condition is not met, the updated first model is sent to the next hop node; wherein the first condition is that the number of times the first node is traversed is greater than or equal to a preset round, or the first condition is that the model prediction accuracy of the first model is greater than or equal to a preset accuracy.
17. The device according to claim 16, characterized in that Each node in the node set is used to update the first model in each round corresponding to the preset round.
18. The device according to any one of claims 10 to 17, characterized in that: The transceiver module is further used to send the updated first model to multiple next-hop nodes.
19. A communication device, characterized in that: The communication device includes a processor; the processor is used to run a computer program or instructions so that the communication device executes the model training method as described in any one of claims 1-9.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions or programs, and when the computer instructions or programs are run on a computer, the model training method as described in any one of claims 1 to 9 is executed.
21. A computer program product, characterized in thatThe computer program product includes computer instructions; when part or all of the computer instructions are run on a computer, the model training method according to any one of claims 1 to 9 is executed.
22. A communication system, characterized in that: include: A first node and a next-hop node of the first node; The first node is used to obtain a first model, update the first model, and obtain an updated first model; wherein the first node is any node in a node set, and the nodes in the node set are used to train the first model; the updated first model converges on the first node; The first node is further configured to send the updated first model to the next hop node; wherein the next hop node is a node in the node set; The next hop node of the first node is used to receive the updated first model from the first node.
23. The system according to claim 22, characterized in that The first node is specifically used for: Determine activation parameters according to the first model; wherein the activation parameters are part or all of the parameters of the first model; The activation parameters are updated to obtain the updated first model.
24. The system according to claim 23, characterized in that The first node is specifically used to determine the activation parameter based on one or more of the following: data characteristics of the first node, computing power of the first node, and update status of parameters of the first model.
25. The system according to claim 23 or 24, characterized in that The activation parameter is a parameter of the first model whose correlation with the data of the first node is greater than or equal to a preset threshold; or The activation parameters are parameters in the first model that have not been updated; or The activation parameter is any one or more parameters in the first model.
26. The system according to any one of claims 22 to 25, characterized in that: The first node is also used to determine the next hop node based on the node information of each node in the node set; wherein the node information includes one or more of the following: first indication information, data characteristics, computing power information, channel status information; the first indication information is used to indicate whether the node has been traversed.
27. The system according to any one of claims 22 to 26, characterized in that: The next hop node is a node in the node set that has not been traversed; or The next hop node is a node in the node set that has the strongest correlation with the data of the first node; or The next hop node is a node in the node set that is closest to the first node; or The next hop node is a node in the node set that has the largest connection power with the first node; or The next hop node is the node with the strongest computing power in the node set; or The next hop node is any node in the node set.
28. The system according to any one of claims 22 to 27, characterized in that: The first node is specifically used to send the updated first model to the next hop node if the first condition is not met; wherein the first condition is that the number of times the first node is traversed is greater than or equal to a preset round, or the first condition is that the model prediction accuracy of the first model is greater than or equal to a preset accuracy.
29. The system according to claim 28, characterized in that Each node in the node set is used to update the first model in each round corresponding to the preset round.
30. The system according to any one of claims 22 to 29, characterized in that: The first node is specifically configured to send the updated first model to the plurality of next-hop nodes.