Method, apparatus, system, device, medium and program product for generating a neural network model

By adjusting the hypernetwork model structure in federated learning, determining the subnetwork model structure and transmitting specific parameters of the subnetwork model, the federated learning efficiency and cost problems between multiple devices are solved, and efficient and accurate neural network model updates are achieved.

CN113570027BActive Publication Date: 2025-06-20HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110704382.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-24
Publication Date
2025-06-20
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

Existing federated learning technologies are difficult to efficiently share and update neural network models between multiple devices, resulting in reduced model accuracy and increased communication and computing costs.

Method used

By adjusting the structure of the hypernetwork model, the structure of the subnetwork model is determined, and specific parameters of the subnetwork model are transmitted between multiple devices, reducing communication costs, and generating a personalized subnetwork model on the device.

Benefits of technology

It improves the federated learning efficiency between multiple devices, improves model accuracy, and reduces communication and computing costs, meeting the resource constraints and data distribution characteristics of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113570027B_ABST
    Figure CN113570027B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, apparatus, system, device, medium, and program product for generating a neural network model. The method includes a first device sending an indication of the structure of a sub-network model to a second device, where the sub-network model is determined by adjusting the structure of a super-network model; the first device receiving parameters of the sub-network model from the second device, where the parameters of the sub-network model are determined by the second device based on the indication and the super-network model; the first device training the sub-network model based on the received parameters of the sub-network model; and the first device sending the parameters of the trained sub-network model to the second device for the second device to update the super-network model. By the above method, an efficient federated learning solution between multiple devices is provided, which improves the model accuracy while reducing the communication cost required for the federated learning process and the computing cost of the devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure mainly relate to the field of artificial intelligence. More specifically, embodiments of the present disclosure relate to methods, devices, systems, equipment, computer-readable storage media, and computer program products for generating neural network models. Background Art

[0002] In recent years, Federated Learning, as an emerging distributed machine learning technology, has been applied to many scenarios, such as mobile services, healthcare, and the Internet of Things. Federated Learning can make full use of the data and computing capabilities at the client side, enabling multiple parties to collaborate in building a general and more robust machine learning model without sharing data. In an environment where data supervision is becoming increasingly strict, Federated Learning can solve key issues such as data ownership, data privacy, and data access rights, and has great commercial value. Summary of the Invention

[0003] In view of the above, embodiments of the present disclosure propose a solution for generating a neural network model applicable to Federated Learning.

[0004] According to a first aspect of the present disclosure, there is provided a method for generating a neural network model, including: a first device sending an indication of the structure of a sub-network model to a second device, where the sub-network model is determined by adjusting the structure of a super-network model; the first device receiving parameters of the sub-network model from the second device, where the parameters of the sub-network model are determined by the second device based on the indication and the sub-network model; and the first device training the sub-network model based on the received parameters of the sub-network model. In this way, an efficient Federated Learning solution between multiple devices is provided, which not only improves the model accuracy but also reduces the communication cost required in the Federated Learning process and the computing cost of the devices.

[0005] In some embodiments, the method may further include the first device obtaining pre-configured parameters of the super-network model, and the first device determining the structure of the sub-network model by adjusting the structure of the super-network model based on the pre-configured parameters. In this way, a personalized sub-network model can be generated on the device.

[0006] In some embodiments, obtaining the pre-configured parameters of the super-network model may include: the first device generating local parameters by training the super-network model; the first device sending the local parameters to the second device; and the first device receiving the pre-configured parameters from the second device, where the pre-configured parameters are determined by the second device based at least on the local parameters received from the first device. In this way, optimized pre-configured parameters can be generated for the super-network model based on the data distribution of the device, which is beneficial to determining the structure of the sub-network model at a lower computing cost.

[0007] In some embodiments, determining the structure of the sub-network model may include: initializing the sub-network model as a super-network model with pre-configured parameters; and iteratively updating the sub-network model by performing the following operations at least once: adjusting the structures of multiple layers of the sub-network model to obtain multiple candidate network models; selecting one candidate network model to update the sub-network model based on the accuracies of the multiple candidate network models; and determining the structure of the sub-network model if the sub-network model meets the constraints of the first device. In this way, the device can simplify the super-network model with accuracy as a measure to obtain a personalized neural network model that meets its own resource limitations and data distribution.

[0008] In some embodiments, adjusting the structures of multiple layers may include: deleting the parameters related to a part of the nodes of one of the multiple layers to obtain one of the multiple candidate network models. In some embodiments, the method further includes: determining a part of the nodes based on a predetermined number ratio. In this way, the nodes and parameters that have less impact on the model accuracy can be removed, thereby ensuring the model accuracy while simplifying the model structure.

[0009] In some embodiments, the constraints include: the computational amount of the sub-network model is lower than a first threshold, or the number of parameters of the sub-network model is lower than a second threshold. In this way, the computational amount and the number of parameters of the sub-network model of the device can be reduced to meet the corresponding resource constraints.

[0010] In some embodiments, both the first threshold and the second threshold are associated with the performance of the first device. In this way, the corresponding resource constraints are set according to the device performance.

[0011] In some embodiments, the indication has a masked form, and the mask indicates whether the sub-network model has the corresponding parameters of the super-network model. In this way, during the federated learning process, only the specific parameters of the sub-network model need to be transmitted, thereby reducing the communication cost of transmission between devices.

[0012] In some embodiments, the method according to the first aspect may further include: the first device determines the change of the parameters by calculating the difference between the parameters of the trained sub-network model and the parameters received from the second device. In this way, during the federated learning process, only the parameters that have changed after training need to be transmitted, thereby reducing the communication cost of transmission between devices.

[0013] According to a second aspect of the present disclosure, a method for generating a neural network model is provided. The method includes: a second device receiving an indication of the structure of a sub-network model from a plurality of first devices, the sub-network model being determined by adjusting the structure of a super-network model; the second device determining the parameters of the sub-network model based on the indication and the super-network model; and the second device sending the parameters of the sub-network model to the plurality of first devices for the plurality of first devices to train the sub-network model respectively. In the above manner, an efficient federated learning solution among multiple devices is provided, which not only improves the model accuracy but also reduces the communication cost required in the federated learning process and the computing cost of the devices.

[0014] In some embodiments, the indication has the form of a mask, and the mask indicates whether the sub-network model has corresponding parameters of the super-network model. In the above manner, in the federated learning process, only the parameters specific to the personalized model need to be transmitted, thereby reducing the communication cost of transmission between devices.

[0015] In some embodiments, the method according to the second aspect may further include: the second device receiving the changes in the parameters of the sub-network model after training from the plurality of first devices. In the above manner, in the federated learning process, only the parameters that have changed after training need to be transmitted, thereby reducing the communication cost of transmission between the server device and the client device.

[0016] In some embodiments, the method according to the second aspect may further include: the second device using the received changes in the parameters to update the super-network model. In the above manner, in the federated learning process, the super-network model at the second device can be iteratively updated, thereby generating a super-network model that better conforms to the data distribution of the first devices.

[0017] In some embodiments, updating the super-network model may include: updating the super-network model based on the update weights of the parameters of the sub-network model, where the update weights depend on the number of sub-network models having the parameters. In the above manner, in the federated learning process, the parameters of the server device can be updated with weights, thereby generating a super-network model that better conforms to the client data distribution.

[0018] In some embodiments, the method according to the second aspect may further include: the second device determining pre-configured parameters of the super-network model for the plurality of first devices to determine their respective sub-network models from the super-network model. In the above manner, the plurality of first devices can start from the super-network model with the same pre-configured parameters and use local data to determine the structure of the sub-network model.

[0019] In some embodiments, determining the preconfigured parameters of the hypernetwork model may include: determining the preconfigured parameters of the hypernetwork model based on the local parameters determined by local training of the hypernetwork model by multiple first devices. In this way, optimized preconfigured parameters can be generated for the hypernetwork model based on the data distributions of multiple first devices, which is beneficial for the first devices to generate personalized neural network models at a relatively low computational cost.

[0020] According to a third aspect of the present disclosure, there is provided an apparatus for generating a neural network model, including: a sending unit configured to send an indication of the structure of a subnetwork model to a second device, the subnetwork model being determined by adjusting the structure of the hypernetwork model; a receiving unit configured to receive the parameters of the subnetwork model from the second device, the parameters of the subnetwork model being determined by the second device based on the indication and the hypernetwork model; and a training unit configured to train the subnetwork model based on the received parameters of the subnetwork model. The sending unit is further configured to send the parameters of the trained subnetwork model to the second device for the second device to update the hypernetwork model. In this way, an efficient federated learning solution among multiple devices is provided, which not only improves the model accuracy but also reduces the communication cost required in the federated learning process and the computational cost of the devices.

[0021] In some embodiments, the receiving unit may further be configured to obtain the preconfigured parameters of the hypernetwork model. In some embodiments, the apparatus according to the third aspect further includes a model determination unit configured to determine the structure of the subnetwork model by adjusting the structure of the hypernetwork model based on the preconfigured parameters. In this way, a personalized subnetwork model can be generated on the device.

[0022] In some embodiments, in the apparatus according to the third aspect, the training unit is configured to locally train the hypernetwork model to determine the local parameters of the hypernetwork model. The sending unit is further configured to send the local parameters to the second device. The receiving unit is further configured to receive the preconfigured parameters from the second device, the preconfigured parameters being determined by the second device at least based on the local parameters received from the first device. In this way, optimized preconfigured parameters can be generated for the hypernetwork model based on the data distribution of the device, which is beneficial for the device to determine the structure of the subnetwork model at a relatively low computational cost.

[0023] In some embodiments, the model determination unit is further configured to initialize the sub-network model as a super-network model with pre-configured parameters. The model determination unit is further configured to iteratively update the sub-network model by performing the following operations at least once: adjusting the structures of multiple layers of the sub-network model to obtain multiple candidate network models; selecting one candidate network model to update the sub-network model based on the accuracies of the multiple candidate network models; and stopping the iterative update if the sub-network model meets the constraints of the first device. In the above manner, the device can simplify the super-network model with accuracy as a measurement index to obtain a personalized neural network model that meets its own resource limitations and data distribution.

[0024] In some embodiments, the model determination unit is further configured to: delete the parameters related to a part of the nodes of one layer among the multiple layers to obtain one of the multiple candidate network models. In some embodiments, the model determination unit is configured to determine a part of the nodes based on a predetermined number ratio. In the above manner, the nodes and parameters with less influence on the model accuracy can be removed, thereby ensuring the model accuracy while simplifying the model structure.

[0025] In some embodiments, the computational amount of the sub-network model is lower than a first threshold, or the number of parameters of the sub-network model is lower than a second threshold. In the above manner, the computational amount and the number of parameters of the sub-network model of the device can be reduced to meet the corresponding resource constraints.

[0026] In some embodiments, both the first threshold and the second threshold are associated with the performance of the first device. In the above manner, the corresponding resource constraints are set according to the performance of the device.

[0027] In some embodiments, the indication has the form of a mask, and the mask indicates whether the sub-network model has the corresponding parameters of the super-network model. In the above manner, during the federated learning process, only the specific parameters of the sub-network model need to be transmitted, thereby reducing the communication cost of transmission between the server device and the client device.

[0028] In some embodiments, the sending unit may further be configured to send the change in the parameters of the trained sub-network model to the second device for the second device to update the super-network model, where the change in the parameters can be determined as the difference between the parameters of the trained sub-network model and the parameters received from the second device. In the above manner, during the federated learning process, only the parameters that have changed after training need to be transmitted, thereby reducing the communication cost of transmission between the devices.

[0029] According to a fourth aspect of the present disclosure, there is provided an apparatus for generating a neural network model, including: a receiving unit configured to receive an indication of the structure of a sub-network model from a plurality of first devices, the sub-network model being determined by adjusting the structure of a super-network model; a sub-network model parameter determination unit configured to determine the parameters of the sub-network model based on the indication and the super-network model; a sending unit configured to send the parameters of the sub-network model to the plurality of first devices for the plurality of first devices to train the sub-network model respectively; and a super-network updating unit, wherein the receiving unit is further configured to receive the parameters of the trained sub-network model from the plurality of first devices; and the super-network updating unit is configured to update the super-network model using the received parameters. In the above manner, an efficient federated learning solution between multiple devices is provided, which not only improves the model accuracy but also reduces the communication cost required in the federated learning process and the computing cost of the devices.

[0030] In some embodiments, the indication has the form of a mask, and the mask indicates whether the sub-network model has the corresponding parameters of the super-network model. In the above manner, in the federated learning process, only the parameters specific to the personalized model need to be transmitted, thereby reducing the communication cost of transmission between devices.

[0031] In some embodiments, the receiving unit may further be configured to receive the change in the parameters of the sub-network model after training from the plurality of first devices. In the above manner, in the federated learning process, only the parameters that have changed after training need to be transmitted, thereby reducing the communication cost of transmission between the server device and the client device.

[0032] In some embodiments, the super-network updating unit may further be configured to: update the super-network model using the received change in the parameters. In the above manner, in the federated learning process, the super-network model at the second device can be iteratively updated, thereby generating a super-network model that better conforms to the data distribution of the first device.

[0033] In some embodiments, the super-network updating unit may further be configured to update the super-network model based on the update weight of the parameters of the sub-network model, and the update weight depends on the number of sub-network models having the parameters. In the above manner, in the federated learning process, the parameters of the server device can be updated with weights, thereby generating a super-network model that better conforms to the client data distribution.

[0034] In some embodiments, the super-network updating unit may further be configured to determine pre-configured parameters of the super-network model for the plurality of first devices to determine their respective sub-network models from the super-network model. In the above manner, the plurality of first devices can start from the super-network model with the same pre-configured parameters and use local data to determine the structure of the sub-network model.

[0035] In some embodiments, the hypernetwork update unit may also be configured to determine pre-configured parameters based on local parameters determined by local training of the hypernetwork model by multiple first devices. In the above manner, optimized pre-configured parameters can be generated for the hypernetwork model based on the data distributions of the multiple first devices, which is beneficial for the first devices to generate personalized neural network models at a relatively low computational cost.

[0036] According to a fifth aspect of the present disclosure, there is provided a system for generating a neural network model, including: a first device configured to execute the method according to the first aspect of the present disclosure; and a second device configured to execute the method according to the second aspect of the present disclosure.

[0037] In a sixth aspect of the present disclosure, there is provided an electronic device, including: at least one computing unit; and at least one memory coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions, when executed by the at least one computing unit, causing the device to implement the method in any one of the implementations of the first aspect or the second aspect.

[0038] In a seventh aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements the method in any one of the implementations of the first aspect or the second aspect.

[0039] In an eighth aspect of the present disclosure, there is provided a computer program product including computer-executable instructions that, when executed by a processor, implement some or all of the steps of the method in any one of the implementations of the first aspect or the second aspect.

[0040] It can be understood that the system provided in the above fifth aspect, the electronic device provided in the sixth aspect, the computer storage medium provided in the seventh aspect, or the computer program product provided in the eighth aspect are all used to execute the methods provided in the first aspect and / or the second aspect. Therefore, the explanations or descriptions regarding the first aspect and / or the second aspect also apply to the fifth aspect, the sixth aspect, the seventh aspect, and the eighth aspect. In addition, the beneficial effects achievable by the fifth aspect, the sixth aspect, the seventh aspect, and the eighth aspect can refer to the beneficial effects in the corresponding methods, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0042] Figure 1 shows an architecture diagram of a federated learning system according to an embodiment of the present disclosure;

[0043] Figure 2 Shows a schematic diagram of communication for federated training according to an embodiment of the present disclosure;

[0044] Figure 3 Shows a flowchart of a process for generating a neural network model according to some embodiments of the present disclosure;

[0045] Figure 4 Shows a flowchart of another method for generating a neural network model according to some embodiments of the present disclosure;

[0046] Figure 5 Shows a schematic diagram of communication for determining the structure of a sub-network model according to an embodiment of the present disclosure;

[0047] Figure 6 Shows a flowchart of a process for initializing a hyper-network according to some embodiments of the present disclosure;

[0048] Figure 7 Shows a flowchart of a process for determining the structure of a sub-network model according to an embodiment of the present disclosure;

[0049] Figure 8 Shows a schematic diagram of pruning a hyper-network according to an embodiment of the present disclosure;

[0050] Figures 9A to 9D Respectively show test results according to some embodiments of the present disclosure;

[0051] Figure 10 Shows a schematic block diagram of an apparatus for generating a neural network model according to some embodiments of the present disclosure;

[0052] Figure 11 Shows a schematic block diagram of another apparatus for generating a neural network model according to some embodiments of the present disclosure; and

[0053] Figure 12 Shows a block diagram of a computing device capable of implementing multiple embodiments of the present disclosure. Detailed Description of Specific Embodiments

[0054] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0055] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0056] As used herein, a "neural network" is capable of processing inputs and providing corresponding outputs, and generally includes an input layer and an output layer, as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications generally include many hidden layers, thus extending the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is used as the final output of the neural network. In this article, the terms "neural network", "network", "neural network model" and "model" can be used interchangeably.

[0057] As used herein, "federated learning" is a machine learning algorithm, usually consisting of a server device and multiple client devices. By transmitting non-sensitive information such as parameters and using the data of multiple clients to train a machine learning model, the purpose of protecting privacy can be achieved.

[0058] As used herein, a "hypernetwork" (also referred to as a hypernetwork model or a parent network) refers to a neural network structure shared by a server and multiple clients in a federated learning environment. A "subnetwork" (also referred to as a subnetwork model or a personalized network) refers to a neural network model maintained, modified, trained, and inferred separately by each client in a federated learning environment. In this article, the subnetwork model and the subnetwork can be used interchangeably, and the hypernetwork model and the hypernetwork can be used interchangeably.

[0059] The subnetwork model can be obtained by pruning the hypernetwork. Here, pruning refers to deleting some parameters and corresponding computational operations in the network model, while the computational process of the remaining part remains unchanged.

[0060] Federated learning is a distributed machine learning technology. The server device and multiple client devices each perform model training locally without collecting data. In recent years, with the increasing demand for user privacy protection, federated learning has received more and more attention.

[0061] However, traditional federated learning requires an expert to design a globally shared model architecture, which may not achieve optimal performance. Recently, some works have used Neural Architecture Search (NAS) methods to automatically search for better model architectures. However, these methods only search for a globally shared model and do not consider the end-side personalization problem. In practical scenarios, the data distributions at each client device may be different (e.g., images taken by different users), and there may also be different resource constraints (e.g., different client devices have different computing capabilities). Therefore, deploying the same model on all clients will lead to a decline in model performance and affect the user experience. On the other hand, the data distributions of different clients are uneven, resulting in a global model not being able to achieve optimal performance on multiple clients simultaneously.

[0062] To address the above problems and other possible problems, embodiments of the present disclosure provide a personalized network architecture search framework and training method based on federated learning. First, in a federated learning system, a personalized model architecture is searched through server-client interaction to provide personalized model architectures for clients with different computing capabilities to meet resource constraints (e.g., model size, floating-point operations FLOPs, inference speed, etc.). Then, effective federated training is performed on different client network architectures. Compared with only local training, the embodiments of the present disclosure improve the model accuracy while reducing the communication cost required in the federated learning process and the computing cost of the device.

[0063] The following describes various example embodiments of the present disclosure with reference to the accompanying drawings.

[0064] Example environment

[0065] Embodiments of the present disclosure propose a personalized network structure search framework for resource-constrained devices (Architecture of Personalization Federated Learning, APFL). This framework can customize model architectures for different devices according to their specific resource requirements and local data distributions. In addition, the framework proposes a communication-friendly sub-network federated training strategy to efficiently complete local machine learning tasks on devices.

[0066] Figure 1The architecture diagram of the federated learning system 100 implementing the above framework according to some embodiments of the present disclosure is shown. The system 200 includes a plurality of first devices 110-1, 110-2... 110-K (collectively referred to as the first devices 110) and a second device 120. Herein, the first device may refer to a client device or simply a client, which may be, for example, a terminal device such as a smart phone, a desktop computer, a laptop computer, a tablet computer, etc. with limited computing resources (e.g., processor computing power, memory, storage space). The second device 120 may refer to a distributed or centralized server device or server cluster implemented in a cloud computing environment, which generally has higher computing resources or performance compared to the terminal device.

[0067] The plurality of first devices 110 and the second device 120 can communicate and transmit data with each other through various wired or wireless means. According to the embodiments of the present disclosure, the first device 110 and the second device 120 can construct and train neural network models for various applications such as image recognition, speech processing, and natural language processing in a collaborative manner, that is, federated learning.

[0068] As shown in the figure, the first device 110 includes a sub-network 111, local data 112, and a control unit 113. The control unit 113 can be a virtual or physical processing unit, which can, for example, use the local data 112 to train the sub-network 111, use the local data 112 to adjust the structure of the sub-network 111, and can also enable the first device 110 to communicate with the second device 120, such as receiving the parameters of the sub-network from the second device or transmitting the sub-network parameters or other information to the server. The second device 120 can include a super-network 121 and a control unit 123. Similarly, the control unit 123 can also be a virtual or physical processing unit, which can maintain the super-network 121, for example, by training or updating the parameters of the super-network 121 based on the parameters of the sub-network 111, and can also enable the second device 120 to communicate with some or all of the first devices 110, such as sending the parameters of the super-network 123, the parameters of the sub-network 111 for different first devices 110, and receiving the parameters of the sub-network 111 from the first device 110. Herein, the sub-network model and the sub-network can be used interchangeably, and the super-network model and the super-network can be used interchangeably.

[0069] The super-network 123 can be shared between the first device 110 and the second device 120, so that the first device 110 can prune the super-network to determine the sub-network 111 of the first device 110. The determined sub-network 111 can meet the resource constraints and data distributions of different first devices 110.

[0070] According to an embodiment of the present disclosure, there is provided a federated training method for a super network 121 and a sub network 111 implemented by a plurality of first devices 110 and a second device. The following is a reference to Figures 2 to 8 for a detailed description.

[0071] Federated training

[0072] According to an embodiment of the present disclosure, a plurality of first devices 110 and a second device 120 transmit parameters and other information of their respective neural network models to each other, thereby implementing federated training.

[0073] Figure 2 FIG. 200 shows a schematic diagram of communication 200 for federated training according to an embodiment of the present disclosure. As shown in the figure, a plurality of first devices 110 send 202 an indication of the structure of their sub network models 111 to the second device 120. According to an embodiment of the present disclosure, the structure of the sub network model 111 is determined based on the super network 123, and the indication of the structure may be in the form of a mask, for example, a 0-1 vector having the same dimension as the super network model dimension. A value of 1 for a component indicates that the corresponding parameter of the sub network model is retained, and a value of 0 for a component indicates that the corresponding parameter is deleted from the super network 123.

[0074] The second device 120 calculates 204 the sub network models 111 for each first device 110 according to the received indication. For example, the second device 120 may multiply the parameters of the super network model 121 by the received mask component by component, and intercept the non-zero components to obtain the parameters of the sub network model 111. Then the second device transmits 206 the obtained parameters of the sub network model 111 to the corresponding first device.

[0075] Next, each first device 110 trains 208 various sub network models 111 based on the received parameters using local data 112, and transmits 210 the parameters of the trained sub network model 111 to the second device 120. The transmitted parameters may be parameter values or changes in the parameters, that is, update amounts.

[0076] The second device 120 updates the super network model 121 using the received parameters of the sub network model 111 trained by the first device 110. For example, the parameters of the super network model 121 are updated by calculating the average value, weighted average or other means of each parameter. This federated learning according to an embodiment of the present disclosure may be iterative, that is, after the second device 120 updates the super network model 121, the actions of 204, 206, 208, 210, 212 are repeated above. In addition, the first devices 110 participating in the above process each time may be variable, for example, any subset of the currently online first devices 110.

[0077] The following is a reference toFigure 3 and Figure 4 describe the processes by which the first device 110 and the second device 120 respectively perform federated learning to generate a neural network model, where the second device 120 maintains a super network The personalized network of each first device 110 is a sub-network of, and the architectures of the respective sub-networks can be different.

[0078] Figure 3 FIG. shows a flowchart of a process 300 for generating a neural network model according to some embodiments of the present disclosure. The process 300 can be adapted to be implemented in each of the K first devices 110, where u k (k = 1,..., K), and more specifically, is executed by the control unit 113.

[0079] In block 310, an indication of the structure of the sub-network model is sent to the second device. The sub-network model is determined by adjusting the structure of the super-network model. In some embodiments, the indication can be in the form of a mask, indicating whether the sub-network model 111 has the corresponding parameters of the super-network model 121. For example, the first device u k sends a mask z k to the second device 120. Where z k is a 0-1 vector having the same dimension as the super-network 121 representing the sub-network structure That is, if the value of the component in z k is 1, it represents that the corresponding parameter in the super-network 121 is retained in , and if the component value is 0, it represents that the corresponding parameter is deleted. Thus, the second device 120 can learn the structure of the sub-network model 111 of the first device 110, and thus can determine the parameters of the sub-network model 111 for each first device. For example, at the start of sub-network federated training, the super-network 121 of the second device 120 has pre-configured parameters, that is, set

[0080] According to an embodiment of the present disclosure, sub-network federated training can be performed iteratively, for example, for T' rounds. In iteration t = 0,..., T' - 1, K' (K' ≤ K) first devices can be selected from the K first devices 110 to participate in this round of iteration. For the selected first devices, the actions as in blocks 720 to 740 are performed.

[0081] In block 320, the parameters of the sub-network model are received from the second device. According to an embodiment of the present disclosure, the second device 120 calculates the parameters of the sub-network model 111 based on the current parameters of the super-network model 121 and the mask of the sub-network model. For example, calculate Then, the second device 120 can intercept the non-zero components and transmit them to the first device u k . Here represents the multiplication of vector components. Thus, the first device 110 receives the parameters of the sub-network model 111.

[0082] Then, at block 330, the sub-network model is trained based on the received parameters. In some embodiments, the first device u k receives and assigns the parameters of the sub-network model 111 to update the sub-network parameters with the local training data 112 to obtain the updated parameters For example, S' steps of gradient descent can be performed using the local training data, as shown in the following equation (1)

[0083]

[0084]

[0085]

[0086] where λ is the learning rate, L is the cross-entropy loss function, is a random sample of the training set .

[0087] At block 340, the parameters of the trained sub-network model are sent to the second device. In some embodiments, the first device 110 sending the parameters of the sub-network model 111 to the second device 120 can include calculating the change in the parameters and transmitting the calculated to the second device. Alternatively, the second device 120 can also directly send the parameters of the trained sub-network model to the second device 120 without calculating the change in the parameters.

[0088] Correspondingly, the second device 120 receives the parameters of the trained sub-network model 111 from K' clients and updates the hyper-network 121 based on the received parameters. As above, the parameters received by the second device 120 can be the change or update amount of the parameters, or the parameters of the updated sub-network model 121.

[0089] Figure 4 shows a flowchart of a process 400 for generating a neural network model according to an embodiment of the present disclosure. The process 400 can be adapted to be implemented on Figure 2 the second device 120 shown, and more specifically, is executed by the control unit 123.

[0090] At block 410, an indication of the structure of the sub-network model 111 is received from multiple first devices 110, and the sub-network model 111 is determined by adjusting the structure of the hyper-network model 121. In some embodiments, the indication has the form of a mask, and the mask indicates whether the sub-network model has the corresponding parameters of the hyper-network model. For example, the mask can be a 0-1 vector. If the value of a component in the vector is 1, it means that the corresponding parameter in the hyper-network is retained in the sub-network model; if the component value is 0, it means that the corresponding parameter is deleted.

[0091] At block 420, the parameters of the sub-network model are determined based on the indication and the hyper-network model. In some embodiments, the second device 120 multiplies the indication in the form of a mask and the parameters of the hyper-network 121 component-wise, and determines the non-zero vector among them as the parameters of the sub-network model 111.

[0092] At block 430, the parameters of the sub-network model are sent to multiple first devices for the multiple first devices to train the sub-network model respectively. Thus, after receiving the parameters of the sub-network model 111, the first device 110 uses the local training data 112 to train the local personalized sub-network model 111 by the gradient descent method. As described above, after training, the first device 110 sends the parameters of the trained sub-network model 111 to the second device 120.

[0093] Correspondingly, at block 440, the parameters of the trained sub-network model are received from multiple first devices. In some embodiments, the parameters of the trained sub-network model received by the second device 120 include the change in the parameters, and the change in the parameters can be the difference between the parameters before and after the training of the sub-network model.

[0094] At block 450, the hyper-network model is updated using the received parameters. In some embodiments, the hyper-network model 121 is updated based on the update weights of the parameters of the sub-network model 111, and the update weights depend on the number of sub-network models having the corresponding parameters in the sub-network models of these first devices 110. For example, the update weight Z t can be set according to the following equation (2):

[0095]

[0096] where Recip(x) takes the reciprocal of each component in the vector x, and if the component value is 0, it is set to 0.

[0097] Then, the second device 120 updates the hyper-network parameters based on the update weights and the received parameters of the sub-network model 111. The hyper-network parameters are updated according to the following equation (3):

[0098]

[0099] where is the hypernetwork parameter after this round of iteration, is the hypernetwork parameter at the start of local iteration, Z t is the updated weight, is the change in the parameter at the first device 110, represents the multiplication of vector components.

[0100] The above-mentioned federated learning process can be repeatedly executed T' times as needed to complete the federated training process between the second device and a number of first devices.

[0101] As a result of generating the neural network model, for each client u k , k = 1,..., K, the second device 120 determines the training result of the sub-network model 111 based on the indication or mask of the structure of the first device 110 in the following specific manner: The second device 120 calculates and intercepts the non-zero components and transmits them to a specific first device u k , this specific first device u k receives and sets the final sub-network parameter as

[0102] Sub-network personalization

[0103] As described above, the sub-network model 111 is determined by the first device 110 by adjusting the structure of the hypernetwork model 123, and the determined sub-network model 111 can meet the resource constraints and data distribution of each first device 110. Embodiments of the present disclosure also provide a method for generating such a personalized sub-network, which will be described below in conjunction with Figures 5 to 8 description.

[0104] Figure 5 FIG. 500 shows a schematic diagram of communication for determining the structure of a sub-network model according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, a personalized sub-network model suitable for the first device 110 is determined based on a federated learning approach.

[0105] First, the second device 120 transmits the parameters of the hypernetwork 121 to a plurality of first devices 110. In some embodiments, the plurality of first devices 110 may be a part of all first devices 110, preferably first devices with higher performance. At this time, the first device 110 only has a shared hypernetwork, and the first device 110 uses the received parameters to train 504 the hypernetwork.

[0106] Then, the first device 110 transmits 506 the parameters of the trained super network to the second device 120. The second device 120 can aggregate 508 these parameters to update the super network 121. The parameters of the updated super network can be transmitted 510 to the first device 110 as preconfigured parameters. Then, the first device 120 can initialize the local super network with the preconfigured parameters and then adjust (e.g., prune) the structure of the super network to determine a sub-network model 111 that meets its own resource constraints and data distribution.

[0107] In some embodiments, the super network 123 of the second device 120 can be iteratively updated, i.e., the second device 120 can repeat operations 502 to 508 after updating the super network 123. And, different first devices 110 can be selected for each iteration.

[0108] The following combines Figures 6 to 8 to describe in detail the process in which the first device 110 and the second device 120 perform federated learning to determine the personalized sub-network 111.

[0109] Figure 6 The flowchart of process 400 for initializing a super network according to some embodiments of the present disclosure is shown. Process 400 can be implemented on the second device 120 and more specifically executed by the control unit 123. Through this process, the preconfigured parameters of the super network 121 can be determined for the first device 110 to subsequently generate a personalized sub-network model 111.

[0110] In block 410, a super network is selected and initial parameters are generated. According to an embodiment of the present disclosure, the second device 21 can select a neural network model from a neural network model library as the super network 121, such as MobileNetV2. The neural network model can include an input layer, multiple hidden layers, and an output layer, and each layer includes multiple nodes. The nodes between layers are connected by edges with weights. The weights can be referred to as model parameters w, which can be learned through training. Depending on the number of layers or the depth of the network model, the number of nodes in each layer, and the number of edges, the neural network model has a corresponding resource budget, including model size, floating-point operations, the number of parameters, inference speed, etc. In some embodiments, the second device 120 can randomly generate the initial parameters w of the super network 121 (0) 。

[0111] As shown in the figure, blocks 620 to 650 can be iteratively executed multiple times (e.g., T rounds, where T is a positive integer). By iterating (t = 0, 1,... T - 1), the parameters of the hypernetwork 121 are updated to determine the preconfigured parameters of the hypernetwork 121. Specifically, in block 620, hypernetwork parameters are sent to a selected plurality of first devices 110 for the plurality of second devices 120 to locally train the hypernetwork. According to an embodiment of the present disclosure, the second device 120 may select a part of the first devices from the online K first devices 110. For example, it may randomly select or preferably select K' clients with higher performance from the K first devices 110 The second device 120 sends the current hypernetwork parameter w (t) to the selected first devices. Then, each selected first device u k obtains the local parameters of the hypernetwork using the local training data 112 Note that at this time, the first device 110 has not yet generated the personalized subnetwork 111. The first device 110 uses the local training data 112 to train the hypernetwork determined in block 610

[0112] In some embodiments, each selected first device uses the local training data 112 to obtain new hypernetwork parameters through S-step gradient descent as shown in Equation (4).

[0113]

[0114]

[0115]

[0116] where λ is the learning rate and L is the cross-entropy loss function, is a random sample of the training set .

[0117] Then, the first device u k transmits the hypernetwork parameters generated in this round of iteration to the second device 120

[0118] Correspondingly, in block 630, the local parameters of the trained hypernetwork are received from the plurality of first devices 110

[0119] Then, in block 640, the global parameters of the hypernetwork are updated. According to an embodiment of the present disclosure, the second device 120 updates the global parameters of the hypernetwork 121 according to the received client local parameters. For example, the second device 120 updates the global parameters according to the following Equation (5):

[0120] where w (t+1) represents the global parameters of the hypernetwork after the (t + 1)-th iteration represents the local hypernetwork parameters generated by the first device 110 participating in this iteration. The Update() function can be a pre-set update algorithm. In some embodiments, the global parameters are calculated based on the weighted average of the training data volume, as shown in the following equation (6)

[0121]

[0122] where is the client u k the size of the training data set Alternatively, the arithmetic mean of the parameters can also be calculated as the global parameters of the hypernetwork

[0123] In block 650, it is determined whether T rounds have been reached. If T rounds have not been reached, it is possible to return to block 620 for the next iteration. If T rounds have been reached, then in block 660, the current global parameters of the hypernetwork 121 are determined as the pre-configured parameters of the hypernetwork 121, as shown in equation (7)

[0124]

[0125] where are the pre-configured parameters of the hypernetwork 121, which can be used by the first device 110 to generate the personalized subnetwork 111

[0126] In some embodiments, when iteratively updating the global parameters of the hypernetwork 121, the second device 120 can select K' first devices 110 participating in the iteration according to the performance level of the first device 110 to accelerate the initialization of the hypernetwork 121

[0127] After completing the global initialization of the hypernetwork, the second device 120 maintains the hypernetwork 121 with parameters A method for a neural network model on a personalized device is provided, which can simplify the structure of the hypernetwork 121 through pruning to generate a personalized subnetwork 111, and the generated personalized subnetwork conforms to the data distribution and resource constraints of the client. The following will be described in detail with reference to Figure 7 and Figure 8 for details

[0128] Figure 7 FIG. shows a flowchart of a process 700 for determining the structure of a subnetwork model according to an embodiment of the present disclosure. The process 700 can be adapted to be executed by each of the K first devices 110. The process 700 simplifies the hypernetwork by iteratively deleting the nodes and parameters of the hypernetwork 121, while also ensuring that the resulting subnetwork 111 has good accuracy

[0129] According to an embodiment of the present disclosure, the supernetwork 121 is pruned based on the accuracy of the network model. For example, parameters that have less impact on the accuracy can be selected according to a certain rule. In addition, resource constraints are regarded as a hard constraint condition, and at least one parameter needs to be deleted in each pruning operation to simplify the supernetwork. As described above, the supernetwork 121 can be initialized by the process described above between multiple first devices 110 and second devices 120, so it can achieve good accuracy and be pruned to generate a personalized subnetwork model 111. Figure 6 Description of the process, so it can be able to achieve good accuracy for being pruned to generate a personalized subnetwork model 111.

[0130] In block 710, the subnetwork model is initialized. According to an embodiment of the present disclosure, the first device 110 u k Receives the preconfigured parameters of the supernetwork 121 And initializes the local subnetwork model 111 as a supernetwork with the preconfigured parameters, that is, Where Represents the subnetwork at client k. That is, at the beginning, the subnetwork model 111 of the first device has the same structure and parameters as the supernetwork of the second device 120, and pruning starts from this. The embodiment of the present disclosure provides a method of pruning layer by layer, which iteratively simplifies the structure of the supernetwork until it meets the resource constraints, and at the same time can ensure that the resulting subnetwork has good accuracy.

[0131] The process of its pruning is described with reference to MobileNetV2, but it should be understood that other types of neural network models are applicable to the pruning method provided by the embodiments of the present disclosure.

[0132] Figure 8 Shows a schematic diagram of pruning the supernetwork according to an embodiment of the present disclosure. MobileNetV2 is used as the supernetwork, and channel (neuron) pruning operations are used, that is, pruning is performed in units of channels. Some parts of the network 800 cannot be pruned, such as the depthwise layer, input layer, output layer, etc., while 34 layers can be pruned. Assume that the prunable layers are L_1,..., L_34, and the pruning layer L_i contains m_i channels C_1,..., C_(m_i). For example, the local data of the first device 110 can be in the picture format of 32x32x3, that is, the input layer has 3 channels corresponding to the RGB values of the picture. As shown in the figure, the solid arrows represent the computational connections between adjacent layers, the marked X represents the pruned channels (nodes), the dashed arrows represent the corresponding deleted connections, and the computational process of the unpruned channels remains unchanged

[0133] At block 720, the structures of multiple layers of the sub-network model are adjusted to obtain multiple candidate neural network models. Specifically, for multiple prunable layers, parameters related to a part of the nodes of each prunable layer of the sub-network model 111 are deleted to obtain multiple candidate network models. In some embodiments, for each layer, a part of the nodes can be determined based on a predetermined number ratio (e.g., 0.1, 0.2, 0.25, etc.), and the parameters related to these nodes are deleted. Then, the local data of the first device can be used to determine the accuracy of the obtained pruned candidate neural network models.

[0134] For example, the first device u k Prunes, and for the i-th prunable layer of MobileNetV2, i = 1,..., 34, the following operations are performed:

[0135] 1. Delete the channels with labels greater than That is, retain the channels

[0136] 2. Let the corresponding deleted parameter be v i .

[0137] 3. Calculate the accuracy of the network On the training set D k , denoted as A i . / / Note that after calculating the accuracy Restore

[0138] Then, at block 730, based on the accuracies of the multiple candidate network models, a candidate network model is selected to update the sub-network model. In some embodiments, the prunable layer i * = argmax i A i with the highest accuracy after pruning can be determined. Thus, the parameter v i* can be deleted from the sub-network model , and

[0139] ​In box 740, determine whether the current sub-network model meets the resource constraint. In some embodiments, the amount of computation such as floating point operations (FLOPs) or the number of parameters of the model can be used as a resource constraint. For example, the first device 110 can be classified as follows: (i) high-budget clients, with an upper limit of 88M FLOPs and an upper limit of 2.2M parameters; (ii) medium-budget clients, with an upper limit of 59M FLOPs and an upper limit of 1.3M parameters; (iii) low-budget clients, with an upper limit of 28M FLOPs and an upper limit of 0.7M parameters. In some embodiments, two federated learning settings can be considered: (1) a high-performance configuration in which the ratio of high, medium, and low-budget clients is 5:3:2; and (2) a low-performance configuration in which the ratio of high, medium, and low-budget clients is 2:3:5. This takes into account both the efficiency and accuracy of subsequent federated training.

[0140] If the current sub-network model does not meet the resource constraint, for example, the computational amount of the sub-network model is higher than or equal to a certain threshold, or the number of parameters of the sub-network model is higher than or equal to another threshold, or both of the above are true, then repeat the above blocks 720 and 730 for further pruning. In some embodiments, the above thresholds on the computational amount and the number of parameters depend on the performance of the first device 110. The higher the performance of the first device 110, the higher the threshold can be set, so that the sub-network model 111 is pruned less, which means that higher accuracy can be achieved after training.

[0141] If the current sub-network model 111 meets the resource constraints, that is, the computational complexity of the sub-network model 111 is lower than the first threshold, or the number of parameters of the sub-network model is lower than the second threshold, or both of the above are true, then in box 750, the structure of the current sub-network model is determined as a personalized sub-network for federated training of the first device 110.

[0142] It is worth noting that in the above process, the structure of the sub-network model 111 is iteratively adjusted, but the parameter value of the sub-network model 111 remains unchanged, that is, the pre-configured parameters received from the second device 120 are maintained. (If not deleted).

[0143] The above describes the process of determining the personalized sub-network model 111. The structure of the determined sub-network model can be indicated by, for example, a mask for federated training.

[0144] Test results

[0145] Figure 9A , Figure 9B , Figure 9C and Figure 9DShows the test results according to some embodiments of the present disclosure. The test results indicate that the solution (FedCS) according to the embodiments of the present disclosure achieves superior performance in the classification task, and the personalized network architecture after federated learning and training has advantages.

[0146] As Figure 9A shown, for high-performance clients, the solution FebCS of the present disclosure has higher accuracy in both IID and NIID types of data compared to other methods (FedUniform). For low-performance clients, the advantage is even more obvious.

[0147] As Figure 9B shown, the communication volume between the server and the client required by the network architecture search framework (FedCS) of the present disclosure is significantly reduced, which indicates a faster convergence speed and saves communication costs.

[0148] At the same time, the searched model has a smaller number of parameters under the condition of equivalent floating-point operation numbers (FLOPs), so it saves the storage space of the client, as follows Figure 9C shown.

[0149] In addition, as Figure 9D shown, compared with the client's individual training strategy, the sub-network federated training algorithm of the present disclosure has a significant improvement in accuracy, indicating that although the sub-network architectures of each client are different, using the masking strategy for federated training through the super network can effectively improve the performance of the sub-network.

[0150] Example device and equipment

[0151] Figure 10 Shows a schematic block diagram of an apparatus 1000 for generating a neural network model according to some embodiments of the present disclosure. The apparatus 1000 can be implemented in a first device 110 as Figure 1 shown.

[0152] The apparatus 1000 includes a sending unit 1010 configured to send an indication of the structure of the sub-network model to a second device, where the sub-network model is determined by adjusting the structure of the super-network model. The apparatus 1000 further includes a receiving unit 1020 configured to receive the parameters of the sub-network model from the second device, where the parameters of the sub-network model are determined by the second device based on the indication and the super-network model. The apparatus 1000 further includes a training unit 1030 configured to train the sub-network model based on the received parameters of the sub-network model. The sending unit 1010 is further configured to send the parameters of the trained sub-network model to the second device for the second device to update the super-network model.

[0153] In some embodiments, the receiving unit 1020 may also be configured to obtain preconfigured parameters of the hypernetwork model. In some embodiments, the apparatus 1000 further includes a model determination unit 1040, which is configured to determine the structure of the subnetwork model by adjusting the structure of the hypernetwork model based on the preconfigured parameters.

[0154] In some embodiments, the training unit 1030 is configured to locally train the hypernetwork model to determine the local parameters of the hypernetwork model. The sending unit 1010 is also configured to send the local parameters to a second device. The receiving unit 1020 is also configured to receive the preconfigured parameters from the second device, where the preconfigured parameters are determined by the second device based at least on the local parameters received from the first device.

[0155] In some embodiments, the model determination unit 1040 is also configured to initialize the subnetwork model as a hypernetwork model with preconfigured parameters. The model determination unit 1040 is also configured to iteratively update the subnetwork model by performing the following operations at least once: adjusting the structures of multiple layers of the subnetwork model to obtain multiple candidate network models; selecting one candidate network model to update the subnetwork model based on the accuracies of the multiple candidate network models; and stopping the iterative update if the subnetwork model meets the constraints of the first device.

[0156] In some embodiments, the model determination unit 1040 is also configured to: delete the parameters related to a part of the nodes of one of the multiple layers to obtain one of the multiple candidate network models. In some embodiments, the model determination unit 1040 is configured to determine a part of the nodes based on a predetermined number ratio. In the above manner,

[0157] In some embodiments, the computational amount of the subnetwork model is lower than a first threshold, or the number of parameters of the subnetwork model is lower than a second threshold.

[0158] In some embodiments, both the first threshold and the second threshold are associated with the performance of the first device.

[0159] In some embodiments, the indication has the form of a mask, and the mask indicates whether the subnetwork model has the corresponding parameters of the hypernetwork model.

[0160] In some embodiments, the sending unit 1010 may also be configured to send the change in the parameters of the trained subnetwork model to the second device for the second device to update the hypernetwork model, where the change in the parameters can be determined as the difference between the parameters of the trained subnetwork model and the parameters received from the second device.

[0161] Figure 11FIG. 0 shows a schematic block diagram of an apparatus 1100 for generating a neural network model according to some embodiments of the present disclosure. The apparatus 1100 may be implemented in a second device 120 as shown in Figure 1 FIG. 1.

[0162] The apparatus 1100 includes a receiving unit 1110 configured to receive an indication of a structure of a sub-network model from a plurality of first devices, the sub-network model being determined by adjusting a structure of a super-network model. The apparatus 1100 further includes a sub-network model parameter determination unit 1120 configured to determine parameters of the sub-network model based on the indication and the super-network model. The apparatus 1100 further includes a sending unit 1130 configured to send the parameters of the sub-network model to the plurality of first devices for the plurality of first devices to train the sub-network model respectively. The apparatus 1100 further includes a super-network updating unit 1140 configured to update the super-network model using the parameters of the trained sub-network model received from the plurality of first devices by the receiving unit 1110.

[0163] In some embodiments, the indication has a form of a mask, and the mask indicates whether the sub-network model has corresponding parameters of the super-network model.

[0164] In some embodiments, the receiving unit 1110 may further be configured to receive changes in the parameters of the sub-network model after being trained from the plurality of first devices. In the above manner,

[0165] In some embodiments, the super-network updating unit 1140 may further be configured to: update the super-network model using the received changes in the parameters.

[0166] In some embodiments, the super-network updating unit 1140 may further be configured to update the super-network model based on an update weight of the parameters of the sub-network model, and the update weight depends on the number of sub-network models having the parameters.

[0167] In some embodiments, the super-network updating unit 1140 may further be configured to determine pre-configured parameters of the super-network model for the plurality of first devices to determine their respective sub-network models from the super-network model.

[0168] In some embodiments, the super-network updating unit 1140 may further be configured to determine the pre-configured parameters based on local parameters determined by local training of the super-network model by the plurality of first devices. In the above manner,

[0169] Figure 12FIG. shows a block diagram of a computing device 1200 capable of implementing multiple embodiments of the present disclosure. The device 1200 can be used to implement the first device 110, the second device 120, and the apparatuses 1000 and 1100. As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to computer program instructions stored in a random access memory (RAM) and / or a read-only memory (ROM) 1202, or computer program instructions loaded from a storage unit 1207 into the RAM and / or the ROM 1202. In the RAM and / or the ROM 1202, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201 and the RAM and / or the ROM 1202 are connected to each other via a bus 1203. An input / output (I / O) interface 1204 is also connected to the bus 1203.

[0170] Multiple components in the device 1200 are connected to the I / O interface 1204, including: an input unit 1205, such as a keyboard, a mouse, etc.; an output unit 1206, such as various types of displays, speakers, etc.; a storage unit 1107, such as a magnetic disk, an optical disc, etc.; and a communication unit 1208, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1208 allows the device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0171] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1201 executes any of the methods and processes described above. For example, in some embodiments, the above process can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1207. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1200 via the RAM and / or the ROM and / or the communication unit 1208. When the computer program is loaded into the RAM and / or the ROM and executed by the computing unit 1201, one or more steps of any of the processes described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute any of the methods and processes described above in any other appropriate manner (e.g., by means of firmware).

[0172] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes may be executed entirely on the machine, partially on the machine, executed partially on the machine as a stand-alone software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0173] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0174] Moreover, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.

[0175] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A training method based on a distributed system, comprising: A first device in the distributed system sends an indication of the structure of a sub-network model to a second device in the distributed system, where the sub-network model is determined by the first device by adjusting the structure of a super-network model; The first device receives the parameters of the sub-network model from the second device, where the parameters of the sub-network model are determined by the second device based on the indication and the super-network model; The first device trains the sub-network model based on the received parameters of the sub-network model; And The first device sends the parameters of the trained sub-network model to the second device for the second device to update the super-network model.

2. The method according to claim 1, further comprising: The first device obtains pre-configured parameters of the super-network model; And Based on the pre-configured parameters, the first device determines the structure of the sub-network model by adjusting the structure of the super-network model.

3. The method according to claim 2, wherein obtaining the pre-configured parameters of the super network model includes The first device locally trains the super network model to determine the local parameters of the super network model; The first device sends the local parameters to the second device; And The first device receives the pre-configured parameters from the second device, where the pre-configured parameters are determined by the second device at least based on the local parameters received from the first device.

4. The method according to claim 2, wherein determining the structure of the sub-network model includes: Initialize the sub-network model as a super-network model with the pre-configured parameters; And Iteratively update the sub-network model by performing the following operations at least once Adjust the structures of multiple layers of the sub-network model to obtain multiple candidate network models; Based on the accuracies of the multiple candidate network models, select one candidate network model to update the sub-network model; and If the sub-network model meets the constraints of the first device, determine the structure of the sub-network model.

5. The method according to claim 4, adjusting the structure of the multiple layers includes: Delete the parameters related to a part of the nodes in one of the multiple layers to obtain one of the multiple candidate network models.

6. The method according to claim 5, further comprising: Determine the part of the nodes based on a predetermined number ratio.

7. The method according to claim 4, wherein the constraint includes: The computational amount of the sub-network model is lower than a first threshold, or the number of parameters of the sub-network model is lower than a second threshold.

8. The method according to claim 7, wherein both the first threshold and the second threshold are associated with the performance of the first device.

9. The method according to claim 1, wherein the indication has a masked form, and the mask indicates whether the sub-network model has the corresponding parameters of the super network model.

10. The method according to claim 1, sending the parameters of the trained sub-network model includes: Determine the change in the parameters by calculating the difference between the parameters of the trained sub-network model and the parameters received from the second device; And Send the change in the parameters to the second device.

11. A training method based on a distributed system, comprising: A second device in the distributed system receives indications of the structures of sub-network models from multiple first devices in the distributed system, where each sub-network model is determined by each of the multiple first devices by adjusting the structure of a super-network model; The second device determines the parameters of the sub-network model based on the indications and the super-network model; The second device sends the parameters of the sub-network model to the multiple first devices for the multiple first devices to train the sub-network models respectively; The second device receives the parameters of the trained sub-network models from the multiple first devices; And The second device uses the received parameters to update the super-network model.

12. The method according to claim 11, wherein the indication has the form of a mask, and the mask indicates whether the sub-network model has the corresponding parameters of the super-network model.

13. The method according to claim 11, wherein receiving the parameters of the trained sub-network model further comprises: Receive the changes in the parameters of the sub-network model after training from the multiple first devices.

14. The method according to claim 11, wherein updating the super-network model comprises: Update the hypernetwork model based on the updated weights of the parameters of the subnetwork model, where the updated weights depend on the number of subnetwork models having the parameters.

15. The method according to claim 11, further comprising: The second device determines preconfigured parameters of the hypernetwork model for the plurality of first devices to determine their respective subnetwork models from the hypernetwork model.

16. The method according to claim 15, wherein determining the pre-configured parameters comprises: Determine the preconfigured parameters based on local parameters determined by local training of the hypernetwork model by the plurality of first devices.

17. A training device, the device is arranged in a first device in a distributed system, and the device comprises: A sending unit, configured to send an indication of the structure of a subnetwork model to a second device in the distributed system, where the subnetwork model is determined by a first device by adjusting the structure of the hypernetwork model; A receiving unit, configured to receive parameters of the subnetwork model from the second device, where the parameters of the subnetwork model are determined by the second device based on the indication and the hypernetwork model; And A training unit, configured to train the subnetwork model based on the received parameters of the subnetwork model; where the sending unit is further configured to send the parameters of the trained subnetwork model to the second device for the second device to update the hypernetwork model.

18. The device according to claim 17, wherein the receiving unit is further configured to obtain pre-configured parameters of a super-network model, and the device further comprises: A model determination unit, configured to determine the structure of the subnetwork model by adjusting the structure of the hypernetwork model based on the preconfigured parameters.

19. The device according to claim 18, wherein the model determination unit is further configured to: Initialize the sub-network model to a super-network model with the pre-configured parameters; and Iteratively update the sub-network model by performing the following operations at least once Adjust the structures of multiple layers of the sub-network model to obtain multiple candidate network models; Select a candidate network model to update the sub-network model based on the accuracy of multiple candidate network models; and If the sub-network model meets the constraints of the first device, determine the structure of the sub-network model.

20. A training device, which is arranged in a second device in a distributed system, and the device includes: A receiving unit, configured to receive an indication of the structure of a subnetwork model from a plurality of first devices in the distributed system, where each subnetwork model is determined by each of the plurality of first devices by adjusting the structure of the hypernetwork model; A subnetwork model parameter determination unit, configured to determine parameters of the subnetwork model based on the indication and the hypernetwork model; A sending unit, configured to send the parameters of the subnetwork model to the plurality of first devices for the plurality of first devices to train the subnetwork model respectively; And A hypernetwork update unit where the receiving unit is further configured to receive the parameters of the trained subnetwork model from the plurality of first devices; and the hypernetwork update unit is configured to update the hypernetwork model using the received parameters.

21. The device according to claim 20, wherein the hyper-network updating unit is further configured to: Update the hyper-network model based on the update weight of the parameters of the sub-network model, and the update weight depends on the number of sub-network models having the parameters.

22. The device according to claim 20, wherein the hyper-network updating unit is further configured to: Determine pre-configured parameters of the hyper-network model for the multiple first devices to determine their respective sub-network models from the hyper-network model.

23. A distributed system, the distributed system includes: A first device, configured to perform the method according to any one of claims 1 to 10; And A second device, configured to perform the method according to any one of claims 11 to 16.

24. An electronic device, including: At least one computing unit; At least one memory, the at least one memory being coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, which when executed by the at least one computing unit cause the electronic device to perform the method according to any one of claims 1 to 10 or the method according to any one of claims 11 to 16.

25. A computer-readable storage medium, on which one or more computer instructions are stored, and when one or more computer instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 10 or the method according to any one of claims 11 to 16.

26. A computer program product, including machine-executable instructions, which when executed by a device cause the device to execute the method according to any one of claims 1 to 10 or the method according to any one of claims 11 to 16.

Citation Information

Patent Citations

  • Network structure searching method and device, equipment, storage medium and program product

    CN112818207A