Method and apparatus for neural network search
By employing a multi-agent system to generate and learn partial neural network structures in parallel, the method addresses the inefficiencies of high search space and long search times in NAS, achieving faster and more efficient neural network construction.
Patent Information
- Application Number
- CN202010075167.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-01-22
AI Technical Summary
The existing neural network structure search methods face problems such as high search space dimensions, long search time, and complex performance evaluation, resulting in low efficiency in establishing neural networks.
The method of generating neural networks by multiple agents is adopted. Each agent is responsible for the generation and learning of some network structures in the neural network. Through the parallel sampling and collaboration mechanism of multiple agents, the search space is reduced, the agent strategy update mechanism is optimized, and the network structure generation efficiency is improved.
It significantly improves the efficiency of neural network structure search, shortens the search time and reduces the number of parameters, while maintaining the accuracy of search results.
Smart Images

Figure CN113159268B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly, to a method and apparatus for neural network search. Background Art
[0002] In recent years, neural networks have developed rapidly, and in some fields, deep neural networks have outperformed humans. However, in practical applications, due to differences in application scenarios, data sets, deployment devices, and metric requirements, etc., it often takes a lot of time and effort for experienced experts to build a neural network that meets the application environment. To improve the efficiency of building neural networks, the industry has proposed using neural architecture search (NAS) to design neural networks in order to obtain a neural network that meets the application environment.
[0003] Neural architecture search can automatically search for a neural network that meets specific constraints and achieves a specific goal using a specific data set, that is, the user can complete the process of modeling using a deep neural network without scenario experience and knowledge and skills of deep learning.
[0004] Currently, neural architecture search faces problems such as a high-dimensional search space and a long search time, resulting in a low efficiency of building neural networks. Improving the efficiency of building neural networks is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a method and apparatus for neural network search, which can improve the efficiency of building neural networks.
[0006] In a first aspect, a method for neural network search is provided. The method is applied to a computing system, and the system includes a plurality of agents, including: determining a plurality of candidate neural networks, the plurality of candidate neural networks having the same network structure, and a first agent among the plurality of agents being used to process the same part in each of the plurality of candidate neural networks, the first agent being one of the plurality of agents; respectively using the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent to obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks, the new candidate neural networks including the partial network structure and the input after being processed by the first agent, where the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; and determining a target neural network according to the plurality of new candidate neural networks.
[0007] In the existing ENAS, a single-agent policy gradient model is adopted to predict the network structure sequence, so as to train each candidate network structure and select a neural network that meets the requirements from them. This leads to a large search space, making it difficult to learn the optimal path. In addition, a large number of network structure samples are required.
[0008] In the embodiments of the present application, a multi-agent generative neural network is adopted. Among them, each agent is responsible for the generation and learning of the same part of the network structure in each candidate neural network. For example, for multiple seven-layer convolutional neural network structures, one agent is only responsible for the generation and learning of the second-layer network structure of multiple seven-layer convolutional neural network structures. Therefore, the search space of each agent can be reduced, which is beneficial to learning the optimal path, and thus the efficiency of establishing a neural network can be improved.
[0009] In other words, in the embodiments of the present application, by adopting multiple agents to be respectively responsible for the generation of different parts of the network, the sequence structure of network generation is improved compared with the existing ENAS. Improving the sequence structure of network generation can be understood as decoupling the timing of network structure generation in the existing network structure search scheme. Therefore, the embodiments of the present application can improve the efficiency of network structure sampling and overall improve the efficiency of neural network structure search.
[0010] Combined with the first aspect, in some implementation manners of the first aspect, the part of the network structure of the neural network responsible by the first agent is an instruction of a node of the neural network.
[0011] In the embodiments of the present application, a multi-agent generative neural network is adopted. Among them, each agent is responsible for the generation and learning of part of the network structure. Each time a network structure is generated, only one agent takes an action, that is, each time during training, only one agent participates in generating the network structure, so that the input of one agent can be independent of the output of other agents, and thus the efficiency of multi-agent parallel sampling can be fully exerted.
[0012] In other words, in the embodiments of the present application, by making the input of one agent independent of the output of other agents, the agent policy update mechanism is optimized.
[0013] Combined with the first aspect, in some implementation manners of the first aspect, selecting the first agent among multiple agents includes: selecting the first agent according to the probability value corresponding to each agent.
[0014] Combined with the first aspect, in some implementation manners of the first aspect, each candidate neural network includes a normal unit and a decay unit, and the probability value of the agent corresponding to the normal unit is higher than the probability value of the agent corresponding to the decay unit.
[0015] Among them, the probability value of the agent is the probability of the agent being selected. In some cases, the agents should be trained more times. For example, in a neural network for image classification, the number of repetitions of normal units is much larger than that of decaying units. Therefore, the agents in normal units should be trained more than the agents in decaying units. In the embodiments of the present application, by setting the probability value of the agent corresponding to the normal unit to be higher than the probability value of the agent corresponding to the decaying unit, each agent can be trained sufficiently.
[0016] In combination with the first aspect, in some implementation manners of the first aspect, the multiple candidate neural networks include k first candidate neural networks, or the multiple candidate neural networks include k first candidate neural networks and k second candidate neural networks, where k is a positive integer; among them, the first instruction of the first candidate neural network is randomly initialized or determined by the result of the previous training of the agent, and the second agent responsible for the second instruction is other agents except the first agent among the multiple agents; the second instruction of the second candidate neural network is obtained by perturbing the first instruction, where both the first instruction and the second instruction are responsible for the second agent.
[0017] In each training of the agent, by perturbing the first instruction responsible for the second agent in the k first candidate neural networks to generate k second candidate neural networks, it can be avoided that the training result of the first agent only reaches the optimal in the k first candidate neural network structures, and it can be avoided that the training of the agent becomes rigid.
[0018] In combination with the first aspect, in some implementation manners of the first aspect, determining multiple candidate neural networks includes: determining multiple candidate neural networks according to a common network structure pool, where the common network structure pool is determined according to the result of the previous training of the agent.
[0019] Updating the common structure pool according to the result of the previous training of the agent can make the agents trained in the candidate neural networks in the common structure pool all have the optimal instructions, and retain the candidate neural networks with good evaluation results, so that the next training of the agent is carried out on the basis of the current optimal candidate neural network, ensuring that the training of the agent is in the direction of the overall network optimum.
[0020] Second aspect, a method for image processing is provided, including: obtaining an image to be processed; classifying the image to be processed according to a neural network to obtain a classification result of the image to be processed; wherein, the neural network is determined according to multiple agents in a computer system, and the determination of the neural network includes: determining multiple candidate neural networks, the multiple candidate neural networks having the same network structure, a first agent among the multiple agents being used to process the same part in each of the multiple candidate neural networks, the first agent being one of the multiple agents; respectively using the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent to obtain multiple new candidate neural networks corresponding to the multiple candidate neural networks, the new candidate neural networks including the part of the network structure and the input after being processed by the first agent, wherein, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; determining a target neural network according to the multiple new candidate neural networks.
[0021] The neural network searched by the method for neural network search provided by the embodiments of the present application can be directly used for image classification processing, and the above neural network is the neural network obtained according to the first aspect and any one of the implementation manners in the first aspect.
[0022] Combined with the second aspect, in some implementation manners of the second aspect, the partial network structure of the neural network responsible by the first agent is an instruction of a node of the neural network.
[0023] Combined with the second aspect, in some implementation manners of the second aspect, selecting the first agent among the multiple agents includes: selecting the first agent according to the probability value corresponding to each agent among the multiple agents.
[0024] Combined with the second aspect, in some implementation manners of the second aspect, each candidate neural network includes a normal unit and a decay unit, the probability value of the agent corresponding to the normal unit being higher than the probability value of the agent corresponding to the decay unit, wherein, the normal unit and the decay unit include multiple nodes.
[0025] Combined with the second aspect, in some implementation manners of the second aspect, the multiple candidate neural network structures include k first candidate neural networks, or, the multiple candidate neural network structures include k first candidate neural networks and k second candidate neural networks, k being a positive integer; wherein, the first instruction of the first candidate neural network is randomly initialized or determined by the result of the previous training of the agent, the second instruction of the second candidate neural network is obtained by perturbing the first instruction, and both the first instruction and the second instruction are responsible by a second agent, the second agent being other agents among the multiple agents except for the first agent.
[0026] In combination with the second aspect, in certain implementations of the second aspect, determining a plurality of candidate neural networks includes: determining a plurality of candidate neural networks according to a common network structure pool, where the common network structure pool is updated according to the training result of the previous training of the agent.
[0027] In a third aspect, there is provided a neural network structure search device, which is applied to a computing system including a plurality of agents. The device includes: a first determination unit for determining a plurality of candidate neural networks having the same network structure, where a first agent among the plurality of agents is used to process the same part in each of the plurality of candidate neural networks, and the first agent is one of the plurality of agents; a processing unit for respectively using the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent, to obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks. The new candidate neural networks include the partial network structure and the input after being processed by the first agent. Here, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; a second determination unit for determining a target neural network according to the plurality of new candidate neural networks.
[0028] In a fourth aspect, there is provided a neural network search device, which includes: a memory for storing a program; a processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method in the first aspect and any one of the implementations in the first aspect.
[0029] In a fifth aspect, there is provided a computer-readable medium storing program code for a device to execute, and the program code includes the method for executing the first aspect and any one of the implementations in the first aspect.
[0030] In a sixth aspect, there is provided a computer program product containing instructions, which when run on a computer, causes the computer to execute the method in the above-mentioned first aspect and any one of the implementations in the first aspect.
[0031] In a seventh aspect, there is provided a chip, which includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in the first aspect and any one of the implementations in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of an automatic machine learning platform architecture provided by an embodiment of the present application.
[0033] Figure 2 It is a schematic block diagram of a neural network search system provided by an embodiment of the present application.
[0034] Figure 3 It is a schematic structural diagram of a classical image classification convolutional neural network provided by an embodiment of the present application.
[0035] Figure 4 It is a schematic diagram of a neural network structure provided by an embodiment of the present application.
[0036] Figure 5 It is a schematic flow diagram of an efficient neural network structure search method provided by an embodiment of the present application.
[0037] Figure 6 It is a schematic diagram of a system architecture provided by an embodiment of the present application.
[0038] Figure 7 It is a flow chart of a neural network search method provided by an embodiment of the present application.
[0039] Figure 8 It is a schematic diagram of multi-agent network generation provided by an embodiment of the present application.
[0040] Figure 9 It is a schematic structural diagram of an image classification network provided by an embodiment of the present application.
[0041] Figure 10 It is a schematic structural diagram of a normal unit network provided by an embodiment of the present application.
[0042] Figure 11 It is a schematic structural diagram of an attenuation unit network provided by an embodiment of the present application.
[0043] Figure 12 It is a schematic flow diagram of a neural network search method provided by an embodiment of the present application.
[0044] Figure 13 It is a schematic block diagram of a neural network search device provided by an embodiment of the present application.
[0045] Figure 14 It is a schematic flow chart of an image processing method provided by an embodiment of the present application.
[0046] Figure 15 It is a schematic block diagram of an image processing device provided by an embodiment of the present application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0048] It should also be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A. B can also be determined according to A and / or other information. It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0049] With the rapid development of neural networks, in some fields, deep neural networks have been superior to humans. Therefore, the demand of enterprises or researchers for using deep neural networks to improve the performance of their own products or service efficiency is becoming more and more urgent. Briefly speaking, the establishment of a neural network is to determine the network structure (or called network architecture) and the corresponding parameters. Determining the network structure can be understood as designing the network structure (or called network architecture) of the neural network. After the network structure is determined, determining the parameters under this network structure can be understood as optimizing the parameters for this network structure. It should be understood that designing the network structure can also be understood as optimizing the network structure parameters.
[0050] For neural networks (especially deep neural networks), adjusting parameters is a huge project, and numerous parameters will generate explosive combinations. Therefore, it is unrealistic to design neural networks manually. Usually, users can design neural networks through an automated machine learning (AutoML) platform (or called tool). For example, the AutoML platform helps users obtain machine learning models that meet the application scenarios through technologies such as transfer learning and automatic hyperparameter tuning, reducing repeated experiments and improving the modeling efficiency. As Figure 1 shown, the AutoML platform provides users with services such as data processing, model construction, model training, evaluation, and hyperparameter tuning. Among them, model construction includes the design of the neural network structure.
[0051] Currently, most AutoML platforms have limited support for the automation of network structure design. For different tasks, such as image classification and natural language understanding, designing a neural network structure usually requires a large amount of structural engineering and technical knowledge, and often requires experts with rich experience to spend a lot of time and effort to design the network structure. Therefore, neural architecture search (NAS) emerged. Its main task is to automate the process of manually designing neural network structures. NAS can automatically search for neural networks that meet specific constraints and achieve specific goals using a specific dataset. That is to say, users can complete the process of using deep neural networks for modeling without scenario experience and knowledge and skills of deep learning. NAS constitutes the differential competitiveness of the AutoML platform.
[0052] A schematic diagram of the system framework of neural architecture search (NAS) is as Figure 2 shown. As Figure 2 shown, the neural architecture search (NAS) system is mainly composed of a search space definition module, a structure parameter learning module, and a sub-network evaluation module.
[0053] Among them, the search space definition module selects a suitable search space according to the application scenario goal and the dataset, and introduces prior knowledge. For example, for the image classification task, the search space is defined as the repeated stacking structure of a classic classification network, such as Figure 3 shown in the schematic diagram of the structure of a classic image classification convolutional neural network, which will be described below. The structure parameter learning module uses a specific strategy to generate samples of the sub-network structure. The sub-network evaluation module obtains the sampled network, evaluates the performance of the sampled network, and feeds it back to the structure parameter learning module to update its strategy, so that the sub-network structure generated next time is closer to the goal. Repeat this process multiple times until a target network with satisfactory performance is obtained.
[0054] Figure 3 A schematic diagram of the structure of a classic image classification convolutional neural network. As Figure 3 shown, in a classic image classification convolutional neural network, the network structure parameters include: multiple discrete variables such as image resolution, width, depth, topology, and operator type. The network structure parameters in each dimension are interdependent and correlated. The performance evaluation of the network structure, especially the accuracy, cannot be achieved through simple calculations. It often requires long-term training of the network and then inference on the validation set of the target device. Therefore, neural architecture search faces problems such as high-dimensional search space, long search time, and complex performance evaluation process.
[0055] To address the problems faced by traditional NAS mentioned above, efficient neural architecture search (ENAS) has been proposed in the industry. Compared with traditional NAS, ENAS accelerates the search process based on reinforcement learning and weight sharing. For example, it can automatically construct a neural network model with better performance than human-designed ones within a day.
[0056] The following combines Figure 4 with Figure 5 to give a brief description of ENAS.
[0057] Figure 4 is a schematic diagram of the network structure. As Figure 4 shown, ENAS represents the neural network structure as a directed acyclic graph, where the nodes ( Figure 3 the circles where 1, 2, 3 and 4 are located in Figure 3 ) represent computational types, and the edges ( Figure 3 the connections between each circle in
[0058] ) represent the data flow between nodes). For example, the computational types represented by the nodes can include: 3*3 convolution, separable 3*3 convolution, 5*5 convolution, separable 5*5 convolution, max pooling, average pooling, direct connection, etc.
[0058] For example, a sub-network structure corresponds to Figure 4 a sub-graph in Figure 4 as shown by the solid line in
[0059] Continuing to refer to Figure 2 , when the neural network architecture search is ENAS, Figure 2 the structure parameter learning module shown in
[0060] uses a recurrent neural networks (RNN) model to sample sub-network structures.
[0061] In other words, ENAS includes two types of neural networks:
[0062] 1) Recurrent neural networks (RNN), which are used for the generation and learning of sub-network structures. The recurrent neural networks (RNN) can be an agent using reinforcement learning. The agent can also be called a controller.
[0062] 2) Sub-network structures, which represent the neural network structures to be established.
[0063] In other words, in ENAS, a neural network is used to train the neural network structure. Among them, the neural network used to train the neural network structure can be called a controller or an agent.
[0064] Figure 5 is a schematic diagram of the process for ENAS to generate sub-network structures. As Figure 5As shown, the output of a cycle forms a sub-network structure, where the output of the previous step is the input of the next step, and the structure of the (N + 1)-th step is predicted based on the states of the previous N steps. Among them, the prediction module uses an agent based on reinforcement learning, such as Figure 5 the agent indicated therein. The agent sequentially determines the input nodes and calculation types of each node. The performance evaluation of the sub-network structure updates the agent policy as a reward. Figure 5 The agents shown in Figure 5 are the same agent. In other words, in
[0065] a single agent is used to sequentially learn and generate the structure of each node.
[0066] The agent used to train the neural network structure is based on a reinforcement learning algorithm. Reinforcement learning guides the policy to approach the target network along a faster path through reward feedback, improving the search efficiency. Figure 4 To avoid the time-consuming evaluation process caused by training each sub-network structure from scratch until the weights converge, ENAS utilizes the structural similarity among sub-networks to maintain a set of weights for the nodes with the same calculation in
[0067] and the sub-network structures share the weights. During each evaluation of the sub-network structure, the relevant parameters are taken from the maintained shared-weight super-network for a short training update and then evaluated on the validation set. Parameter sharing can quickly obtain the rewards of the sub-network.
[0068] Compared with traditional NAS algorithms, ENAS has a 1000-fold improvement in search efficiency. However, in a single image classification scenario, the network structure search still takes more than ten hours, limiting its application as a general service in various scenarios.
[0069] To address the above problems, the present application proposes a neural network search solution, which can effectively improve the overall search efficiency of neural network structures compared with the prior art.
[0070] Referring to Figure 1 , the present application can improve the model construction in the AutoML platform. Referring to Figure 2 , the present application can improve the structure parameter learning module and the interaction between the structure parameter learning module and the sub-network evaluation module in the existing neural network structure search system. The embodiments of the present application will be described below.
[0071] It should be noted that in the embodiments of the present application, a neural network used for training a neural network structure is taken as an example of an agent for description.
[0072] It should be noted that in some descriptions herein, the neural network structure will be abbreviated as the network structure. In other words, the network structure mentioned herein refers to the neural network structure.
[0073] As Figure 6 shown, an embodiment of the present application provides a system architecture 600. The system architecture includes a local device 630, a local device 640, an execution device 610, and a data storage system 620, wherein the local device 630 and the local device 640 are connected to the execution device 610 through a communication network.
[0074] The execution device 610 can be implemented by one or more servers. Optionally, the execution device 610 can be used in cooperation with other computing devices, such as: devices such as data memories, routers, load balancers, etc. The execution device 610 can be arranged on one physical site or distributed on multiple physical sites. The execution device 610 can use the data in the data storage system 620 or call the program code in the data storage system 620 to implement the method for searching the neural network structure in the embodiments of the present application.
[0075] Specifically, the execution device 610 can execute the following process: determine a plurality of candidate neural networks, the plurality of candidate neural networks having the same network structure, a first agent among the plurality of agents being used to process the same part in each of the plurality of candidate neural networks, the first agent being one of the plurality of agents; respectively use the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent to obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks, the new candidate neural networks including the part of the network structure processed by the first agent and the input, wherein the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; determine a target neural network according to the plurality of new candidate neural networks.
[0076] Through the above process, the execution device 610 can obtain at least one target neural network, and the target neural network can be used for image classification or image processing, etc.
[0077] Users can operate their respective user devices (such as the local device 630 and the local device 640) to interact with the execution device 610. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc.
[0078] The local device of each user can interact with the execution device 610 through a communication network using any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.
[0079] In one implementation, the local devices 630 and 640 obtain the relevant parameters of the target neural network from the execution device 610, deploy the target neural network on the local devices 630 and 640, and use the target neural network for image classification or image processing, etc.
[0080] In another implementation, the target neural network can be directly deployed on the execution device 610. The execution device 610 obtains the images to be processed from the local devices 630 and 640, and classifies or performs other types of image processing on the images to be processed according to the target neural network.
[0081] The above-mentioned execution device 610 can also be referred to as a cloud device. At this time, the execution device 610 is generally deployed in the cloud.
[0082] Next, in combination with Figure 7 A detailed introduction to the method of neural network search in the embodiments of the present application will be given. Figure 7 The method shown can be executed by a neural network search device, which can be a device such as a computer, a server, or a cloud device with sufficient computing power to implement the search of the neural network.
[0083] Figure 7 The method shown includes steps 701 to 704. This method is applied to a computing system, which includes multiple agents. The following will describe these steps in detail.
[0084] S701, determine multiple candidate neural networks. The multiple candidate neural networks have the same network structure. The first agent among the multiple agents is used to process the same part in each candidate neural network among the multiple candidate neural networks. The first agent is one of the multiple agents.
[0085] S702, respectively use the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent, and obtain multiple new candidate neural networks corresponding to the multiple candidate neural networks. The new candidate neural networks include the part of the network structure processed by the first agent and the input. Among them, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent.
[0086] This step describes the process of training the first agent using multiple candidate neural networks. Among them, taking the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent means that in a candidate neural network, taking the context of this candidate neural network as the input of the first agent, and the above operation is performed for each candidate neural network among the multiple candidate neural networks.
[0087] S703. Determine the target neural network according to multiple new candidate neural networks.
[0088] Among them, that multiple candidate neural networks have the same network structure means that the structures of multiple candidate neural networks are network structures applied to the same scenario, have the same network structure type, for example, they are all multi-layer convolutional network structures, and have the same number of network layers, and each network layer has the same number of nodes.
[0089] Among them, for the description of the network structure context, taking one of the multiple agents (denoted as agent A) as an example, assume that agent A is responsible for part of the network structure a in the neural network. The input of agent A is the other part of the network structure in the neural network except part of the network structure a, and the output of agent A is the new network structure formed by part of the network structure a and the input. Among them, the other part of the network structure in the neural network except part of the network structure a can be called the network structure context of agent A.
[0090] In the prior art, when training an agent, only the association between the current structure and the previous structure is considered.
[0091] In the embodiments of the present application, a multi-agent generated neural network is adopted. Among them, the input of each agent is the context network structure. Therefore, in the process of training an agent responsible for part of the network structure, not only the association between the current network structure and the previous network structure is considered, but also the association between the current network structure and the subsequent network structure is considered, so as to fully consider the influence of the network structure context on the current network structure, and the probability of searching for a better network structure can be improved.
[0092] In the embodiments of the present application, a multi-agent generated neural network is adopted. Among them, each agent is responsible for the generation and learning of part of the network structure in the neural network.
[0093] For the description of part of the network structure, taking one of the multiple agents (denoted as agent A) as an example, assume that agent A is responsible for part of the network structure a in the neural network. The input of agent A is the other part of the network structure in the neural network except part of the network structure a, and the output of agent A is the new network structure formed by part of the network structure a and the input.
[0094] In the neural network mentioned in the above example, the other partial network structures except for the partial network structure a can be called the context network structure.
[0095] In the existing ENAS, a single-agent policy gradient model is used to predict the network structure sequence, resulting in a large search space, making it difficult to learn the optimal path. In addition, a large number of network structure samples are required.
[0096] In the embodiments of the present application, a multi-agent generative neural network is adopted, where each agent is responsible for the generation and learning of a partial network structure. Therefore, the search space of each agent can be reduced, which is beneficial to learning the optimal path, thereby improving the efficiency of establishing a neural network.
[0097] In other words, in the embodiments of the present application, by adopting multiple agents to be respectively responsible for the generation of different parts of the network, compared with the existing ENAS, the sequence structure of network generation is improved. Improving the sequence structure of network generation can be understood as decoupling the timing of network structure generation in the existing network structure search scheme. Therefore, the embodiments of the present application can improve the efficiency of network structure sampling and overall improve the efficiency of neural network structure search.
[0098] The following uses an example to illustrate step 703. As an example, assume that the network structure X to be established includes partial network structure A + partial network structure B + partial network structure C. Three agents are used to generate this network structure, where agent 1 is responsible for the generation and learning of partial network structure A, agent 2 is responsible for the generation and learning of partial network structure B, and agent 3 is responsible for the generation and learning of partial network structure C.
[0099] The input of agent 1 is partial network structure B + partial network structure C, and the output is the input + the partial network structure responsible by agent 1, that is, the output is partial network structure A + partial network structure B + partial network structure C. The input of agent 2 is partial network structure A + partial network structure C, and the output is the input + the partial network structure responsible by agent 2, that is, the output is partial network structure A + partial network structure B + partial network structure C. The input of agent 3 is partial network structure A + partial network structure B, and the output is the input + the partial network structure responsible by agent 3, that is, the output is partial network structure A + partial network structure B + partial network structure C. Among them, partial network structure B + partial network structure C can be called the network structure context of agent 1, partial network structure A + partial network structure C can be called the network structure context of agent 2, and partial network structure A + partial network structure B can be called the network structure context of agent 3.
[0100] The multiple candidate neural networks are multiple samples of the above network structure X. For example, the multiple samples of the network structure X can be expressed as {{a0+b0+c0},{a1+b1+c1},…,{ai+bi+ci},…}.
[0101] An agent is selected from the three agents for training, for example, agent 2 is selected as an example. The training of agent 2 includes: training agent 2 according to multiple candidate neural networks, that is, multiple samples of network structure X.
[0102] By training agent 2, the optimization of part of network structure B in network structure X is completed.
[0103] If the convergence condition is met after completing the training of the training agent 2, the establishment of the network structure X is completed.
[0104] If the convergence condition is not met after completing the training of the training agent 2, continue to select an agent from the three agents for training until the convergence condition is met, and complete the establishment of the network structure X. In the embodiment of the present application, a multi-agent neural network is generated, wherein each agent is responsible for the generation and learning of a part of the network structure, and the input of each agent is the context network structure corresponding to the part of the network structure it is responsible for.
[0105] Each agent is responsible for generating and learning a portion of the network structure in the neural network. For example, the portion of the network structure can be called a portion of the network layer in the neural network, or can be called a portion of the node in the neural network, or can be called a portion of the instructions of the node in the neural network.
[0106] As an example, in a scenario where the search space of the neural network structure search is a macro search, the partial network structure may be a partial network layer in the neural network. Figure 8 In the multi-agent network generation diagram shown in the figure, multiple agents are responsible for Figure 8 Generation of a multi-layer neural network as shown, where each agent is responsible for one layer of the neural network.
[0107] As another example, in a scenario where the search space of the neural network structure search is a micro-search, the partial network structure may be a partial node in the neural network, or may be a partial instruction of a node in the neural network.
[0108] For example, in Figure 9 In the neural network for image classification shown in the figure, the network structure can be as follows Figure 9 As shown in the figure, it is composed of multiple identical normal units and multiple decay units repeatedly stacked, so the structure of the neural network can be determined by simply determining the structure of one normal unit and one decay unit. Among them, the network structure of the normal unit is as follows Figure 10As shown, the network structure of the attenuation unit is as follows Figure 11 As shown, the normal unit and the attenuation unit have similar network structures. Hereinafter, Figure 11 the network structure of the attenuation unit shown below is used to illustrate the nodes and instructions in the network structure.
[0109] As Figure 11 shown, the attenuation unit may include the inputs C[k - 2] and C[k - 1] from the previous two units and the output C[k] of this unit. Nodes 0, 1, 2, and 3 are the nodes referred to in the embodiments of the present application. A node represents the constituent structure of a unit in a candidate neural network. In Figure 11 the attenuation unit shown, it can receive the data output from the previous two units C[k - 2] and C[k - 1], and nodes 0, 1, 2, and 3 respectively process the input data, and C[k] finally outputs the data processed by this attenuation unit.
[0110] The arrows connecting the nodes represent inputs, and the operations corresponding to the arrows can be operators in a convolutional neural network. In the embodiments of the present application, the operators may include five types: 3×3 convolution (3×3conv), 5×5 convolution (5×5conv), max-pooling, avg-pooling, and identify operation. In the embodiments of the present application, the input and operator of a node are collectively referred to as the instruction of this node.
[0111] Optionally, in the embodiments of the present application, an agent is responsible for one instruction of one node of each candidate neural network.
[0112] Hereinafter, taking one candidate neural network among multiple candidate neural networks as an example, the meaning that one instruction of one node of each candidate neural network is responsible for by one agent among multiple agents is described.
[0113] One agent is responsible for each instruction of each node in the construction unit of a candidate neural network.
[0114] As an example, assume that the construction unit of a candidate neural network includes B nodes, and each node has 4 instructions. Then the 4B instructions are respectively responsible for by 4B agents, that is, this candidate neural network corresponds to 4B agents. The construction unit of the candidate neural network is, for example, a normal unit or an attenuation unit.
[0115] As another example, assume that the building block of a candidate neural network includes 2B nodes, and each node has 4 instructions. Then, 8B instructions are respectively responsible by 8B agents, that is, the candidate neural network corresponds to 8B agents. In this example, the building block of the candidate neural network includes, for example, two normal units, or one normal unit or one decay unit.
[0116] In some existing technologies, a single-agent policy gradient model is used for network structure sequence prediction, resulting in a large search space, making it difficult to learn the optimal path. In addition, a large number of network structure samples are required.
[0117] In the embodiments of the present application, a multi-agent generative neural network is adopted. Among them, each agent is responsible for the generation and learning of part of the network structure. When generating the network structure each time, only one agent takes an action, that is, only one agent participates in generating the network structure each time during training, so that the input of one agent can be independent of the output of other agents, thereby fully exerting the efficiency of multi-agent parallel sampling.
[0118] In other words, in the embodiments of the present application, by making the input of one agent independent of the output of other agents, the agent policy update mechanism is optimized.
[0119] In the following embodiments, part of the network structure is taken as an example of part of the instructions of the nodes in the neural network for description.
[0120] (1) Determine multiple candidate neural networks. Each candidate neural network among the multiple candidate neural networks includes multiple nodes, and one instruction of one node of each candidate neural network is responsible by one agent among multiple agents.
[0121] The instructions responsible by each agent correspond to part of the network structure in the neural network described above.
[0122] An agent is a predefined recurrent neural network that controls or directly constructs the structure of a neural network by making decisions, and these decisions are various operation types (such as convolution, pooling, etc.) that make up a specific network layer in the neural network. The neural network can be determined through these decisions.
[0123] (2) Select the first agent among the multiple agents.
[0124] Since one agent is responsible for one instruction of one node of each candidate neural network, therefore, an agent needs to be selected from multiple agents for training first. The selected agent is called the first agent here. The selection method of the first agent can be randomly selected from multiple agents, or the agents can be selected in sequence according to any set order.
[0125] (3) Train a first agent according to multiple candidate neural networks, where the first agents corresponding to different candidate neural networks among the multiple candidate neural networks are the same.
[0126] This step may correspond to the above S702 step, that is, the process of training the first agent.
[0127] Specifically, the multiple candidate neural networks include k first candidate neural networks, where k is a positive integer. The first agent determines the instruction it is responsible for for the first time. For example, if the instruction the first agent is responsible for is an input, determine the two nodes before and after the connection of the input; if the instruction the first agent is responsible for is a policy, determine that the policy is one of 3×3 convolution, 5×5 convolution, max pooling, average pooling, and direct connection operation. Among them, the first agents corresponding to different candidate neural networks are the same, that is, the instructions the first agent is responsible for in the k first candidate neural networks are all determined to be the same one.
[0128] Optionally, the multiple candidate neural networks include k first candidate neural networks and k second candidate neural networks, and the instructions the first agents in the k first candidate neural networks and the k second candidate neural networks are responsible for are all determined to be the same one. The other agents except the first agent in the k first candidate neural networks and the k second candidate neural networks are here called second agents. The instruction the second agent is responsible for in the first candidate neural network is called the first instruction, and the instruction the second agent is responsible for in the second candidate neural network is called the second instruction. Among them, the second instruction is obtained by perturbing the first instruction, and the perturbation here can be a random adjustment of the first instruction. Among them, the second instruction can be obtained by perturbing the first instruction after the first agent determines the instruction it is responsible for each time. This can avoid the training of the first agent only achieving the optimum in the k first candidate neural networks.
[0129] In the embodiment of the present application, when training an agent, randomly perturb the states of other agents in the network to obtain a group of new sub-network samples, which can increase the generalization ability of the network structure samples and the agent policies. The generalization ability represents the adaptability of the machine learning algorithm to fresh samples. This concept is prior art and will not be elaborated herein.
[0130] After the first intelligent agent determines the instruction it is responsible for for the first time, multiple candidate neural networks are evaluated. The evaluation method can be an existing method for evaluating neural networks, that is, training multiple candidate neural networks on a training set and then verifying the accuracy of multiple candidate neural networks on a validation set. According to the accuracy of multiple candidate neural networks, it can be determined whether there is a neural network that meets the requirements among multiple candidate neural networks. If not, the reward corresponding to the instruction determined by the first intelligent agent for the first time can be calculated based on the accuracy of multiple candidate neural networks.
[0131] The first intelligent agent determines the instruction it is responsible for for the second time. For example, if the instruction that the first intelligent agent is responsible for is an input, the two nodes before and after the connection of this input are re-determined; if the instruction that the first intelligent agent is responsible for is a policy, one is re-selected from 3×3 convolution, 5×5 convolution, max pooling, average pooling, and direct connection operations. Multiple candidate neural networks after the first intelligent agent determines the instruction it is responsible for for the second time are evaluated to determine the accuracy of multiple candidate neural networks. According to the accuracy of multiple candidate neural networks, it can be determined whether there is a neural network that meets the requirements among multiple selected neural networks. If not, the reward corresponding to the instruction determined by the first intelligent agent for the second time can be calculated based on the accuracy of multiple candidate neural networks.
[0132] Similarly, the reward corresponding to the instruction determined by the first intelligent agent for the third time can be determined. Among them, the number of times the first intelligent agent determines the instruction it is responsible for can be preset manually, or it can end after obtaining a neural network that meets the requirements, or it can end after determining all possible instructions it is responsible for.
[0133] According to the instructions determined by the first intelligent agent multiple times and the corresponding rewards, the parameters of the policy network are determined. Then, taking the instruction determined by the first intelligent agent for the first time and the other instructions in multiple candidate neural networks except those the first intelligent agent is responsible for as the input of the policy network, the output of the policy network corresponding to the instruction determined by the first intelligent agent for the first time can be obtained; taking the instruction determined by the first intelligent agent for the second time and the other instructions in multiple candidate neural networks except those the first intelligent agent is responsible for as the input of the policy network, the output of the policy network corresponding to the instruction determined by the first intelligent agent for the second time can be obtained; similarly, the output of the policy network corresponding to the instruction determined by the first intelligent agent for the third time can be obtained... According to the output of the policy network, the optimal instruction of the first intelligent agent is determined. Among them, the instruction corresponding to the largest output value of the policy network is the optimal instruction of the first intelligent agent.
[0134] In the prior art, when training an intelligent agent, only the association between the current structure and the previous structure is considered.
[0135] In the embodiment of the present application, a multi-agent generation neural network is adopted. Among them, the input of each agent is the context network structure. Therefore, in the process of training an agent responsible for a partial network structure, not only the association between the current network structure and the previous network structure is considered, but also the association between the current network structure and the subsequent network structure is considered, so as to fully consider the influence of the network structure context on the current network structure and improve the probability of searching for a better network structure.
[0136] (4) Determine the neural network according to the trained first agent.
[0137] In (3), the optimal instruction of the first agent has been determined. Then, the first agents in multiple candidate neural networks are all updated to the optimal instruction, and then the accuracy rates of the multiple candidate neural networks are evaluated. Determine whether there is a neural network that meets the requirements according to the accuracy rates of the multiple candidate neural networks.
[0138] If there is no neural network that meets the requirements, reselect the first agent from multiple agents, train the first agent, and determine the neural network that meets the requirements.
[0139] In the embodiment of the present application, each instruction of a node in the network structure is responsible for by one agent among multiple agents. Train the first agent among multiple agents according to multiple candidate neural networks. Among them, for the first agent, other agents are obtained by sampling from multiple candidate neural networks.
[0140] Optionally, when selecting the first agent, the first agent can also be selected according to the probability value corresponding to each agent among multiple agents. Among them, the probability value corresponding to each agent is the probability that each agent is selected.
[0141] In the embodiment of the present application, when the candidate neural network is Figure 9 repeatedly stacked by multiple normal units and multiple attenuation units as shown, the number of repetitions of the normal units is greater than the number of repetitions of the attenuation units. In order to fully train the agents in the normal units and the agents in the attenuation units, the probability that the agents in the normal units are selected can be preset to be greater than the probability that the agents in the attenuation units are selected. In addition, as Figure 11In the structure of the attenuation unit shown, the input of node 0 may come from the outputs of C[k - 1] and C[k - 2], the input of node 1 may come from the outputs of C[k - 1], C[k - 2] and node 0, the input of node 2 may come from the outputs of C[k - 1], C[k - 2], node 0 and node 1, and so on. The input of a node with a later order may come from any one or more of the outputs of all the previous nodes. Therefore, more training data is required during training. Thus, in the embodiments of the present application, it can be pre - specified that the probability of selecting the agent corresponding to the instruction of a node with a later order is greater, so as to ensure that each agent can be fully trained. The method of the embodiments of the present application further includes that when determining multiple candidate neural networks, multiple candidate neural networks can be determined according to a common network structure pool. The common network structure pool can be determined according to the results of the previous training of the agents. For example, update the first agent in multiple candidate neural networks to the optimal instruction, then evaluate the accuracy rates of multiple candidate neural networks, sort multiple candidate neural networks according to the high - low order of the accuracy rates, retain some candidate neural networks with high accuracy rates in the common network structure pool, and then randomly generate some new candidate neural networks in the common network structure pool.
[0142] The current prior art has also proposed a multi - agent mechanism, but each agent does not cooperate well, and it is necessary to allocate the overall evaluation reward of the network to each agent through a transformation matrix.
[0143] In the embodiments of the present application, in order to achieve the mutual cooperation between agents, a cooperation mechanism based on genetic evolution is proposed, maintaining a common structure pool to save the optimal networks generated by agents. Each agent samples in the common structure pool, inherits the optimal decisions of other agents, and updates the structure pool by training itself on this basis. This cooperation mechanism can make the overall network update in a direction with better performance while the training of multiple agents is updated in a better direction respectively. In other words, this cooperation mechanism can make the training of multiple agents converge to the optimal direction of the overall network, rather than just the local optimum.
[0144] During the training of the first agent, if one or more candidate neural networks that meet the preset requirements appear, output the one or more candidate neural networks that meet the preset requirements as the neural networks that meet the requirements. If the iteration times reach the preset value and no candidate neural network that meets the preset requirements has been determined, output the existing candidate neural networks sorted by the high - low order of the accuracy rates.
[0145] Optionally, the method provided in the embodiments of the present application can also select multiple agents simultaneously and train the multiple agents.
[0146] Next, takeFigure 12 Specific examples are used to elaborate in detail on the neural network search method provided by the embodiments of the present application.
[0147] The input of neural network search includes sample data, which is used for the training of the neural network. Neural network search first needs to define a search space, which is the complete set of possible network structures. Neural network search is to find a neural network that meets the requirements from the complete set. Since in the embodiments of the present application, the candidate neural network can be Figure 9 repeatedly stacked by multiple normal units and multiple attenuation units as shown, that is, each normal unit in the candidate neural network is the same, and each attenuation unit is the same. Therefore, the search space can be simplified to normal units and attenuation units. Among them, the network structures of normal units and attenuation units are similar. The network structure of the attenuation unit is as shown in Figure 11 Each unit has 2 inputs, which are the outputs of the previous two units C[k - 1] and C[k - 2] respectively. For ease of description, it is defined here that each normal unit and each attenuation unit include L intermediate nodes, each intermediate node has two inputs j1, j2, and each intermediate node has two operators o1, o2. Thus, it can be encoded into a 4-bit code (j1, j2, o1, o2). The output of each node can be expressed as The output of each unit is the sum of the outputs of all nodes.
[0148] The following is the process of neural network search using the method of the embodiments of the present application.
[0149] Randomly generate k candidate neural networks to form a candidate neural network pool P. Each candidate neural network in the k candidate neural networks is composed of normal units and attenuation units. Since each unit includes L intermediate nodes and each intermediate node includes a 4-bit code (j1, j2, o1, o2), each unit can be encoded with 4L bits. Furthermore, since each unit is divided into two cases: normal unit and attenuation unit, each candidate neural network in the k candidate neural networks can be encoded with 8L bits. Correspondingly, 8L agents can be generated to be responsible for optimizing the inputs and operators (j1, j2, o1, o2) of the L intermediate nodes of each candidate neural network in the k candidate neural networks respectively. The selection of each input and each operator for each intermediate node is regarded as the action of the agent. Among them, each agent sets a stability value v, and the initial stability value can be set to 0.5.
[0150] The sampling agent can randomly select an agent for training during the training process of the agent. However, in each candidate neural network, the number of repetitions of normal units is greater than that of decay units. Therefore, the agents in normal units should receive more training times. Thus, in the embodiments of this application, it is stipulated that when selecting an agent, the probability of an agent in a normal unit being selected is greater than the probability of an agent in a decay unit being selected. Exemplarily, it can be stipulated that the probability of a normal unit being selected is 0.7, and correspondingly, the probability of a decay unit being selected is 0.3. Since within each unit, the selection of the input for each agent includes the outputs of all previous nodes. For example, the intermediate node 2 can select the outputs of the intermediate node 0 and / or the intermediate node 1 as inputs, and the intermediate node 3 can select the outputs of the intermediate node 0, the intermediate node 1, and / or the intermediate node 2 as inputs. Therefore, the later the order of the agent, the more actions are available for selection, the more complex the policy network is, and the more sample data is required. Thus, the later the order of the agent, the more training times it should receive. Therefore, in the embodiments of this application, it is stipulated that among the 4L agents within each unit, the later the position of the agent, the greater the probability of being selected. Specifically, the probability of the agent at the i-th position being selected can be:
[0151]
[0152] Thus, it can be ensured that each agent can receive sufficient training as much as possible.
[0153] After selecting an agent, the agent makes predictions on k candidate neural networks, that is, the agent randomly takes an action, and can generate a candidate neural network group p(i,x). Here, the agent taking an action is the agent determining the instruction it is responsible for in the above text. For example, select Figure 11The agent a responsible for one of the operators of node 0 in the shown unit. Assume the actions that agent a can take include avg - pooling and max - pooling. In this training, agent a randomly takes an action, for example, avg - pooling. Then, for each candidate neural network in the k candidate neural networks, one of the operators of node 0 in the decay unit of each candidate neural network is avg - pooling. These k candidate neural networks form the candidate neural network group p(i,x). If agent a is trained only based on the candidate neural network group p(i,x), the optimal action of agent a is only obtained according to the candidate neural network group p(i,x). Therefore, to avoid the training result of agent a being only optimal in p(i,x), and also to evaluate the influence of the states of other agents on the selection of the optimal action of agent a, the states of other agents except agent a in the candidate neural network group p(i,x) are randomly adjusted, that is, the input of node 0, another operator, and the operators and inputs of other nodes are randomly adjusted. The probability of adjustment for each agent is (1 - v). After adjustment, the candidate neural network group p'(i,x) can be obtained. The candidate neural network group p(i,x) and the candidate neural network group p'(i,x) form the candidate neural network set Q to be evaluated.
[0154] Evaluate the accuracy of the candidate neural networks in the candidate neural network set Q. To improve the search efficiency, in the embodiments of the present application, the parameter sharing method is adopted, that is, the parameters of all candidate neural networks are updated to the super network. For each training of each candidate neural network, a candidate neural network is randomly selected from the candidate neural network set Q to be trained on the training data set and the corresponding parameters are updated. Then, using the shared parameters in the super network, the accuracy of the candidate neural network is verified on the validation data set, which can greatly save time.
[0155] Calculate the reward when agent a takes the action of avg - pooling. This reward can be calculated according to the following formula:
[0156]
[0157] where Acc(p(i,x)) is the accuracy of each candidate neural network in the candidate neural network group p(i,x), Acc(p'(i,x)) is the accuracy of each candidate neural network in the candidate neural network group p'(i,x), and β is a value set artificially.
[0158] Thus, the reward when agent a takes the action of avg - pooling can be obtained. Repeating the above steps, the reward when agent a takes the action of max - pooling can be obtained.
[0159] Repeat the above steps to train the agent, and the actions that the agent may take and the corresponding rewards for taking the actions can be obtained.
[0160] According to the actions taken by the above agent a and the corresponding rewards, calculate the loss of the policy network and update the parameters of the policy network.
[0161] Update the structure pool. Determine the candidate neural network set Q'. The candidate neural network set Q' includes the candidate neural network group p(i,x) and other candidate neural network groups p'(i,x) in p(i,x). Among them, p(i,x) can be the candidate neural network group generated by the agent a taking the action avg - pooling or the action max - pooling, and p'(i,x) can be the candidate neural network group generated by randomly adjusting other agents in p(i,x) except the agent a. According to the parameters of the updated policy network, determine the optimal action of the agent a in each candidate neural network in the candidate neural network set Q. Specifically, keep the states of other agents except the agent a in each candidate neural network unchanged, use the states of other agents except the agent a in each candidate neural network as the input of the policy network, calculate the output of the policy network when the agent a takes each action, and evaluate the optimal action of one of the operators of the decay unit in node 0 in each candidate neural network in the candidate neural network set Q' according to the output result of the policy network. Thus, the candidate neural network group p”(i,x) and the candidate neural network group p”'(i,x) can be obtained, where p”(i,x) is obtained by the agent a in each candidate neural network in p(i,x) taking the optimal action, and p”'(i,x) is obtained by the agent a in each candidate neural network in p'(i,x) taking the optimal action. p”(i,x) and p”'(i,x) form the new candidate neural network set Q”.
[0162] According to the validation data, verify the accuracy of each candidate neural network in the candidate neural network set Q” on the validation set, and sort the candidate neural networks according to the accuracy from high to low. In the embodiment of the present application, select the candidate neural networks ranked in the top k / 2 in terms of accuracy, and randomly regenerate k / 2 new candidate neural networks to form the new candidate neural network pool P'.
[0163] Update the stability value of the agent a. Randomly generate m candidate neural networks as the policy input of the agent a, calculate the policy output when the agent a takes the determined optimal action, and this output is a probability value. Take the maximum value as the new stability value of the agent a.
[0164] If the maximum number of iterations is not reached or no candidate neural network that meets the requirements appears, another agent, such as agent b, is selected. The optimal action of agent b is determined according to the new candidate neural network pool P', and a new candidate neural network pool is generated. Iteration is performed according to the above steps until the maximum number of iteration steps is completed, for example, the optimal actions of all agents are determined, or a candidate neural network that meets the performance requirements appears in the structure pool.
[0165] The candidate neural network that meets the performance requirements can be retrained, and this retraining is to train the network structure according to the training samples in the prior art to determine the network parameters. For the sake of simplicity, the embodiments of the present application will not elaborate further.
[0166] Output. If a candidate neural network that meets the performance requirements appears in the structure pool, the candidate neural network is retrained and then the neural network is output, which includes the network structure and network parameters. If no candidate neural network that meets the performance requirements appears in the structure pool but the maximum number of iteration steps is reached, the candidate neural networks in the structure pool are sorted according to the accuracy rate, and then the candidate neural networks in the structure pool are retrained respectively, or a part of the candidate neural networks with high accuracy rate is selected for retraining, and finally the neural networks can be output according to the accuracy rate from high to low.
[0167] To illustrate the effect of the neural network structure search method in the embodiments of the present application, the neural network structure search method in the embodiments of the present application is compared with the existing solutions below. Table 1 shows the results after the neural network structures searched by the neural network structure search method in the embodiments of the present application and other existing search methods under similar constraint conditions are trained on the CIFAR-10 dataset and then tested on the test set. Other search methods include neural architecture search (NAS), efficient neural architecture search (ENAS), differentiable architecture search (DARTS), AmoebaNet neural network structure search, neural architecture optimization (NAO), and multi-agent neural network search.
[0168] Table 1 Performance comparison of neural network structure search algorithms
[0169] Search algorithm Error rate (%) Search time (GPU days) Number of parameters (M) NAS 3.65 22400 37.4 ENAS 2.89 0.67 4.6 DARTS 2.76 4 3.3 AmoebaNet 2.13 3150 34.9 NAO 2.11 200 128 Multi-agent neural network search 2.52 4 3.4 This application 2.84 0.19 3.02
[0170] As can be seen from Table 1, compared with other neural network architecture search methods, the neural network architecture search method provided by the embodiments of the present application searches for neural networks with comparable accuracy, takes the shortest time, and also uses fewer parameters. Compared with the ENAS algorithm based on single-agent reinforcement learning, the neural network architecture search method provided by the embodiments of the present application searches for neural networks with comparable accuracy, the search time is shortened by 67%, and the number of parameters is reduced by 8%.
[0171] Figure 13 It is a schematic hardware structure diagram of the neural network search device provided by the embodiments of the present application. Figure 13 The shown neural network search device 1300 (this device 1300 can specifically be a computer device) includes a memory 1301, a processor 1302, a communication interface 1303, and a bus 1304. Among them, the memory 1301, the processor 1302, and the communication interface 1303 are communicatively connected to each other through the bus 1304.
[0172] The memory 1301 can be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1301 can store a program. When the program stored in the memory 1301 is executed by the processor 1302, the processor 1302 is used to execute each step of the neural network search method of the embodiments of the present application.
[0173] The processor 1302 can adopt a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits, and is used to execute relevant programs to implement the neural network search method of the method embodiments of the present application.
[0174] The processor 1302 can also be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the neural network search method of the present application can be completed by the integrated logic circuit in the hardware of the processor 1302 or by instructions in software form.
[0175] The above-mentioned processor 1302 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1301, and the processor 1302 reads the information in the memory 1301 and combines its hardware to complete the functions required to be executed by the units included in this neural network structure search device, or executes the neural network search method in the method embodiment of the present application.
[0176] The communication interface 1303 uses a transceiver device such as, but not limited to, a transceiver to implement the communication between the device 1300 and other devices or communication networks. For example, information about the neural network to be constructed and the training data required during the construction of the neural network can be obtained through the communication interface 1303.
[0177] The bus 1304 can include a path for transmitting information between various components of the device 1300 (for example, the memory 1301, the processor 1302, the communication interface 1303).
[0178] Figure 14 It is a schematic flowchart of the image processing method in the embodiments of the present application. It should be understood that the relevant content limitations, explanations, and expansions of the method shown above Figure 7 also apply to the method shown in Figure 14 appropriately omitting repeated descriptions when introducing the method shown in Figure 14 below. Figure 14 The method shown in
[0179] 1401. Obtain the image to be processed;
[0180] 1402. Classify the image to be processed according to the target neural network to obtain the classification result of the image to be processed.
[0181] Among them, the neural network is determined according to the trained first agent. The training of the first agent includes: determining a plurality of candidate neural networks; selecting the first agent from a plurality of agents. Among them, each agent in the plurality of agents is responsible for a partial network structure of the neural network. The input of each agent is the network structure context, and the output is the new network structure formed by the partial network structure it is responsible for and the input; training the first agent according to the plurality of candidate neural networks.
[0182] Optionally, the partial network structure of the neural network that the first agent is responsible for is an instruction of a node of the neural network.
[0183] Among them, selecting the first agent from a plurality of agents includes: selecting the first agent according to the probability values corresponding to each agent in the plurality of agents.
[0184] Each candidate neural network includes a normal unit and a decay unit. The probability value of the agent corresponding to the normal unit is higher than the probability value of the agent corresponding to the decay unit. Among them, the normal unit and the decay unit include multiple nodes.
[0185] The plurality of candidate neural networks includes k first candidate neural networks, or the plurality of candidate neural networks includes k first candidate neural networks and k second candidate neural networks, where k is a positive integer; among them, the first instruction of the first candidate neural network is randomly initialized or determined by the result of the previous training of the agent, and the second instruction of the second candidate neural network is obtained by perturbing the first instruction. Among them, both the first instruction and the second instruction are responsible for the second agent, and the second agent is other agents except the first agent among the plurality of agents.
[0186] Determining a plurality of candidate neural networks includes: determining a plurality of candidate neural networks according to the common network structure pool, where the common network structure pool is updated according to the training result of the previous training of the agent.
[0187] Figure 15 It is a schematic diagram of the hardware structure of the image processing device according to an embodiment of the present application. Figure 15 The illustrated image processing device 1500 includes a memory 1501, a processor 1502, a communication interface 1503, and a bus 1504. Among them, the memory 1501, the processor 1502, and the communication interface 1503 are communicatively connected to each other through the bus 1504.
[0188] The memory 1501 can be a ROM, a static storage device, and a RAM. The memory 1501 can store a program. When the program stored in the memory 1501 is executed by the processor 1502, the processor 1502 and the communication interface 1503 are used to execute each step of the image processing method according to an embodiment of the present application.
[0189] The processor 1502 can be a general-purpose CPU, microprocessor, ASIC, GPU, or one or more integrated circuits, which are used to execute relevant programs to implement the functions required by the units in the image processing apparatus according to the embodiments of the present application, or to execute the image processing method according to the method embodiments of the present application.
[0190] The processor 1502 can also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the image processing method according to the embodiments of the present application can be completed by the integrated logic circuit in the hardware of the processor 1502 or by instructions in the form of software.
[0191] The above-mentioned processor 1502 can also be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly implemented by the execution of the hardware decoding processor, or can be implemented by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 1501, and the processor 1502 reads the information in the memory 1501 and combines its hardware to complete the functions required by the units included in the image processing apparatus according to the embodiments of the present application, or to execute the image processing method according to the method embodiments of the present application.
[0192] The communication interface 1503 uses a transceiver device such as, but not limited to, a transceiver to implement the communication between the device 1500 and other devices or communication networks. For example, the image to be processed can be obtained through the communication interface 1503.
[0193] The bus 1504 can include a path for transmitting information between various components of the device 1500 (for example, the memory 1501, the processor 1502, the communication interface 1503).
[0194] It should be noted that although the above-described apparatuses 1300 and 1500 only show a memory, a processor, and a communication interface, in actual implementation, those skilled in the art should understand that apparatuses 1300 and 1500 may also include other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that apparatuses 1300 and 1500 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that apparatuses 1300 and 1500 may also only include the devices necessary for implementing the embodiments of the present application, and do not necessarily include Figure 13 and Figure 15 all the devices shown in
[0195] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0196] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0197] In several embodiments provided by the present application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.
[0198] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0199] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.
[0200] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0201] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for neural network search, the method being applied to a computing system, the system including a plurality of agents, characterized in that, Including: Determine a plurality of candidate neural networks, the plurality of candidate neural networks having the same network structure, a first agent among the plurality of agents being used to process the same part in each of the plurality of candidate neural networks, the first agent being one of the plurality of agents; Respectively use the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent, to obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks, the new candidate neural networks including the part of the network structure processed by the first agent and the input, wherein, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; Determine a target neural network according to the plurality of new candidate neural networks, the target neural network being used to classify the image to be processed to obtain the classification result of the image to be processed.
2. The method according to claim 1, characterized in that, The partial network structure of the neural network responsible for by the first agent is an instruction of a node of the neural network.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Select the first agent according to the probability values corresponding to each agent among the plurality of agents.
4. The method according to claim 1 or 2, characterized in that, Each of the candidate neural networks includes a normal unit and a decay unit, the probability value of the agent corresponding to the normal unit being higher than the probability value of the agent corresponding to the decay unit, wherein, the normal unit and the decay unit include a plurality of nodes.
5. The method according to claim 1 or 2, characterized in that, The plurality of candidate neural networks include k first candidate neural networks, or, the plurality of candidate neural networks include k first candidate neural networks and k second candidate neural networks, k being a positive integer; Wherein, the first instruction of the first candidate neural network is randomly initialized or determined by the result of the previous training of the agent, and the second instruction of the second candidate neural network is obtained by perturbing the first instruction, Wherein, both the first instruction and the second instruction are responsible for by a second agent, the second agent being other agents among the plurality of agents except for the first agent.
6. The method according to claim 1 or 2, characterized in that, The determining the plurality of candidate neural networks includes: Determine the plurality of candidate neural networks according to a common network structure pool, wherein, the common network structure pool is updated according to the training result of the previous training of the agent.
7. An image processing method, characterized in that, Including: Obtain the image to be processed; Classify the image to be processed according to the neural network to obtain the classification result of the image to be processed; Wherein, the neural network is determined according to a plurality of agents in a computer system, the determination of the neural network including: Determine a plurality of candidate neural networks, the plurality of candidate neural networks having the same network structure, a first agent among the plurality of agents being used to process the same part in each of the plurality of candidate neural networks, the first agent being one of the plurality of agents; For each of the candidate neural networks, use the context of the part of the neural network processed by the first agent as the input to the first agent, to obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks. The new candidate neural networks include the part of the network structure processed by the first agent and the input. Wherein, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; Determine a target neural network according to the plurality of new candidate neural networks.
8. The method according to claim 7, wherein The partial network structure of the neural network responsible for by the first agent is an instruction of a node of the neural network.
9. The method according to claim 7 or 8, characterized in that, The method further includes: Select the first agent according to the probability values corresponding to each agent in the plurality of agents.
10. The method according to claim 7 or 8, characterized in that Each of the candidate neural networks includes a normal unit and a decay unit, and the probability value of the agent corresponding to the normal unit is higher than the probability value of the agent corresponding to the decay unit. Wherein, the normal unit and the decay unit include a plurality of nodes.
11. The method according to claim 7 or 8, characterized in that The plurality of candidate neural networks include k first candidate neural networks, or the plurality of candidate neural networks include k first candidate neural networks and k second candidate neural networks, where k is a positive integer; Wherein, the first instruction of the first candidate neural network is randomly initialized or determined by the result of the previous training of the agent, and the second instruction of the second candidate neural network is obtained by perturbing the first instruction, Wherein, both the first instruction and the second instruction are responsible for by a second agent, and the second agent is other agents in the plurality of agents except for the first agent.
12. The method according to claim 7 or 8, characterized in that, The determining the plurality of candidate neural networks includes: Determine the plurality of candidate neural networks according to a common network structure pool, where the common network structure pool is updated according to the training result of the previous training of the agent.
13. An apparatus for neural network search, the apparatus being applied to a computing system, the system including a plurality of agents, characterized in that, Includes: A first determining unit, configured to determine a plurality of candidate neural networks, the plurality of candidate neural networks having the same network structure, and a first agent in the plurality of agents is used to process the same part in each of the plurality of candidate neural networks, and the first agent is one of the plurality of agents; A processing unit, configured to use the context of the part of the neural network processed by the first agent in each of the candidate neural networks as the input to the first agent, to obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks. The new candidate neural networks include the part of the network structure processed by the first agent and the input. Wherein, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except for the part processed by the first agent; A second determining unit, configured to determine a target neural network according to the plurality of new candidate neural networks, and the target neural network is used to classify the image to be processed to obtain a classification result of the image to be processed.
14. The device according to claim 13, characterized in that, A partial network structure of the neural network responsible by the first agent is an instruction of a node of the neural network.
15. The device according to claim 13 or 14, characterized in that, The device further includes a selection unit, and the selection unit is configured to: select the first agent according to the probability values corresponding to each agent among the multiple agents.
16. The device according to claim 13 or 14, characterized in that Each of the candidate neural networks includes a normal unit and a decay unit, and the probability value of the agent corresponding to the normal unit is higher than the probability value of the agent corresponding to the decay unit, wherein the normal unit and the decay unit include multiple nodes.
17. The device according to claim 13 or 14, characterized in that, The multiple candidate neural networks include k first candidate neural networks, or the multiple candidate neural networks include k first candidate neural networks and k second candidate neural networks, where k is a positive integer; wherein, the first instruction of the first candidate neural network is randomly initialized or determined by the result of the previous training of the agent, and the second instruction of the second candidate neural network is obtained by perturbing the first instruction. wherein, both the first instruction and the second instruction are responsible by a second agent, and the second agent is other agents among the multiple agents except the first agent.
18. The device according to claim 13 or 14, characterized in that, The first determining unit is configured to: determine the multiple candidate neural networks according to a common network structure pool, wherein the common network structure pool is updated according to the training result of the previous training of the agent.
19. An apparatus for neural network search, the apparatus being applied to a computing system, the system including a plurality of agents, characterized in that, including: a memory for storing programs; a processor for executing the programs stored in the memory, and when the programs stored in the memory are executed, the processor is configured to execute the following processes: determine multiple candidate neural networks, the multiple candidate neural networks have the same network structure, and a first agent among the multiple agents is used to process the same part in each of the multiple candidate neural networks, and the first agent is one of the multiple agents; respectively use the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent, to obtain multiple new candidate neural networks corresponding to the multiple candidate neural networks, the new candidate neural networks include the processed part network structure and the input by the first agent, wherein the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except the part processed by the first agent; determine a target neural network according to the multiple new candidate neural networks, and the target neural network is used to classify the image to be processed to obtain a classification result of the image to be processed.
20. An image processing apparatus, characterized in that, including: a memory for storing programs; a processor for executing the programs stored in the memory, and when the programs stored in the memory are executed, the processor is configured to execute the following processes: obtain an image to be processed; classify the image to be processed according to a neural network to obtain a classification result of the image to be processed; wherein, the neural network is determined according to multiple agents in a computer system, and the determination of the neural network includes: Determine a plurality of candidate neural networks, the plurality of candidate neural networks having the same network structure, and a first agent among the plurality of agents being configured to process the same part in each of the plurality of candidate neural networks, the first agent being one of the plurality of agents; Respectively use the context of the part of the neural network processed by the first agent in each candidate neural network as the input of the first agent, and obtain a plurality of new candidate neural networks corresponding to the plurality of candidate neural networks. The new candidate neural networks include the part of the network structure processed by the first agent and the input. Wherein, the context of the part of the neural network processed by the first agent is the remaining candidate neural network in a candidate neural network except the part processed by the first agent; Determine a target neural network according to the plurality of new candidate neural networks.
21. A computer-readable storage medium, characterized in that, The computer-readable medium stores program code for a device to execute, and the program code includes methods for executing any one of claims 1-6 or 7-12.
22. A chip, characterized in that, The chip includes a processor and a data interface, and the processor reads instructions stored on a memory through the data interface to execute the method according to any one of claims 1-6 or 7-12.
Citation Information
Patent Citations
Ship electric-power system reconstruction method based on BDI (belief-desire-intention theory) multi-agent
CN103745120A
Multi-agent collaborative optimization-based photovoltaic micro source-containing active distribution network topology reconfiguration method
CN104036329A