Neural network structure search method, device, computer equipment and storage medium

By encoding the discrete features of the training neural network structure into continuous features, constructing a continuous search space, and training related networks based on reconstruction loss, the problem of low efficiency in traditional neural network structure search in discrete space is solved, and a more efficient search process is achieved.

CN113408721BActive Publication Date: 2025-05-16INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011567991.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-25
Publication Date
2025-05-16
Estimated Expiration
2040-12-25

AI Technical Summary

Technical Problem

Traditional neural network structure search methods search in discrete spaces, resulting in slow convergence speed and low efficiency.

Method used

By obtaining the training neural network structure, discrete structural features are extracted from the input graph neural network, encoded into continuous structural features, and constructing a continuous search space. Based on the reconstruction loss training graph neural network, encoding network and decoding network, the target search space is determined for search.

Benefits of technology

Searching in continuous search space is easier to converge than discrete spaces, improving the efficiency of neural network structure search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113408721B_ABST
    Figure CN113408721B_ABST
Patent Text Reader

Abstract

The present application relates to a neural network structure search method, device, computer equipment and storage medium. The method includes: inputting the training neural network structure into a graph neural network to obtain the corresponding discrete structural features; inputting the discrete structural features into the encoding network, encoding the discrete structural features into continuous structural features through the encoding network; decoding according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure; based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, training the graph neural network, the encoding network and the decoding network until the training stop condition is met, obtaining the target encoding network and the target decoding network, and determining the latent space corresponding to the target encoding network as the target search space; searching from the target search space according to the target search strategy to obtain the target structural features, and decoding through the target decoding network to obtain the target neural network structure. The use of this method can improve the efficiency of neural network structure search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a neural network structure search method, device, computer equipment and storage medium. Background Art

[0002] With the development of artificial intelligence technology, the design of neural network structures is transforming from manual design to machine automatic design. Neural Architecture Search (NAS) can help developers automatically search for the optimal neural network structure.

[0003] In traditional technology, neural network structure search is usually performed in discrete space. However, searching in discrete space is a black box optimization problem with a relatively slow convergence speed, resulting in low efficiency of neural network structure search. Summary of the invention

[0004] Based on this, it is necessary to provide a neural network structure search method, device, computer equipment and storage medium that can improve the efficiency of neural network structure search in response to the above technical problems.

[0005] A neural network structure search method, characterized in that the method comprises:

[0006] Get the training neural network structure;

[0007] Inputting the training neural network structure into a graph neural network to obtain discrete structural features corresponding to the training neural network structure;

[0008] Inputting the discrete structural features into an encoding network, and encoding the discrete structural features into corresponding continuous structural features through the encoding network;

[0009] Decoding is performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure;

[0010] Based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, the graph neural network, the encoding network and the decoding network are trained until a training stop condition is met, a target encoding network and a target decoding network are obtained, and a latent space corresponding to the target encoding network is determined as a target search space;

[0011] The target structural features are obtained by searching the target search space according to the target search strategy, and the target structural features are decoded by the target decoding network to obtain the target neural network structure.

[0012] A neural network structure search method and device, characterized in that the device comprises:

[0013] A training data acquisition module is used to acquire the training neural network structure;

[0014] A discrete encoding module, used for inputting the training neural network structure into a graph neural network to obtain discrete structural features corresponding to the training neural network structure;

[0015] A continuous encoding module, used for inputting the discrete structural features into an encoding network, and encoding the discrete structural features into corresponding continuous structural features through the encoding network;

[0016] A decoding template, used for decoding according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure;

[0017] A training module, used for training the graph neural network, the encoding network and the decoding network based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, until a training stop condition is met, a target encoding network and a target decoding network are obtained, and a latent space corresponding to the target encoding network is determined as a target search space;

[0018] A search module is used to search from the target search space according to a target search strategy to obtain target structural features, and to decode the target structural features through the target decoding network to obtain a target neural network structure.

[0019] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0020] Get the training neural network structure;

[0021] Inputting the training neural network structure into a graph neural network to obtain discrete structural features corresponding to the training neural network structure;

[0022] Inputting the discrete structural features into an encoding network, and encoding the discrete structural features into corresponding continuous structural features through the encoding network;

[0023] Decoding is performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure;

[0024] Based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, the graph neural network, the encoding network and the decoding network are trained until a training stop condition is met, a target encoding network and a target decoding network are obtained, and a latent space corresponding to the target encoding network is determined as a target search space;

[0025] The target structural features are obtained by searching the target search space according to the target search strategy, and the target structural features are decoded by the target decoding network to obtain the target neural network structure.

[0026] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0027] Get the training neural network structure;

[0028] Inputting the training neural network structure into a graph neural network to obtain discrete structural features corresponding to the training neural network structure;

[0029] Inputting the discrete structural features into an encoding network, and encoding the discrete structural features into corresponding continuous structural features through the encoding network;

[0030] Decoding is performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure;

[0031] Based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, the graph neural network, the encoding network and the decoding network are trained until a training stop condition is met, a target encoding network and a target decoding network are obtained, and a latent space corresponding to the target encoding network is determined as a target search space;

[0032] The target structural features are obtained by searching the target search space according to the target search strategy, and the target structural features are decoded by the target decoding network to obtain the target neural network structure.

[0033] The above-mentioned neural network structure search method, device, computer equipment and storage medium, by inputting the training neural network structure into the graph neural network, the structural information of the training neural network can be effectively extracted through the graph neural network to obtain the discrete structural features corresponding to the training neural network structure, and by inputting the discrete structural features into the encoding network, the discrete structural features can be encoded into continuous structural features through the encoding network, thereby constructing a continuous search space, decoding and reconstructing according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure, and finally based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the graph neural network, the encoding network and the decoding network are trained until the training stop condition is met to obtain the target encoding network and the target decoding network. Since the target encoding network and the target decoding network are obtained by training based on the reconstruction loss, the target encoding network can well learn the structural features of the training neural network, and the target decoding network can accurately decode and reconstruct the structural features from the target search space to obtain the target neural network structure, so the latent space of the target encoding network can be determined as the target search space. Since the target search space is a continuous space, searching in the continuous space is easier to converge than searching in the discrete space, thereby improving the efficiency of the neural network structure search. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A schematic diagram of a process flow of a neural network structure search method in one embodiment;

[0035] Figure 2 A schematic diagram of a flow chart of a neural network structure search method in another embodiment;

[0036] Figure 2A is a schematic diagram of a training process in one embodiment;

[0037] Figure 2B A schematic diagram of a training neural network structure in one embodiment;

[0038] Figure 2C In one embodiment, Figure 2B A schematic diagram of a reconstructed neural network structure obtained by reconstructing the training neural network structure in;

[0039] Figure 3 A schematic diagram of a process of searching from a target search space to obtain target structural features in one embodiment;

[0040] Figure 4 It is an architecture diagram of a neural network structure search method in a specific embodiment;

[0041] Figure 5 is a structural block diagram of a neural network structure search device in one embodiment;

[0042] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0045] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0046] Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to the use of cameras and computers to replace human eyes to identify, follow and measure targets, and further perform graphics processing so that the computer processing becomes an image that is more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0047] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0048] The solution provided in the embodiments of the present application involves technologies such as machine learning of artificial intelligence, which is specifically described by the following embodiments:

[0049] In one embodiment, Figure 1 As shown, a neural network structure search method is provided. This embodiment uses the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server.

[0050] In this embodiment, the method includes the following steps:

[0051] Step 102, obtaining a training neural network structure.

[0052] Among them, training a neural network refers to the structural information of the neural network given in the training data set. The structural information of the neural network includes: 1) network topology, such as the number of layers, layer connection relationships, etc.; 2) layer types, such as convolutional layers, pooling layers, fully connected layers, activation layers, etc.; 3) hyperparameters within the layer, such as the number of convolution kernels, number of channels, and step size in the convolutional layer.

[0053] Specifically, the terminal can obtain the neural network structure information from the training data set, and use the obtained neural network structure information as the training neural network structure.

[0054] In one embodiment, the training data set may be a data set pre-stored locally by the terminal. In another embodiment, the terminal may obtain the training data set from other computer devices, such as a server, through a network or the like.

[0055] In a specific embodiment, the training data set may be an existing data set, such as NAS-Bench-101, NAS-Bench-201, and the like.

[0056] It is understandable that in different task scenarios, the types of training neural network structures obtained are different. For example, when it is necessary to search for a suitable neural network structure for an image recognition task in computer vision, the selected training neural network structure is the neural network structure used for image recognition. In other task scenarios, the neural network structure used is definitely different from the image recognition task. In this case, it is necessary to select a neural network structure corresponding to the other task scenarios for learning. Therefore, it is necessary to reselect the neural network structure corresponding to the task scenario as the training neural network structure.

[0057] Step 104, input the training neural network structure into the graph neural network to obtain discrete structural features corresponding to the training neural network structure.

[0058] Among them, Graph Neural Networks (GNN) refers to a neural network used to process graph data. The discrete structural features corresponding to the training neural network structure refer to the structural features of the neural network represented in discrete form.

[0059] Specifically, the neural network structure can be regarded as a directed acyclic graph (DAG), each layer of the neural network is a node in the DAG, and the connection relationship between layers is the edge in the DAG. Then, the neural network structure can be feature extracted through the graph neural network. After obtaining the training neural network structure, the terminal can input the training neural network structure into the graph neural network, extract the network topology and node content information in the training neural network DAG through the graph neural network, and encode all nodes in the DAG in sequence starting from the input node. The encoding of each node comprehensively uses the encoding information (network topology information) of all preceding nodes (i.e., the preceding nodes connected to the current node through directed edges) and the current node type (node ​​content information), and recursively encodes the subsequent nodes. Since the neural network is a directed acyclic graph, the encoding of the output node will eventually be obtained. Since the encoding of the output node contains the topology information and node information of the entire training neural network DAG, the encoding of the output node is used as the encoding representation of the entire network. At the same time, the structural feature representations of different neural networks are discretely distributed, so the obtained encoding representation of the training neural network is a discrete expression.

[0060] For example, a training neural network is a three-layer neural network, corresponding to three nodes X1, X2, and X3, and its topological structure is: X1→X2→X3, X1→X3. Use a graph neural network to encode the training neural network, use the aggregation function A to aggregate the encoded expressions of all predecessor nodes, and use the update function U to encode the current node in combination with the predecessor node representation and the current node type. Use a graph neural network to express the training neural network, and there are the following steps: 1) X1 has no predecessor node, and only needs to consider its node type T1, then X1 is expressed as X1=U(T1,0); 2) X2's predecessor node is X1, and X2's node type is T2, so X2 is expressed as X2=U(T2,A(X1)); 3) X3's predecessor nodes are X1 and X2, and X3's node type is T3, so X3 is expressed as X3=U(T3,A(X1,X2)).

[0061] In one embodiment, before inputting the training neural network structure into the graph neural network, the graph neural network needs to be initialized, including determining the model structure information and model parameters of the graph neural network. In a specific embodiment, the graph neural network can be a convolutional neural network. Since the convolutional neural network is a multi-layer neural network, each layer is composed of multiple two-dimensional planes, and each plane is composed of multiple independent neurons, it is necessary to determine which layers the graph neural network includes (for example, convolutional layers, pooling layers, etc.), the connection order relationship between layers, and which parameters each layer includes (for example, weights, bias terms, convolution step length), etc.

[0062] In a specific embodiment, the graph neural network includes two parts: 1) aggregation function A, which is used to synthesize all previous node encoding representations, and uses gated sum to fit the aggregation function A; 2) update function U, which is used to give the current node representation according to the node type and the previous node expression, and uses gated recurrent unit (GRU) to fit the update function U. The graph neural network uses the network topology information and all node information of the training neural network structure as input to obtain discrete structural features, which are specifically expressed as follows:

[0063] h v =U(T v ,A({h u :v→v}))

[0064] Among them, h v is the feature representation of node v, h u is the feature representation of node u, and node u is the predecessor node of node v. It can be seen that network A aggregates the incoming edge information of node v and the feature representation of the predecessor node; network U updates the feature representation of node v according to the node type of node v.v It is a vector representation of the node type (i.e., layer type). For example, (0,0,1,0,0) means there are five layer types in total, and this layer belongs to the third layer type, such as a convolutional layer.

[0065] Furthermore, the network parameters of the graph neural network can be initialized. In practice, the various network parameters of the graph neural network can be initialized with some different small random numbers. "Small random numbers" are used to ensure that the network will not enter a saturated state due to excessive weights, thereby causing training failure, and "different" is used to ensure that the network can learn normally.

[0066] Step 106: input the discrete structural features into the encoding network, and encode the discrete structural features into corresponding continuous structural features through the encoding network.

[0067] The encoding network refers to the machine learning module used for encoding, and encoding refers to the process of converting information from one form or format to another. The continuous structural feature refers to the form of the feature being expressed as a continuous probability distribution in the latent space. It can be understood that the latent space where the continuous structural feature is located is a continuous space.

[0068] Specifically, after the terminal inputs the obtained discrete structural features into the encoding network, the encoding network encodes the discrete structural features to obtain continuous structural features. At this time, the latent space corresponding to the encoder is a continuous space.

[0069] In a specific embodiment, the encoding network may be the encoding part of a variational auto-encoder (VAE). As a form of deep generative model, the variational auto-encoder is a generative network structure based on variational Bayes (VB) inference proposed by Kingma et al. in 2014. Unlike traditional autoencoders that describe latent space in a numerical way, it describes the observation of latent space in a probabilistic way.

[0070] Step 108: input the discrete structural features into the encoding network, and encode the discrete structural features into corresponding continuous structural features through the encoding network.

[0071] Among them, the decoding network refers to the machine learning module used for decoding. Decoding is the inverse process of encoding. Decoding restores data expressed in another form to its original form or format, and reconstructs data in the same form or format as the original data.

[0072] Specifically, the terminal can obtain a reconstructed neural network structure based on the continuous structural features and by decoding and reconstructing through a decoding network.

[0073] In one embodiment, the terminal may sample from the continuous structural features, input the sample points obtained by the sampling into a decoding network, perform decoding and reconstruction through the decoding network, and obtain a reconstructed neural network.

[0074] In another embodiment, since the continuous structural feature is a continuous probability distribution in the latent space, the terminal can also calculate the mean of the continuous probability distribution, input the calculation result into the decoding network, and obtain the reconstructed neural network structure through decoding and reconstruction of the decoding network.

[0075] Step 110, based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, train the graph neural network, the encoding network and the decoding network until the training stop condition is met, obtain the target encoding network and the target decoding network, and determine the latent space corresponding to the target encoding network as the target search space.

[0076] The reconstruction loss is related to the difference between the reconstructed neural network structure and the training neural network structure. The smaller the difference between the reconstructed neural network structure and the training neural network structure, the smaller the reconstruction loss, and the closer the obtained reconstructed neural network is to the training neural network. Training stop conditions include but are not limited to training time exceeding a preset time threshold, training times exceeding a preset number of times, reconstruction loss less than a preset threshold, etc.

[0077] In one embodiment, the reconstruction loss between the training neural network structure and the reconstructing neural network structure may be the reconstruction loss of the hidden layer of the neural network; in another embodiment, the reconstruction loss between the training neural network structure and the reconstructing neural network structure may also be the reconstruction loss of the neural network edges; in other embodiments, the reconstruction loss between the training neural network structure and the reconstructing neural network structure may also be the cumulative value of the reconstruction loss of the hidden layer of the neural network and the reconstruction loss of the neural network edges.

[0078] Specifically, the terminal determines the reconstruction loss between the training neural network structure and the reconstructed neural network structure based on the difference between the training neural network structure and the reconstructed neural network structure, and adjusts the network parameters of the graph neural network, encoding network and decoding network by backpropagation according to the reconstruction loss until the training stop condition is met, and the training is terminated to obtain the trained target encoding network, target decoding network and target graph neural network. At this time, since the network parameters of the target encoding network have been determined, the latent space corresponding to the target encoding network is also determined, and the latent space can be determined as the target search space, and the neural network structure can be searched based on the target search space.

[0079] Step 112, searching from the target search space according to the target search strategy to obtain the target structural features, and decoding the target structural features through the target decoding network to obtain the target neural network structure.

[0080] The target search strategy refers to the search strategy for searching the neural network structure. Different target search strategies can be defined according to different search requirements. For example, under the requirement of high search efficiency, the search strategy can be a random search in the target search space.

[0081] In one embodiment, the target search strategy refers to a search strategy for searching for an optimal neural network structure. Different target search strategies may be used depending on the evaluation index of the optimal neural network structure to be searched. For example, the optimal neural network may be the neural network with the lowest model complexity, so the target search strategy is to search by minimizing the model complexity. For another example, the optimal neural network may also be the neural network with the greatest generalization performance, so the target search strategy is to search by maximizing the generalization performance. The target structural feature refers to the structural feature of the optimal neural network structure searched under the target search strategy.

[0082] Specifically, the terminal can perform an iterative search in the target search space according to the target search strategy until the structural features of the optimal neural network structure are searched as the target structural features. The searched target structural features are input into the trained target decoding network. Since the target decoding network is trained based on the reconstruction loss, the neural network structure corresponding to the target structural features can be accurately reconstructed to obtain the target neural network structure. The obtained target neural network structure can be used for machine learning tasks in the current scenario. For example, if the trained neural network structure in the current scenario is a neural network structure for liveness detection, the obtained target neural network structure can be used for liveness detection.

[0083] In one embodiment, when searching in the target search space, the terminal may first sample sample points in the target search space, and then use a gradient method based on the sample points to iteratively search for neural network structural features corresponding to the target search strategy until the target structural features are obtained.

[0084] In the above neural network structure search method, by inputting the training neural network structure into the graph neural network, the graph neural network can effectively extract the structural information of the training neural network, and obtain the discrete structural features corresponding to the training neural network structure. By inputting the discrete structural features into the encoding network, the discrete structural features can be encoded into continuous structural features through the encoding network, thereby constructing a continuous search space, and decoding and reconstruction are performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure. Finally, based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the graph neural network, the encoding network and the decoding network are trained until the training stop condition is met to obtain the target encoding network and the target decoding network. Since the target encoding network and the target decoding network are obtained based on the reconstruction loss training, the target encoding network can well learn the structural features of the training neural network, and the target decoding network can accurately decode and reconstruct the structural features from the target search space to obtain the target neural network structure. Therefore, the latent space of the target encoding network can be determined as the target search space. Since the target search space is a continuous space, searching in the continuous space is easier to converge than searching in the discrete space, thereby improving the efficiency of the neural network structure search.

[0085] In one embodiment, before training the graph neural network, encoding network and decoding network based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the neural network structure search method also includes: obtaining at least one evaluation indicator label value corresponding to the training neural network structure; making predictions based on the continuous structural features corresponding to the training neural network structure and each evaluation indicator prediction network, and obtaining evaluation indicator training values ​​corresponding to each evaluation indicator prediction network; determining the evaluation indicator loss corresponding to each evaluation indicator label value based on each evaluation indicator label value and the corresponding evaluation indicator training value; and each evaluation indicator loss is used to train the corresponding target evaluation indicator prediction network.

[0086] Among them, the evaluation index refers to the index used to evaluate the quality of the neural network structure. The evaluation index includes but is not limited to generalization performance, model complexity, resource constraints, etc. Among them, generalization performance refers to the performance of the neural network on unknown data. Generalization performance can be, for example, model accuracy. The model accuracy is used to describe the accuracy of the searched neural network structure under the corresponding machine learning task. For example, the model accuracy of the neural network structure used for image classification, and its corresponding model accuracy describes the classification accuracy of the neural network structure when used for image classification tasks. Model complexity is used to characterize the complexity of the model structure, for example, it can be the number of parameters and training time.

[0087] The evaluation index label value refers to the corresponding value of the evaluation index of the training neural network in the training data set, which is used as the training label of the training neural network during the training process, that is, the expected output value of the training neural network. The evaluation index prediction network refers to a network used to predict the evaluation index value of an unknown neural network structure. In one embodiment, the evaluation index label value includes at least one of model complexity and generalization performance.

[0088] It is understandable that, depending on the search strategy for the neural network structure, the evaluation index can be one or more, and different evaluation indexes correspond to different evaluation index prediction networks, so correspondingly, the evaluation index prediction network can also be one or more. For example, when the evaluation index is model complexity, the corresponding evaluation index prediction network is the model complexity prediction network, which is used to predict the model complexity of the neural network structure; for another example, when the evaluation index is generalization performance, the corresponding evaluation index prediction network is the generalization performance prediction network, which is used to predict the generalization performance of the neural network structure.

[0089] Specifically, in order to enable the evaluation index prediction network to accurately predict the evaluation index value of the searched neural network structure, the terminal needs to train the evaluation index prediction network. In this application, since the search is performed in the latent space corresponding to the encoding network, the final evaluation index prediction network needs to predict the evaluation index corresponding to the neural network structure characteristics in the latent space. Then, during the training process, the terminal can be trained based on the latent space. The terminal predicts the continuous structural features obtained by encoding the training neural network into the latent space and each evaluation index prediction network, and obtains the evaluation index value corresponding to each evaluation index prediction network. The evaluation index value is the real value output during the training process, so it is called the evaluation index training value. The purpose of training the evaluation indicator prediction network is to make the real value of the evaluation indicator prediction network output fit the expected output value corresponding to the evaluation indicator prediction network. During the fitting process, the parameters of the evaluation indicator prediction network can be adjusted based on the evaluation indicator loss determined according to the difference between the expected output value corresponding to the evaluation indicator prediction network (i.e., the evaluation indicator label value) and the real output value (i.e., the evaluation indicator training value). In other words, each evaluation indicator loss determined according to the difference between the expected output value corresponding to each evaluation indicator prediction network (i.e., the evaluation indicator label value) and the real output value (i.e., the evaluation indicator training value) is used to train the corresponding target evaluation indicator prediction network. The target evaluation indicator prediction network here refers to the evaluation indicator prediction network obtained when the training is completed.

[0090] In one embodiment, the terminal performs prediction based on the continuous structural features encoded into the latent space by the training neural network and each evaluation index prediction network to obtain the evaluation index value corresponding to each evaluation index prediction network. Specifically, the terminal samples the continuous structural features, inputs the sampling results into each evaluation index prediction network respectively, and predicts the evaluation index value corresponding to each sampling result through each evaluation index prediction network.

[0091] In another embodiment, the terminal performs prediction based on the continuous structural features encoded into the latent space by the training neural network and each evaluation index prediction network to obtain the evaluation index value corresponding to each evaluation index prediction network. Specifically, the terminal calculates the mean of the continuous structural features, inputs the calculation results into each evaluation index prediction network respectively, and predicts the corresponding evaluation index value of each calculation result through each evaluation index prediction network.

[0092] In a specific embodiment, in order to emphasize the positive correlation between the difference between the expected output value and the actual output value and the evaluation index loss, the square loss between the expected output value and the actual output value may be used as the evaluation index loss.

[0093] In one embodiment, Figure 2 As shown, a neural network structure search method is provided, comprising the following steps 202-216:

[0094] Step 202, obtaining a training neural network structure.

[0095] Step 204: input the training neural network structure into the graph neural network to obtain discrete structural features corresponding to the training neural network structure.

[0096] Step 206: input the discrete structural features into the encoding network, and encode the discrete structural features into corresponding continuous structural features through the encoding network.

[0097] Step 208, decoding is performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure.

[0098] Step 210, obtaining at least one evaluation index label value corresponding to the training neural network structure.

[0099] Step 212, prediction is performed based on the continuous structural features corresponding to the training neural network structure and each evaluation index prediction network, and evaluation index training values ​​corresponding to each evaluation index prediction network are obtained respectively.

[0100] Step 214, determining the evaluation indicator loss corresponding to each evaluation indicator label value according to each evaluation indicator label value and the corresponding evaluation indicator training value.

[0101] Regarding steps 202 to 214, reference may be made to the description in the above embodiment, and this application will not elaborate on them here.

[0102] Step 216, based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and the losses of various evaluation indicators, jointly train the graph neural network, various evaluation indicator prediction networks, encoding network and decoding network until the training stop conditions are met, and obtain the target evaluation indicator prediction network, target encoding network and target decoding network.

[0103] It can be understood that in order to achieve the purpose of searching for the optimal neural network structure in the continuous search space, it is necessary to train the graph neural network, the encoding network, the decoding network and at least one evaluation index prediction network. Then, a supervised training loss function for jointly training these networks can be constructed. According to the supervised training loss function, the graph neural network, the encoding network, the decoding network and at least one evaluation index prediction network can be jointly trained, thereby improving the training efficiency.

[0104] Specifically, the terminal can construct a supervised training loss function based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and the losses of various evaluation indicators, and then jointly train the graph neural network, various evaluation indicator prediction networks, encoding network and decoding network based on the supervised training loss function. The training process is also the process of adjusting the network parameters of the graph neural network, various evaluation indicator prediction networks, encoding network and decoding network, until the training stop conditions are met. At this time, the network parameters of the graph neural network, various evaluation indicator prediction networks, encoding network and decoding network have been determined, so the graph neural network, various evaluation indicator prediction networks, encoding network and decoding network at this time can be determined as the target evaluation indicator prediction network, target encoding network and target decoding network.

[0105] like Figure 2A As shown in FIG. 1 , a schematic diagram of a training process in a specific embodiment is shown. In this embodiment, the graph neural network and the encoding network are collectively referred to as the encoding network. Figure 2A The terminal first inputs the training neural network structure into the encoding network, encodes the training neural network structure into continuous structural features in the latent space through the encoding network, and then samples based on the continuous structural features, and inputs the sampling results into the evaluation index prediction network, so that the evaluation index prediction network can fit the evaluation index of the training neural network structure, and the decoding network restores the continuous structural features to obtain a specific neural network structure, that is, reconstructs the neural network structure.

[0106] refer to Figure 2B, which is a schematic diagram of a training neural network structure in an embodiment. In this embodiment, the training neural network structure specifically includes an input layer, five convolutional layers, a 3x3 maximum pooling layer, and an output layer, wherein the five convolutional layers are, from top to bottom, a 5x5 depthwise separable convolution layer, a 3x3 depthwise separable convolution layer, a 3x3 ordinary convolution layer, a 3x3 ordinary convolution layer, and a 3x3 depthwise separable convolution layer.

[0107] refer to Figure 2C As shown, in one embodiment Figure 2B Schematic diagram of the reconstructed neural network structure obtained by reconstructing the training neural network structure in . The reconstructed neural network structure specifically includes an input layer, four convolutional layers, a 3x3 average pooling layer, a 3x3 maximum pooling layer and an output layer, wherein the four convolutional layers are, from top to bottom, a 5x5 depthwise separable convolutional layer, a 5x5 ordinary convolutional layer, a 3x3 depthwise separable convolutional layer and a 3x3 ordinary convolutional layer.

[0108] Understandably, Figure 2B , Figure 2C The neural network structure shown is only an example. In a specific application, a training neural network structure with a different structure is selected according to the requirements of the application scenario, and a corresponding reconstructed neural network structure is reconstructed.

[0109] It is understandable that in other embodiments, the terminal may first train the graph neural network, encoding network and decoding network based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and after training to obtain the target graph neural network, the target encoding network and the target decoding network, fix the parameters of the three networks, and then train the evaluation index prediction network based on the training sample set. Specifically, the terminal may input the training neural network structure in the training sample set into the target graph neural network to obtain the corresponding discrete structural features, and then input the obtained discrete structural features into the target encoding network, and encode the corresponding continuous structural features through the target encoding network, and obtain the true output value according to the obtained continuous structural features and the evaluation index prediction network, and determine the evaluation index loss based on the difference between the true output value and the expected output value, and adjust the model parameters of the evaluation index prediction network based on the loss by back propagation, until the training stop condition is met to terminate the training and obtain the target evaluation index prediction network.

[0110] In the above embodiment, by training the evaluation index prediction network, in the model search stage, the searched neural network structure can be predicted based on the evaluation index prediction network, so that the neural network structure with the most optimized evaluation index can be searched.

[0111] In one embodiment, Figure 3 As shown, according to the target search strategy, the target structural features are obtained by searching from the target search space, including:

[0112] Step 302: Sampling is performed in the target search space to obtain sample points.

[0113] In one embodiment, the terminal may perform random sampling in the target search space to obtain sample points.

[0114] In another embodiment, the encoding network is the encoding network of a variational autoencoder, and the decoding network is the decoding network of a variational autoencoder. When the terminal performs sampling, it can sample from the prior distribution of the variational autoencoder to obtain sample points. Due to the special structure of the variational autoencoder, the prior distribution of the variational autoencoder is a standard normal distribution. The terminal samples from the standard normal distribution, which can reduce the search range and further improve the search efficiency.

[0115] Step 304, searching for neural network structural features using the sample point as a starting point.

[0116] Step 306: input the searched neural network structure features into each trained target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network.

[0117] Step 308, obtaining a comprehensive evaluation value according to the predicted values ​​of each evaluation index and the target search strategy.

[0118] Specifically, the terminal starts searching for the neural network structure from the sample point, inputs the searched neural network structure features into each trained target evaluation index prediction network, predicts the searched neural network structure features through each target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network, and then the terminal can calculate each evaluation index prediction value according to the target search strategy to obtain a comprehensive evaluation value, which is used to comprehensively evaluate the advantages and disadvantages of the searched neural network structure.

[0119] It is understandable that when there is only one evaluation index, there is also only one corresponding evaluation index prediction network, and the final evaluation index prediction value of the neural network structure is also only one. At this time, the evaluation index prediction value is also the comprehensive evaluation value.

[0120] In one embodiment, the evaluation index label value includes model complexity and generalization performance, wherein the trained target evaluation index prediction network corresponding to the model complexity is the target model complexity prediction network, and the trained evaluation index prediction network corresponding to the generalization performance is the target generalization performance prediction network. Then, when the terminal searches in the target search space, the searched neural network structure features can be input into the target model complexity prediction network and the target generalization performance prediction network respectively, and the model complexity of the searched neural network structure features is predicted by the target model complexity prediction network to obtain the corresponding model complexity prediction value, and the generalization performance of the searched neural network structure features is predicted by the target generalization performance prediction network to obtain the generalization performance prediction value.

[0121] Step 310, iteratively searching for the neural network structural features along the direction of optimizing the comprehensive evaluation value until the target structural features are obtained.

[0122] Specifically, the terminal can start from the sample starting point and use the gradient method to iteratively search the neural network structure features in the direction of the optimized comprehensive evaluation value until the target structure features are obtained. The gradient method can be gradient ascent or gradient descent, which is determined according to the target search strategy. The optimization can be to minimize the comprehensive evaluation value or to maximize the comprehensive evaluation value, which is specifically determined according to the target search strategy. It is understandable that in other embodiments, other methods can also be used to optimize the comprehensive average value, and this application is not limited here.

[0123] In one embodiment, the target search strategy is to maximize the generalization performance while reducing the model complexity. Then, after obtaining the generalization performance prediction value and the model complexity prediction value, the terminal can subtract the model complexity prediction value from the generalization performance prediction value to obtain a comprehensive evaluation value, and then use the gradient ascent method to iteratively search for the neural network structure feature in the direction of maximizing the comprehensive evaluation value until the target structure feature is obtained. The specific formula is as follows:

[0124]

[0125] Among them, f(s) is the comprehensive evaluation value, f perf (s) is the predicted value of generalization performance, f comp (s) is the predicted value of model complexity. In this embodiment, by optimizing the comprehensive evaluation value, a neural network structure with excellent generalization performance and low model complexity can be searched.

[0126] In another embodiment, the target search strategy is to minimize the model complexity while improving the generalization performance. Then, after obtaining the generalization performance prediction value and the model complexity prediction value, the terminal can subtract the generalization performance prediction value from the model complexity prediction value to obtain a comprehensive evaluation value, and then use the gradient descent method to iteratively search for the neural network structure features in the direction of minimizing the comprehensive evaluation value until the target structure features are obtained. In this embodiment, by optimizing the comprehensive evaluation value, a neural network structure with low model complexity and excellent generalization performance can be searched.

[0127] In one embodiment, based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the graph neural network, the encoding network and the decoding network are trained, including: determining the reconstruction loss of the hidden layer according to the hidden layer structure of the training neural network structure and the hidden layer structure of the reconstructed neural network structure; based on the reconstruction loss of the hidden layer, training the graph neural network, the encoding network and the decoding network.

[0128] Specifically, the terminal determines the reconstruction loss of the hidden layer according to the following formula:

[0129]

[0130] Among them, L n is the reconstruction loss of the hidden layer, T i Represents the layer type of the i-th layer of the training neural network structure, for example, it can be (0,0,1,0,0), T i ' represents the layer type of the i-th layer of the reconstructed neural network structure, and CE represents the use of cross-entropy to measure the difference between the layer type of the i-th layer of the trained neural network structure and the layer type of the i-th layer of the reconstructed neural network structure.

[0131] Specifically, after calculating the reconstruction loss of the hidden layer, the terminal can adjust the network parameters of the graph neural network, the encoding network and the decoding network by backpropagation based on the reconstruction loss of the hidden layer.

[0132] In one embodiment, before training the graph neural network, encoding network and decoding network based on the reconstruction loss of the hidden layer, the method further includes: determining the reconstruction loss of the edges according to the edges of the training neural network structure and the edges of the reconstructed neural network structure; training the graph neural network, encoding network and decoding network based on the reconstruction loss of the hidden layer, including: training the graph neural network, encoding network and decoding network based on the reconstruction loss of the hidden layer and the reconstruction loss of the edges.

[0133] Specifically, the terminal calculates the reconstruction loss of the edge according to the following formula:

[0134]

[0135] Among them, Le is the edge reconstruction loss, E i Represents all edge types of the i-th node of the training neural network structure, for example, it can be (1,1,0,0,0,1), E i ' represents the edge type of the i-th node of the reconstructed neural network structure, and CE represents the use of cross-entropy to measure the difference between the edge type of the i-th node of the trained neural network structure and the edge type of the i-th node of the reconstructed neural network structure.

[0136] Specifically, after calculating the reconstruction loss of the hidden layer and the reconstruction loss of the edge, the terminal superimposes the two losses to obtain the sum of losses, and adjusts the network parameters of the graph neural network, encoding network and decoding network based on the loss and back propagation.

[0137] It is understandable that in some other embodiments, after calculating the reconstruction loss of the edge, the terminal may adjust the network parameters of the graph neural network, the encoding network, and the decoding network based on the reconstruction loss of the edge by back-propagation.

[0138] In one embodiment, the encoding network is an encoding network of a variational autoencoder, and the decoding network is a decoding network of a variational autoencoder: based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the graph neural network, the encoding network, and the decoding network are trained, including: based on relative entropy and the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the graph neural network, the encoding network, and the decoding network are trained.

[0139] Among them, relative entropy, also known as Kullback-Leibler divergence (KL divergence) or information divergence, is an asymmetric measure of the difference between two probability distributions.

[0140] Specifically, since the variational autoencoder encodes a continuous probability distribution, and in practice it is expected that the continuous probability distribution obtained by the variational autoencoder is a standard normal distribution, because under the standard normal distribution, the mean is 0 and the variance is 1, so the latent space obtained has better continuity. In order to make the distribution of the variational autoencoder fit the standard normal distribution, during training, a loss function can be constructed based on the reconstruction loss and the regularization term. The reconstruction term tends to make the encoding and decoding scheme as high-performance as possible, and the regularization term normalizes the organization of the latent space by making the distribution returned by the encoding network close to the standard normal distribution. Specifically in this embodiment, the reconstruction loss is the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and the regularization term is the KL divergence between the returned distribution and the standard Gaussian.

[0141] In a specific embodiment, when performing training, the terminal can construct a loss function based on the KL divergence and the reconstruction loss, perform training based on the loss function, adjust the network parameters of the graph neural network, the encoding network, and the decoding network, and use the gradient descent method to optimize the loss function during the training process, and the optimization goal is to minimize the loss function.

[0142] In the above embodiment, by using variational autoencoders as the encoding network and decoding network, during the training process, training is performed based on relative entropy and reconstruction loss, which not only makes the obtained encoding network and decoding network have high performance, but also makes the obtained latent space have good continuity, thereby making it easier to converge during the search process and further improving the search efficiency.

[0143] In a specific embodiment, the network architecture used in the neural network structure search method provided in the embodiment of the present application is as follows: Figure 4 Reference Figure 4The network architecture includes an encoder, a predictor, a decoder, and a latent space: wherein the encoder uses a graph neural network (GNN) to represent the structure of the training neural network, represents the structural information of the training neural network in discrete form, and uses the encoding part (i.e., the encoding network) of the variational autoencoder (VAE) to map the discrete expression of the neural network structure (i.e., discrete structural features) to a suitable latent space, and needs to minimize the KL divergence during the training process to learn the decoder (i.e., the decoding network); the predictor includes two regression models, a generalization performance predictor (i.e., a generalization performance prediction network) and a model complexity predictor (i.e., a model complexity prediction network), which are respectively used to fit the potential relationship between the continuous expression of the neural network structure (i.e., continuous structural features) and the corresponding generalization performance and the corresponding model complexity. In the model search stage, the continuous expression of the optimal neural network is obtained by using the optimization algorithm on the predictor; the decoder uses the decoder part of the variational autoencoder (VAE) to decode the structural features in the continuous space into the neural network structure. In the training process, it is necessary to minimize the reconstruction error between the decoded neural network structure and the training neural network, and is used to decode the searched neural network structure features in the model search stage. This network architecture can be used to construct a continuous search space for neural network structure search, and ultimately obtain a neural network structure with excellent generalization performance and low model complexity.

[0144] In a specific embodiment, a neural network structure search method is provided, comprising the following steps:

[0145] 1. The training process includes the following steps

[0146] 1. Obtain the training neural network structure and the model complexity label value and generalization performance label value corresponding to the training neural network structure;

[0147] 2. Input the training neural network structure into the graph neural network to obtain the discrete structural features corresponding to the training neural network structure;

[0148] 3. Input the discrete structure features into the encoding part of the variational autoencoder, and encode the discrete structure features into a continuous expression in the continuous differentiable feature space through the encoding part;

[0149] 4. Sampling is performed from the continuous expression, and the sample points obtained by sampling are input into the decoding part of the variational autoencoder to obtain the reconstructed neural network structure;

[0150] 5. Determine the reconstruction loss of the hidden layer according to the hidden layer structure of the training neural network structure and the hidden layer structure of the reconstructed neural network structure;

[0151] 6. Determine the reconstruction loss of the edge according to the edge of the training neural network structure and the edge of the reconstructed neural network structure;

[0152] 7. Input the sample point inputs obtained by sampling into the generalization performance predictor and the model complexity predictor respectively to obtain the generalization performance training value and the model complexity training value;

[0153] 8. Determine the model complexity loss based on the square of the difference between the model complexity label value and the model complexity training value;

[0154] 9. Determine the generalization performance loss based on the square of the difference between the generalization performance label value and the generalization performance training value;

[0155] 10. According to the reconstruction loss of the hidden layer, the reconstruction loss of the edge, the model complexity loss, the generalization performance loss and the relative entropy (KL divergence), a loss function for jointly training the graph neural network, the encoding part of the variational autoencoder, the decoding part of the variational autoencoder, the generalization performance predictor and the model complexity predictor is constructed. The gradient descent method is used to train with the goal of minimizing the loss function, and the network parameters of the graph neural network, the encoding part of the variational autoencoder, the decoding part of the variational autoencoder, the generalization performance predictor and the model complexity predictor are adjusted until the training process converges, and the target graph neural network, the target variational autoencoder, the target generalization performance prediction network and the target model complexity prediction network are obtained. During the training process, the generalization performance predictor and the model complexity predictor fit the continuous expression of the training neural network structure in the continuously differentiable feature space.

[0156] Among them, the loss function is as follows:

[0157]

[0158] Among them, φ and θ are the parameters of the encoding network and decoding network respectively, y is the generalization performance training value, and z is the model complexity training value.

[0159] 2. The search process specifically includes the following steps:

[0160] 11. Determine the continuous differentiable feature space corresponding to the target variational autoencoder as the target search space, and perform sampling in the target search space based on the prior distribution of the variational autoencoder to obtain sample points;

[0161] 12. Using the sample point as the starting point, search for the neural network structure features. During the search, input the searched neural network structure features into the trained target generalization performance predictor and target model complexity predictor to obtain the generalization performance prediction value and the model complexity prediction value. Optimize the optimization target according to the obtained generalization performance prediction value and the model complexity prediction value. The optimization target is to maximize the generalization performance while reducing the model complexity. The optimization method can use the gradient ascent method. The specific optimization targets are as follows;

[0162]

[0163] Among them, f perf (s) is the predicted value of generalization performance, f comp (s) is the predicted value of model complexity.

[0164] 15. When the optimization target reaches the minimum value, the searched neural network structure is determined as the target structure feature, the target structure feature is input into the target decoding network, and the target neural network structure is obtained by decoding the target structure feature through the target decoding network.

[0165] In this embodiment, by inputting the training neural network structure into the graph neural network to construct discrete structural features, the structural information in the training neural network can be effectively extracted; by using the variational autoencoder, the neural network structure can be mapped to a continuous search space in this embodiment; by using the gradient method to optimize each graph neural network, the encoding part of the variational autoencoder, the decoding part of the variational autoencoder, the generalization performance predictor and the model complexity predictor in the continuous differentiable feature space, the neural network structure search method provided in this embodiment has a high training efficiency in the training stage; by simultaneously fitting the generalization performance and model complexity of the neural network, the neural network structure search method provided in this embodiment has both excellent generalization performance and low model complexity in the search stage.

[0166] The overall effect of the neural network structure search method provided in the embodiment of the present application in practical application is shown in Table 1 below:

[0167] Table 1

[0168]

[0169] It can be seen from the above table that when training samples are selected from the same data set, the neural network structure search method provided in the embodiment of the present application achieves excellent results in terms of training time, generalization performance, and complexity of the searched model compared with other methods.

[0170] The present application also provides an application scenario, which applies the above-mentioned neural network structure search method. Specifically, the application of the neural network structure search method in this application scenario is as follows:

[0171] In this application scenario, it is necessary to select a suitable deep learning model network architecture for the image classification task in computer vision. For example, the image classification task can be to classify product images, classify the gender of images containing human faces, and so on. In this application scenario, the neural network structure for image classification is first selected from the NAS-Bench-101 dataset as the training neural network structure. The parameter quantity and model accuracy corresponding to the selected neural network structure are used as evaluation indicators to predict the training label of the network. When searching for the neural network structure, the neural network structure for image classification is first input into the graph neural network to obtain the corresponding discrete structural features. Then, the discrete structural features are encoded into the continuous space through the encoding part of the variational autoencoder to obtain a continuous probability distribution. Samples are taken from the continuous probability distribution, and the sample points are input into the decoding part of the variational autoencoder. The decoding part decodes and reconstructs to obtain a reconstructed neural network structure, determines the reconstruction loss of the hidden layer according to the hidden layer structure of the reconstructed neural network structure and the hidden layer structure of the selected training neural network structure, determines the reconstruction loss of the edge according to the edge of the training neural network structure and the edge of the reconstructed neural network structure, and inputs the sample point into the generalization performance predictor and the model complexity predictor at the same time, obtains the accuracy prediction value through the generalization performance predictor, obtains the model parameter prediction value through the model complexity predictor, and then superimposes the KL divergence, the reconstruction loss of the hidden layer, the reconstruction loss of the edge, the average of the accuracy prediction value and the model accuracy in the training label. The comprehensive loss is obtained by combining the square loss of the predicted value of the model parameter quantity and the square loss of the parameter quantity in the training label. Based on the comprehensive loss, the graph artifact network, the encoding part of the variational autoencoder, the decoding part of the variational autoencoder, the generalization performance predictor and the model complexity predictor are jointly trained. During the training process, the gradient descent method is used to adjust the parameters of the above networks with the goal of minimizing the comprehensive loss until the training process converges. The trained graph neural network, variational autoencoder, generalization performance predictor and model complexity predictor are obtained. Next, the search can be carried out based on the trained networks. Specifically, first, the prior distribution of the variational autoencoder is Sampling is performed in the cloth, starting from the sample points obtained by sampling, the gradient ascent method is used to iteratively search the neural network structure, and the searched neural network structure features are input into the trained generalization performance predictor and model complexity predictor to obtain the parameter quantity prediction value and model accuracy prediction value of the searched neural network structure features, until the value obtained by subtracting the parameter quantity prediction value from the model accuracy prediction value is the largest, the searched neural network structure features are determined as the target structure features, and the target structure features are decoded and reconstructed through the decoding network of the trained variational autoencoder to obtain the neural network structure, which can be used for image classification tasks.

[0172] Take the classification of product images as an example. The classification categories can be clothing, fresh food, food, toiletries, etc. When performing image classification tasks, you first need to obtain labeled product image data, use these labeled product images as training sample images, and use the corresponding category labels as training labels to conduct supervised training on the neural network. After training, you can classify product images of unknown classification. In practical applications, by classifying product images, you can recommend products of the same type to users.

[0173] The present application also provides an application scenario, which applies the above-mentioned neural network structure search method. Specifically, the application of the neural network structure search method in this application scenario is as follows:

[0174] In this application scenario, it is necessary to select a suitable deep learning model network architecture for the text classification task in natural language processing. Text classification tasks can be, for example, news classification, sentiment classification, public opinion classification, spam classification, etc. In this application scenario, the neural network structure for text classification is first selected from the ENAS-CIFAR-10 dataset as the training neural network structure. The training time and model accuracy corresponding to the selected neural network structure are used as evaluation indicators to predict the training label of the network. When searching for the neural network structure, the neural network structure for text classification is first input into the graph neural network to obtain the corresponding discrete structural features. Then, the discrete structural features are encoded into the continuous space through the encoding part of the variational autoencoder to obtain a continuous probability distribution. Samples are taken from the continuous probability distribution, and the sample points are input into the decoding part of the variational autoencoder. The decoding part decodes and reconstructs to obtain a reconstructed neural network structure, determines the reconstruction loss of the hidden layer according to the hidden layer structure of the reconstructed neural network structure and the hidden layer structure of the selected training neural network structure, determines the reconstruction loss of the edge according to the edge of the training neural network structure and the edge of the reconstructed neural network structure, and inputs the sample point into the generalization performance predictor and the model complexity predictor at the same time, obtains the accuracy prediction value through the generalization performance predictor, obtains the model training time prediction value through the model complexity predictor, and then superimposes the KL divergence, the reconstruction loss of the hidden layer, the reconstruction loss of the edge, the square of the accuracy prediction value and the model accuracy in the training label The comprehensive loss is obtained by combining the square loss between the loss, the model training time prediction value and the training time in the training label. Based on the comprehensive loss, the graph artifact network, the encoding part of the variational autoencoder, the decoding part of the variational autoencoder, the generalization performance predictor and the model complexity predictor are jointly trained. During the training process, the gradient descent method is used to adjust the parameters of the above networks with the goal of minimizing the comprehensive loss until the training process converges. The trained graph neural network, variational autoencoder, generalization performance predictor and model complexity predictor are obtained. Next, the search can be performed based on the trained networks. Specifically, first, in the prior distribution of the variational autoencoder Sampling is performed in the method, and starting from the sample points obtained by sampling, the gradient ascent method is used to iteratively search the neural network structure. The searched neural network structure features are input into the trained generalization performance predictor and model complexity predictor to obtain the training time prediction value and model accuracy prediction value of the searched neural network structure features. When the value obtained by subtracting the training time prediction value from the model accuracy prediction value is the largest, the searched neural network structure features are determined as the target structure features. The target structure features are decoded and reconstructed through the decoding network of the trained variational autoencoder to obtain the neural network structure, which can be used for text classification tasks.

[0175] Taking news classification as an example, the classification categories can be international news, domestic news, and local news. When performing text classification tasks, first obtain labeled news texts, use these labeled news texts as training samples, and use the corresponding category labels as training labels to train the neural network. After training, news texts of unknown classification can be classified. In practical applications, by classifying news, users can be pushed news of the same type based on the text type of frequently browsed news.

[0176] It should be understood that although Figure 1 , Figure 3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 , Figure 3 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0177] In one embodiment, Figure 5 As shown, a neural network structure search device 500 is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: a training data acquisition module 502, a discrete encoding module 504, a continuous encoding module 506, a decoding template 508, a training module 510 and a search module 512, wherein:

[0178] A training data acquisition module 502 is used to acquire a training neural network structure;

[0179] A discrete encoding module 504 is used to input the training neural network structure into the graph neural network to obtain discrete structural features corresponding to the training neural network structure;

[0180] A continuous encoding module 506 is used to input the discrete structural features into the encoding network, and encode the discrete structural features into corresponding continuous structural features through the encoding network;

[0181] A decoding template 508, used to perform decoding according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure;

[0182] A training module 510 is used to train the graph neural network, the encoding network and the decoding network based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure until a training stop condition is met, thereby obtaining a target encoding network and a target decoding network, and determining a latent space corresponding to the target encoding network as a target search space;

[0183] The search module 512 is used to search from the target search space according to the target search strategy to obtain the target structure features, and decode the target structure features through the target decoding network to obtain the target neural network structure.

[0184] The above-mentioned neural network structure search device, by inputting the training neural network structure into the graph neural network, can effectively extract the structural information of the training neural network through the graph neural network, and obtain the discrete structural features corresponding to the training neural network structure. By inputting the discrete structural features into the encoding network, the discrete structural features can be encoded into continuous structural features through the encoding network, thereby constructing a continuous search space, and decoding and reconstruction are performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure. Finally, based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, the graph neural network, the encoding network and the decoding network are trained until the training stop condition is met to obtain the target encoding network and the target decoding network. Since the target encoding network and the target decoding network are obtained based on the reconstruction loss training, the target encoding network can well learn the structural features of the training neural network, and the target decoding network can accurately decode and reconstruct the structural features in the target search space to obtain the target neural network structure. Therefore, the latent space of the target encoding network can be determined as the target search space. Since the target search space is a continuous space, searching in the continuous space is easier to converge than searching in the discrete space, thereby improving the efficiency of the neural network structure search.

[0185] In one embodiment, the above-mentioned device also includes: an evaluation index input module, which is used to obtain at least one evaluation index label value corresponding to the training neural network structure; predicting according to the continuous structural characteristics corresponding to the training neural network structure and each evaluation index prediction network, and obtaining the evaluation index training values ​​corresponding to each evaluation index prediction network; determining the evaluation index loss corresponding to each evaluation index label value according to each evaluation index label value and the corresponding evaluation index training value; each evaluation index loss is used to train the corresponding target evaluation index prediction network.

[0186] In one embodiment, the training module 510 is also used to: based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and the losses of various evaluation indicators, jointly train the graph neural network, various evaluation indicator prediction networks, encoding networks and decoding networks until the training stop conditions are met, thereby obtaining the target evaluation indicator prediction network, target encoding network and target decoding network.

[0187] In one embodiment, the search module 512 is also used to sample in the target search space to obtain sample points; search for neural network structural features with the sample points as the starting point; input the searched neural network structural features into each trained target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network; obtain a comprehensive evaluation value based on each evaluation index prediction value and the target search strategy; iteratively search the neural network structural features along the direction of optimizing the comprehensive evaluation value until the target structural features are obtained.

[0188] In one embodiment, at least one evaluation indicator label value includes at least one of model complexity and generalization performance.

[0189] In one embodiment, the evaluation index label value includes model complexity and generalization performance; the search module 512 is also used to input the searched neural network structure features into the target model complexity prediction network and the target generalization performance prediction network respectively to obtain the generalization performance prediction value and the model complexity prediction value; subtract the model complexity prediction value from the generalization performance prediction value to obtain a comprehensive evaluation value; iteratively search the neural network structure features in the direction of maximizing the comprehensive evaluation value until the target structure features are obtained.

[0190] In one embodiment, the evaluation index label value includes model complexity and generalization performance; the search module 512 is also used to input the searched neural network structure features into the target model complexity prediction network and the target generalization performance prediction network respectively to obtain the generalization performance prediction value and the model complexity prediction value; subtract the generalization performance prediction value from the model complexity prediction value to obtain a comprehensive evaluation value; iteratively search the neural network structure features in the direction of minimizing the comprehensive evaluation value until the target structure features are obtained.

[0191] In one embodiment, the training module 510 is also used to determine the reconstruction loss of the hidden layer according to the hidden layer structure of the training neural network structure and the hidden layer structure of the reconstructed neural network structure; based on the reconstruction loss of the hidden layer, train the graph neural network, encoding network and decoding network.

[0192] In one embodiment, the training module 510 is also used to determine the reconstruction loss of the edges based on the edges of the training neural network structure and the edges of the reconstructed neural network structure; based on the reconstruction loss of the hidden layer and the reconstruction loss of the edges, train the graph neural network, the encoding network and the decoding network.

[0193] In one embodiment, the search module 512 is further used to sample in the prior distribution of the variational autoencoder to obtain sample points.

[0194] In one embodiment, the encoding network is an encoding network of a variational autoencoder, and the decoding network is a decoding network of a variational autoencoder: the training module 510 is also used to train the graph neural network, the encoding network and the decoding network based on relative entropy and the reconstruction loss between the training neural network structure and the reconstructed neural network structure.

[0195] In one embodiment, the decoding template 508 is also used to perform sampling according to the continuous structural features to obtain the sampled structural features; and input the sampled structural features into the decoding network to obtain the reconstructed neural network structure.

[0196] For the specific definition of the neural network structure search device, please refer to the definition of the neural network structure search method above, which will not be repeated here. Each module in the above-mentioned neural network structure search device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0197] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a neural network structure search method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0198] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0199] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0200] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0201] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.

[0202] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0203] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0204] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A neural network structure search method, characterized in that: The method comprises: Obtain a neural network structure for text classification or image classification as a training neural network structure; Inputting the training neural network structure into a graph neural network to obtain discrete structural features corresponding to the training neural network structure; Inputting the discrete structural features into an encoding network, and encoding the discrete structural features into corresponding continuous structural features through the encoding network; Decoding is performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure; Based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, the graph neural network, the encoding network and the decoding network are trained until a training stop condition is met, a target encoding network and a target decoding network are obtained, and a latent space corresponding to the target encoding network is determined as a target search space; According to the target search strategy, a target structural feature is searched from the target search space to obtain the target structural feature, and the target neural network structure is obtained by decoding the target structural feature through the target decoding network. The target neural network structure is used for text classification tasks or image classification tasks.

2. The method according to claim 1, characterized in that Before training the graph neural network, the encoding network, and the decoding network based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, the method further includes: Obtaining at least one evaluation index label value corresponding to the training neural network structure; Predictions are made based on the continuous structural features corresponding to the training neural network structure and each evaluation index prediction network, to obtain evaluation index training values ​​corresponding to each evaluation index prediction network; The evaluation indicator loss corresponding to each evaluation indicator label value is determined according to each evaluation indicator label value and each corresponding evaluation indicator training value; each evaluation indicator loss is used to train a corresponding target evaluation indicator prediction network.

3. The method according to claim 2, characterized in that The step of training the graph neural network, the encoding network, and the decoding network based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure until a training stop condition is met to obtain a target encoding network and a target decoding network comprises: Based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and the losses of each evaluation index, the graph neural network, each evaluation index prediction network, the encoding network and the decoding network are jointly trained until the training stop condition is met, thereby obtaining the target evaluation index prediction network, the target encoding network and the target decoding network.

4. The method according to claim 2, characterized in that: The step of searching the target search space according to the target search strategy to obtain the target structure feature includes: Sampling in the target search space to obtain sample points; Searching for neural network structural features using the sample point as a starting point; Input the searched neural network structure features into each trained target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network; Obtaining a comprehensive evaluation value according to the predicted values ​​of each evaluation index and the target search strategy; The neural network structural features are iteratively searched in the direction of optimizing the comprehensive evaluation value until the target structural features are obtained.

5. The method according to claim 2, characterized in that: The at least one evaluation index label value includes at least one of model complexity and generalization performance.

6. The method according to claim 4, characterized in that The evaluation index label values ​​include model complexity and generalization performance; The searched neural network structure features are input into each trained target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network, including: The searched neural network structure features are respectively input into the target model complexity prediction network and the target generalization performance prediction network to obtain the generalization performance prediction value and the model complexity prediction value; The comprehensive evaluation value is obtained according to the predicted values ​​of each evaluation index and the target search strategy, including: Subtract the model complexity prediction value from the generalization performance prediction value to obtain a comprehensive evaluation value; The iterative search for the neural network structural features along the direction of optimizing the comprehensive evaluation value until the target structural features are obtained includes: The neural network structural features are iteratively searched along the direction of maximizing the comprehensive evaluation value until the target structural features are obtained.

7. The method according to claim 4, characterized in that The evaluation index label values ​​include model complexity and generalization performance; The searched neural network structure features are input into each trained target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network, including: The searched neural network structure features are respectively input into the target model complexity prediction network and the target generalization performance prediction network to obtain the generalization performance prediction value and the model complexity prediction value; According to the predicted values ​​of each evaluation index and the target search strategy, a comprehensive evaluation value is obtained, including: Subtract the generalization performance prediction value from the model complexity prediction value to obtain a comprehensive evaluation value; The iterative search for the neural network structural features along the direction of optimizing the comprehensive evaluation value until the target structural features are obtained includes: The neural network structural features are iteratively searched in the direction of minimizing the comprehensive evaluation value until the target structural features are obtained.

8. The method according to claim 1, characterized in that The step of training the graph neural network, the encoding network, and the decoding network based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure includes: Determine the reconstruction loss of the hidden layer according to the hidden layer structure of the training neural network structure and the hidden layer structure of the reconstructed neural network structure; Based on the reconstruction loss of the hidden layer, the graph neural network, the encoding network and the decoding network are trained.

9. The method according to claim 8, characterized in that Before training the graph neural network, the encoding network, and the decoding network based on the reconstruction loss of the hidden layer, the method further includes: Determine the reconstruction loss of the edge according to the edge of the training neural network structure and the edge of the reconstructed neural network structure; The step of training the graph neural network, the encoding network, and the decoding network based on the reconstruction loss of the hidden layer includes: The graph neural network, the encoding network and the decoding network are trained based on the reconstruction loss of the hidden layer and the reconstruction loss of the edge.

10. The method according to claim 4, characterized in that The encoding network is an encoding network of a variational autoencoder, and the decoding network is a decoding network of a variational autoencoder: the sampling in the target search space to obtain sample points includes: Sampling is performed in the prior distribution of the variational autoencoder to obtain sample points.

11. The method according to any one of claims 1 to 10, characterized in that: The encoding network is an encoding network of a variational autoencoder, and the decoding network is a decoding network of a variational autoencoder: the reconstruction loss between the training neural network structure and the reconstruction neural network structure is used to train the graph neural network, the encoding network, and the decoding network, including: The graph neural network, the encoding network and the decoding network are trained based on relative entropy and the reconstruction loss between the training neural network structure and the reconstruction neural network structure.

12. The method according to any one of claims 1 to 10, characterized in that: The encoding network is an encoding network of a variational autoencoder, and the decoding network is a decoding network of a variational autoencoder: the decoding is performed according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure, including: Sampling is performed according to the continuous structural features to obtain sampling structural features; The sampling structure features are input into the decoding network to obtain the reconstructed neural network structure.

13. A neural network structure search device, characterized in that: The device comprises: A training data acquisition module, used to acquire a neural network structure for text classification or image classification as a training neural network structure; A discrete encoding module, used to input the training neural network structure into a graph neural network to obtain discrete structural features corresponding to the training neural network structure; A continuous encoding module, used for inputting the discrete structural features into an encoding network, and encoding the discrete structural features into corresponding continuous structural features through the encoding network; A decoding module, used for decoding according to the continuous structural features and the decoding network to obtain a reconstructed neural network structure; A training module, used for training the graph neural network, the encoding network and the decoding network based on the reconstruction loss between the training neural network structure and the reconstruction neural network structure, until a training stop condition is met, a target encoding network and a target decoding network are obtained, and a latent space corresponding to the target encoding network is determined as a target search space; A search module is used to search from the target search space according to a target search strategy to obtain target structural features, and decode the target structural features through the target decoding network to obtain a target neural network structure, and the target neural network structure is used for text classification tasks or image classification tasks.

14. The device according to claim 13, characterized in that The device also includes: An evaluation index input module is used to obtain at least one evaluation index label value corresponding to the training neural network structure; predict according to the continuous structural features corresponding to the training neural network structure and each evaluation index prediction network, and obtain the evaluation index training values ​​corresponding to each evaluation index prediction network; determine the evaluation index loss corresponding to each evaluation index label value according to each evaluation index label value and the corresponding evaluation index training value; each evaluation index loss is used to train the corresponding target evaluation index prediction network.

15. The device according to claim 14, characterized in that The training module is also used to jointly train the graph neural network, each evaluation index prediction network, the encoding network and the decoding network based on the reconstruction loss between the training neural network structure and the reconstructed neural network structure, and each evaluation index loss, until the training stop condition is met, thereby obtaining the target evaluation index prediction network, the target encoding network and the target decoding network.

16. The device according to claim 14, characterized in that The search module is also used to sample in the target search space to obtain sample points; search for neural network structural features with the sample points as starting points; input the searched neural network structural features into each trained target evaluation index prediction network to obtain the evaluation index prediction value corresponding to each target evaluation index prediction network; obtain a comprehensive evaluation value based on each evaluation index prediction value and the target search strategy; iteratively search for neural network structural features along the direction of optimizing the comprehensive evaluation value until the target structural features are obtained.

17. The device according to claim 14, characterized in that The at least one evaluation index label value includes at least one of model complexity and generalization performance.

18. The device according to claim 16, characterized in that The evaluation index label values ​​include model complexity and generalization performance; The search module is also used to input the searched neural network structure features into the target model complexity prediction network and the target generalization performance prediction network respectively to obtain a generalization performance prediction value and a model complexity prediction value; subtract the model complexity prediction value from the generalization performance prediction value to obtain a comprehensive evaluation value; The neural network structural features are iteratively searched along the direction of maximizing the comprehensive evaluation value until the target structural features are obtained.

19. The device according to claim 16, characterized in that The evaluation index label values ​​include model complexity and generalization performance; The search module is also used to input the searched neural network structure features into the target model complexity prediction network and the target generalization performance prediction network respectively to obtain a generalization performance prediction value and a model complexity prediction value; subtract the generalization performance prediction value from the model complexity prediction value to obtain a comprehensive evaluation value; The neural network structural features are iteratively searched in the direction of minimizing the comprehensive evaluation value until the target structural features are obtained.

20. The device according to claim 13, characterized in that The training module is also used to determine the reconstruction loss of the hidden layer according to the hidden layer structure of the training neural network structure and the hidden layer structure of the reconstructed neural network structure; based on the reconstruction loss of the hidden layer, train the graph neural network, the encoding network and the decoding network.

21. The device according to claim 20, characterized in that The training module is also used to determine the reconstruction loss of the edges according to the edges of the training neural network structure and the edges of the reconstructed neural network structure; based on the reconstruction loss of the hidden layer and the reconstruction loss of the edges, train the graph neural network, the encoding network and the decoding network.

22. The device according to claim 16, characterized in that The search module is also used to sample in the prior distribution of the variational autoencoder to obtain sample points.

23. The device according to any one of claims 13 to 22, characterized in that The encoding network is the encoding network of a variational autoencoder, and the decoding network is the decoding network of a variational autoencoder: the training module is also used to train the graph neural network, the encoding network and the decoding network based on relative entropy and the reconstruction loss between the training neural network structure and the reconstructed neural network structure.

24. The device according to any one of claims 13 to 22, characterized in that The encoding network is the encoding network of a variational autoencoder, and the decoding network is the decoding network of a variational autoencoder: the decoding module is also used to sample according to the continuous structural features to obtain sampling structural features; the sampling structural features are input into the decoding network to obtain the reconstructed neural network structure.

25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

26. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

27. A computer program product, comprising computer instructions, which when executed by a processor implement the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Neural network architecture evaluation method based on attribute graph optimization

    CN110232434A

  • Neural network structure search method and device, computer equipment and storage medium

    CN112116090A