Neural network architecture search method, device, storage medium and electronic device

By using an initial neural search network composed of multiple search model units for architecture search training, a neural network model that is adapted to different depths is generated, which solves the problem of low search efficiency of neural network architecture in the existing technology, and realizes more efficient network model search.

CN114118403BActive Publication Date: 2025-05-16SHANGHAI JINSHENG COMM TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111213860.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2025-05-16
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

The existing technology requires a lot of research and experiments in the design of neural network model architecture, and the network structure is complex, resulting in low architectural search efficiency.

Method used

By obtaining the initial neural search network composed of at least two types of search model units, performing architecture search training based on business sample data, generating the trained first neural network model, and generating the second neural network model based on the model, avoiding repeated stacking of search model units of the same type of architectural parameters.

Benefits of technology

It adapts to the architectural search needs of different neural network depths, avoids the inefficiency caused by a single network structure, greatly improves the efficiency of neural network search architecture and saves network model search time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114118403B_ABST
    Figure CN114118403B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a neural network architecture search method, device, storage medium and electronic device, wherein the method comprises: obtaining an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type, performing architecture search training processing on the initial neural search network based on business sample data to obtain a trained first neural network model, and generating a second neural network model based on the model parameters corresponding to the first neural network model. By adopting the embodiment of the present application, the efficiency of architecture search can be improved and the search time can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a neural network architecture search method, device, storage medium and electronic device. Background Art

[0002] With the development of computer technology, most neural network model architectures are designed manually. In the process of designing the neural network model architecture, a lot of research and experiments are needed to try and explore the effects of different network model structures. In addition, the structure of the neural network is optimized year by year, and new network structures continue to emerge and become more and more complex.

[0003] Neural Architecture Search (NAS), as a technology that can automatically design neural network structures, has attracted the attention of more and more researchers. The best architecture designed by NAS has achieved performance that exceeds that of manually designed network architectures in a variety of tasks, such as image classification, semantic segmentation, object detection, etc. Summary of the invention

[0004] The embodiment of the present application provides a neural network architecture search method, device, storage medium and electronic device, which can allocate a service thread to a suitable processor cluster. The technical solution of the embodiment of the present application is as follows:

[0005] In a first aspect, an embodiment of the present application provides a neural network architecture search method, the method comprising:

[0006] Acquire an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type;

[0007] Performing architecture search training on the initial neural search network based on the business sample data to obtain a trained first neural network model;

[0008] A second neural network model is generated based on the model parameters corresponding to the first neural network model, and the number of search model units corresponding to the second neural network model is greater than or equal to the number of search model units corresponding to the first neural network model.

[0009] In a second aspect, an embodiment of the present application provides a neural network architecture search device, the device comprising:

[0010] A network acquisition module, used to acquire an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type;

[0011] A search training module, used to perform architecture search training processing on the initial neural search network based on business sample data to obtain a trained first neural network model;

[0012] A model determination module is used to generate a second neural network model based on the model parameters corresponding to the first neural network model, and the number of search model units corresponding to the second neural network model is greater than or equal to the number of search model units corresponding to the first neural network model.

[0013] In a third aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0014] In a fourth aspect, an embodiment of the present application provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0015] The beneficial effects brought about by the technical solutions provided by some embodiments of the present application include at least:

[0016] In one or more embodiments of the present application, an electronic device obtains an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters of the same architecture type, and then performs architecture search training processing on the initial neural search network based on business sample data to obtain a trained first neural network model, and finally generates a second neural network model based on the model parameters corresponding to the first neural network model. By avoiding using the same type of search model units with the same architecture parameters to build the initial neural search network, the architecture search requirements of different neural network depths can be adapted, and the low efficiency and long time of architecture search caused by a single network structure can be avoided, thereby greatly improving the efficiency of the neural network search architecture and saving network model search time. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a flowchart of a neural network architecture search method provided in an embodiment of the present application;

[0019] Figure 2 It is a scene graph of a neural network architecture search involved in the neural network architecture search method provided in an embodiment of the present application;

[0020] Figure 3 It is a schematic diagram of the internal structure of a search model unit involved in the neural network architecture search provided in an embodiment of the present application;

[0021] Figure 4 It is a schematic diagram of the internal structure of another search model unit involved in the neural network architecture search provided in an embodiment of the present application;

[0022] Figure 5 It is a flowchart of a target type determination module provided in an embodiment of the present application;

[0023] Figure 6 It is a schematic diagram of the architecture of an initial neural search network involved in the neural network architecture search method provided in an embodiment of the present application;

[0024] Figure 7 It is a schematic diagram of the architecture of a sub-network involved in the neural network architecture search method provided in an embodiment of the present application;

[0025] Figure 8 It is a schematic diagram of the network architecture search scenario of an initial neural search network;

[0026] Fig. 9 is a structural schematic diagram of another neural network architecture search device provided in an embodiment of the present application;

[0027] Fig.10 is a structural schematic diagram of another neural network architecture search device provided in an embodiment of the present application;

[0028] Fig.11 is a structural diagram of a network acquisition module provided in an embodiment of the present application;

[0029] Fig.12 It is a schematic diagram of the structure of the operating system and user space provided in the embodiment of the present application;

[0030] Fig.13 yes Fig.11 The architecture diagram of the Android operating system;

[0031] Fig.14 yes Fig.11 Architecture diagram of the IOS operating system. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0033] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood in specific circumstances. In addition, in the description of the present application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are an "or" relationship.

[0034] The present application is described in detail below with reference to specific embodiments.

[0035] In one embodiment, Figure 1 As shown, a neural network architecture search method is proposed, which can be implemented by a computer program and can be run on a neural network architecture search device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0036] Specifically, the neural network architecture search method includes:

[0037] Step S101: obtaining an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type.

[0038] In practical applications, the process of neural network architecture search can be simplified as follows Figure 2 As shown, Figure 2It is a scene graph of neural network architecture search. The neural network architecture search constructs a suitable neural search network for architecture search, that is, the initial neural search network of this application, and then obtains a candidate neural network architecture (such as Figure 2 The network architecture x shown in the figure) and the candidate neural network architecture (such as Figure 2 The network architecture X shown in the figure is used for network performance evaluation. Based on the evaluation result (performance of network architecture X), feedback is given to the architecture search method and corresponding adjustments are made to optimize the architecture search. The above process is repeated until the expected neural network architecture is found. It can be understood that a specific network architecture x is selected from the pre-built initial neural search network. The network architecture x will be evaluated for performance, and the performance estimation structure of the network architecture x will be obtained. Then, the performance estimation structure is fed back to the search method (module, which can also be understood as a search strategy). In the process of searching for a neural network architecture with excellent performance, sampling will be continuously performed from the initial neural search network, and the whole process of normal neural network model training will be performed. Finally, the best model that meets expectations is selected as the output of the algorithm in the whole process, which usually requires a lot of computing resources and computing time. Thus, a neural network model corresponding to the neural network architecture is finally generated.

[0039] In the present application, the initial neural search network used for neural network architecture search is not generated by repeated superposition of search model units of the same type (cell). In the related art, cells of the same type, i.e., cells of the same architecture type and with the same (initial) architecture parameters, are often used to repeatedly stack to form the initial neural search network for neural network architecture search. Such a search network search space often encounters bottlenecks in the model search process during network search and even network training, making it difficult to achieve good network performance. In the present application, it is found through practice that the requirements for network architecture or structure are actually different during network architecture search at different network depths. The initial neural search network generated by repeated superposition of search model units of the same type (cell) usually presents the same network architecture or network structure at different network depths. In the present application, the aforementioned concept is not used in the stage of creating the initial neural search network, but cells of different types, i.e., search model units corresponding to different architecture parameters, are used to adapt to the network structure requirements at different network depths. It can be understood that cells of different network depths have different architecture parameters.

[0040] The search model unit (cell) is the basic unit for constructing the initial neural search network. The search model unit can be used as a directed acyclic graph. Each search model unit consists of i (i is an integer greater than 1) ordered nodes, each node represents a feature graph, and each directed edge represents an operator, which is a number of candidate operations (such as pooling, convolution, etc.) used to process the input feature graph. For example, the directed edge (i, j) represents the connection relationship from node i to node j, and the operator o∈O on the directed edge (i, j) is used to convert the feature graph x_i input by node i into the feature graph x_j. Among them, O represents all candidate operations in the search space. The structure based on the search model unit (cell) can make the network "deep" and "wide" due to its specific hierarchical evolution. Based on this concept, different types of neural networks can often be generated through neural network architecture search, such as convolutional neural network CNN and recurrent neural network RNN.

[0041] Indicatively, Figure 3 As shown, Figure 3 is a schematic diagram of the internal structure of a search model unit. Figure 3 The cell shown contains 7 nodes. The first two nodes are input nodes (i.e. input 1 and input 2), which are obtained from the outputs of the previous two cells respectively. The next 4 nodes are intermediate nodes (i.e. the 4 nodes shown as "0", "1", "2", and "3" in the figure). Each intermediate node is calculated by all the previous nodes. It can be understood that the input of any intermediate node comes from all the forward nodes of the node in this cell. The last node is the output node, which is the connection of the feature vectors of the intermediate nodes and represents the output of the entire cell.

[0042] Indicatively, Figure 4 As shown, Figure 4 This is another schematic diagram of the internal structure of the search model unit. Figure 4 The cell shown contains 4 nodes, which are composed of 4 nodes (1 input node x1, 2 intermediate nodes x2 and x3, and 1 output node x4). If the final output node x4 is to be obtained, the values ​​of nodes x2 and x3 must first be calculated through node x1 and operations O1 and O2. Nodes x2 and x3 take O3 and O4 operations respectively to obtain the result of x4.

[0043] It is understandable that the lines between the two nodes in the figure represent the operations between the nodes. The candidate operations can be one or more of the following forms:

[0044] 3x3 separable convolution, 5x5 separable convolution, 3x3 dilated separable convolution operation, 5x5 dilated separable convolution operation, 3x3 maximum pooling layer operation, 3x3 average pooling layer operation, identity transformation operation, zero, i.e. no connection operation, etc. These operations correspond to weight parameters in the initial neural search network, which can be weighted by the softmax function and relaxed to the continuous space; in the subsequent training process, the structure is mainly optimized by adjusting the weights of these operations (i.e. updating parameters based on back propagation). The last node is the output node, which is a cascade operation of 4 intermediate nodes.

[0045] Furthermore, the search model unit (cell) is divided into two types, namely, ordinary unit NC and compression unit RC. The difference is: NC: the number of channels before and after NC remains unchanged, and the length and width of its input and output data features are the same; RC: after the compression unit RC, the number of channels becomes 1 / 2 of the original, and the length and width of its input data features are twice that of the output.

[0046] It should be noted that the "two types" in the aforementioned "at least two types of search model units" does not mean "two architectural types of ordinary units NC and compressed units RC". "At least two types of search model units" can be understood as: at least two types of ordinary units NC and / or at least two types of compressed units RC when broken down into specific ordinary units NC and compressed units RC. The difference is that the architectural parameters of ordinary units NC or compressed units RC of the same architectural type in the entire initial neural search network can be different.

[0047] The (unit) architectural parameters of the search model unit include, but are not limited to: the candidate operation types corresponding to each node in the search model unit, the operation parameters of the candidate operations corresponding to each node in the search model unit, the weights (also understood as weights) of the candidate operations corresponding to each node in the search model unit, and the unit model parameters of the search model unit (a single search model unit can be regarded as a micro neural network). In some embodiments, the architectural parameters of the search model unit can also be the activation function of the initial state (in some implementation scenarios, the activation function can be further neurally evolved with training).

[0048] The at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type.

[0049] It can be understood that at least two types of search model units are stacked to construct at least two sub-networks of different network depth types, and an initial neural search network including each of the sub-networks is generated;

[0050] Among them, the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of the same network depth type are the same; and the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of different network depth types are different.

[0051] Optionally, for a single search model unit (NC or RC), at least two small-scale data sets of the same type as the business sample data can be used to train the search model units respectively; usually, the architecture parameters of the search model units trained using different data sets are usually different, that is, the search model units with different aforementioned architecture parameters can be obtained by training the model units based on this.

[0052] Optionally, one or more architectural parameters in the above-mentioned trained search model unit may be modified, such as by using an expert end to perform the modification, so that different types of search model units may be obtained. The specific modification strategy is quantitatively determined based on the actual business form.

[0053] Furthermore, the above-mentioned construction process of the neural search network can be completed on the electronic device itself; it can also be completed and stored on the server. The electronic device only needs to send a request to the server to obtain the initial neural search network on the server composed of at least two types of search model units.

[0054] Step S102: Performing architecture search training processing on the initial neural search network to obtain a trained first neural network model.

[0055] According to some implementation steps, the neural network architecture search is performed by obtaining an initial neural search network, which is composed of at least two types of search model units to adapt to the structural requirements at different network depths. Then, a candidate neural network architecture (such as Figure 2 The network architecture x shown in the figure) and the candidate neural network architecture (such as Figure 2 The network performance is evaluated based on the network architecture X shown in the figure. Based on the evaluation result (the performance of the network architecture X), the feedback is given to the architecture search method and corresponding adjustments are made to optimize the architecture search. The above process is repeated until a neural network architecture that meets the expectations is found. It can be understood that a specific network architecture x is selected from the pre-built initial neural search network, and the network architecture x will be evaluated for performance to obtain a performance estimation structure of the network architecture x, which is then fed back to the search method (module, which can also be understood as a search strategy). In the process of searching for a neural network architecture with excellent performance, sampling is continuously performed from the initial neural search network, and the entire process of normal neural network model training is performed. Finally, the best model that meets expectations is selected as the output of the algorithm in the entire process, thereby finally generating a neural network model corresponding to the neural network architecture.

[0056] In a specific implementation scenario, a neural network architecture search can be implemented in a differentiable manner, that is, the model architecture parameters and model weight parameters of the neural network are searched simultaneously. Since in the initial neural search network stage, for each search model unit in at least two types of search model units: a set of candidate operations (including multiple candidate operations) constituting the network is set between two nodes in each search model unit, and the corresponding weight parameters of these operations in the initial neural search network are specifically weighted by the softmax function, and relaxed to the continuous space corresponding to the initial neural search network. In the specific search stage, business sample data is input into the initial neural search network for network training and neural network architecture search, and the candidate operations are weighted using the softmax function during the search process. The performance evaluation strategy is used to evaluate the network performance of the neural network architecture obtained each time by weighting, and the back propagation algorithm (BP algorithm) is used to continuously adjust the search strategy based on the evaluation results, and the network structure and model parameters are jointly optimized, and then the softmax function is used for weighting during the optimization process, and the parameters of the entire network are updated through back propagation. In some implementation methods, only the operation with the largest corresponding weight on each connection of the network cell is retained until the initial neural search network converges to obtain the first neural network model.

[0057] In some embodiments, the first neural network model can be regarded as a proxy model of the neural network model finally obtained. It can be understood that in order to improve the network search efficiency and save the network architecture search time, the performance evaluation strategy adopts the strategy of the proxy model, and the first neural network model is obtained based on the proxy model strategy. The whole process does not focus on the verification error determination of the network, but adopts the concept of approximate task substitution, replaces the time verification task, and replaces the actual network error with the result obtained by the proxy model in the proxy task using the business sample data. This can save a lot of evaluation time for evaluating the model at each stage. In this application, the model obtained on the proxy task corresponding to the proxy task (which can be regarded as an approximate replacement business for the actual business) based on the proxy model strategy can be migrated to the target task based on the means of model data migration. Finally, the neural network model that is expected to be obtained is obtained through training.

[0058] Step S103: generating a second neural network model based on the model parameters corresponding to the first neural network model.

[0059] In some embodiments, the number of search model units corresponding to the second neural network model is greater than or equal to the number of search model units corresponding to the first neural network model.

[0060] In practical applications, model parameters of a neural network model include but are not limited to model architecture parameters and model weight parameters.

[0061] It is understandable that the number of search model units corresponding to the second neural network model can be determined based on the model application business. One way can be to establish a mapping relationship between the model application business and the number of reference search model units in advance. The model application business can feedback the situation of the neural network model expected to be established to a certain extent, such as image classification business, semantic segmentation business, object recognition business, speech recognition business, etc. In the aforementioned different model application businesses, the number of reference search model units can be a reference range used in the same type of business based on actual model development experience or a determined parameter value. The mapping relationship can be represented in the form of a mapping set, a mapping linked list, a table, an array, etc. In the actual application process, based on the corresponding model application business, the user of the electronic device can usually determine the number of search model units required for the final generation of the second neural network model based on the actual model development experience. For example, if the CNN network model is often used for a certain type of business, the number of search model units usually required for the CNN network model can be determined based on experience as a reference value or reference range.

[0062] After determining the number of search model units corresponding to the second neural network model, the model parameters of the first neural network can be determined based on the neural network architecture search method to initialize the second neural network model.

[0063] Optionally, when the number of search model units corresponding to the second neural network model is equal to the number of search model units corresponding to the first neural network model, the first neural network model may be used as the second neural network model.

[0064] Optionally, when the number of search model units corresponding to the second neural network model is greater than the number of search model units corresponding to the first neural network model, the model can be expanded based on the model parameters of the first neural network model to expand the first neural network model into the second neural network model. The specific expansion process is mainly based on the proportional expansion of the number of search model units of the first neural network model to increase the number of cells. For example, the number of search model units is expanded by n times (n is a positive integer). The expanded search model unit can inherit the model architecture parameters of the search model unit of the aforementioned first neural network model, so that the expanded initial second neural network model can be obtained. Then, the business sample data corresponding to the model application business can be used to train the initial second neural network model. The model weight parameters of the second neural network model are trained through model training. It can be understood that the model architecture parameters have been determined during the architecture search training process. In the model training process, the model architecture parameters are usually unchanged, and only the model weight parameters are optimized. The specific model training method can be implemented using conventional neural network model training technology, which will not be described in detail here.

[0065] Furthermore, in the process of generating the initial second neural network, the model weight parameters of the expanded initial second neural network model may also inherit the first neural network model; the model weight parameters of the expanded initial second neural network model may also not inherit the first neural network model, but only inherit the architecture parameters of the search model unit when expanding and generating the initial second neural network model.

[0066] In the actual application process, considering that different model application businesses will involve different numbers of search model units corresponding to the final expected neural network model, the model application business with a higher degree of computational processing requires more search model units. In order to save computing resources and reduce computing time, a proxy model with fewer search model units can be trained first to obtain the first neural network model. At this time, the model architecture of the first neural network model is usually highly similar to the model architecture of the final expected neural network model. The initial second neural network model is regenerated based on the model parameters of the first neural network model, and then the sample is used to continue the model training to optimize the model weights of the initial second neural network model. After the training is completed, the final second neural network model can be obtained, completing the entire neural network architecture search process. The architecture search time is greatly reduced.

[0067] It can be understood that the electronic device obtains the model architecture parameters corresponding to the first neural network model in the aforementioned manner, and determines the unit expansion ratio for the first neural network model; then generates an initial second neural network model based on the unit expansion ratio, the model architecture parameters and the first neural network model. Furthermore, in a specific application, the electronic device expands the number of search model units of the first neural network model based on the unit expansion ratio to obtain an initial neural network model; updates the model architecture of the initial neural network model based on the model architecture parameters to generate an initial second neural network model; wherein, the training of the initial second neural network model mainly involves model training of the model weight parameters of the initial second neural network model, and after model training of the initial second neural network model, a trained second neural network model is obtained.

[0068] In an embodiment of the present application, an electronic device obtains an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters of the same architecture type, and then performs architecture search training processing on the initial neural search network based on business sample data to obtain a trained first neural network model, and finally generates a second neural network model based on the model parameters corresponding to the first neural network model. By avoiding using the same type of search model units with the same architecture parameters to build the initial neural search network, the architecture search requirements of different neural network depths can be adapted, and the low efficiency and long time of architecture search caused by a single network structure can be avoided, thereby greatly improving the efficiency of the neural network search architecture and saving network model search time.

[0069] See also Figure 5 , Figure 5 This is a flowchart of another embodiment of a neural network architecture search method proposed in this application. Specifically:

[0070] Step S201: stacking at least two types of search model units to construct at least two sub-networks of different network depth types.

[0071] Step S202: Generate an initial neural search network including each of the sub-networks.

[0072] In the present application, the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of the same network depth type are the same; and the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of different network depth types are different.

[0073] The network depth type is determined based on the network layer depth of the sub-network in the initial neural search network. For example, the initial neural search network can be divided into a shallow sub-network with a shallow network layer depth, a middle sub-network with a medium network layer depth, and a deep sub-network with a deep network layer depth. For another example, the initial neural search network can be divided into a first sub-network of a first network depth type, a second sub-network of a second network depth type...and an Nth sub-network of an Nth network depth type (N is an integer greater than 2).

[0074] Furthermore, the search model unit of the i-th subnetwork of the i-th (i is an integer less than N) network depth type can be called cell-i. There are two types of search model units, namely, two architectural types: common unit NC and compression unit RC. The common unit NC of the i-th subnetwork can be NC-i, and the common unit RC of the i-th subnetwork can be RC-i; for example, the common unit NC-1 of the 1st subnetwork, the common unit RC-4 of the 4th subnetwork... and so on.

[0075] "The architectural parameters corresponding to the search model units of the same architectural type in the sub-network of the same network depth type are the same" can be understood as: the architectural parameters of all common units NC-i of the i-th sub-network of the i-th network depth type are the same, and / or, the architectural parameters of all compression units RC-i of the i-th sub-network of the i-th network depth type are the same.

[0076] “The architectural parameters corresponding to the search model units of the same architectural type in the sub-networks of different network depth types are different” can be understood as: the architectural parameters corresponding to cell-i of the i-th sub-network of the i-th network depth type and cell-j of the j-th sub-network of the j-th (j is an integer less than N, and i and j are different) network depth type are different, for example, the architectural parameters of RC-i and RC-j are different, and the architectural parameters of NC-i and NC-j are different.

[0077] The (unit) architectural parameters of the search model unit include, but are not limited to: the candidate operation types corresponding to each node in the search model unit, the operation parameters of the candidate operations corresponding to each node in the search model unit, the weights (also understood as weights) of the candidate operations corresponding to each node in the search model unit, and the unit model parameters of the search model unit (a single search model unit can be regarded as a micro neural network). In some embodiments, the architectural parameters of the search model unit can also be the activation function of the initial state (in some implementation scenarios, the activation function can be further neurally evolved with training).

[0078] In addition, the (unit) architecture parameters of various search model units in the initial neural search network are set in combination with different network depths according to the actual situation of the model application task. For example, the reference neural network model developed for the same type of model application task is often analyzed and processed based on the model development experience, and the unit architecture commonality of the search model unit of the reference neural network model that processes such model application tasks is set based on a large number of reference neural network models. The unit architecture commonality is mainly reflected in the weight characteristics of the candidate operations of each node at the same network depth, the unit model parameter characteristics of the search model unit, the final candidate operation type characteristics corresponding to the nodes, and so on. The construction of the entire initial neural search network involved in this application does not adopt the method of simply stacking cells with the same architecture parameters, so as to avoid a single architecture that is not suitable for unit architecture requirements at different network depths. For example: in the deep-level architecture corresponding to the network depth, the trained reference neural network model often sets the initial candidate operations of the nodes in the search model unit to more jump connection operations, because the gradient drops fastest in the jump connection, which can accelerate the back propagation process.

[0079] It should be noted that the architectural parameters corresponding to the cells of sub-networks of different network depth types are determined based on actual business application situations. The examples involved in this application are only for illustrative purposes. Those skilled in the art should understand that the examples involved do not impose any limitations on this application.

[0080] In a feasible implementation, the timing of the neural network architecture search of the sub-networks of the initial neural search network with different network depth types is different, and the timing of the first network architecture search of the sub-network with a shallower network depth is earlier. For example, the timing of the network architecture search of the i-th sub-network is before the timing of the network architecture search of the i+1-th sub-network.

[0081] In a feasible implementation, the network architecture search method for the initial neural search network is:

[0082] Phase 1 network architecture search: Perform network architecture search on the first sub-network;

[0083] Phase 2 network architecture search: Perform network architecture search on the first sub-network and the second sub-network trained in the first phase. This can be understood as adding the first sub-network trained in the first phase to the second sub-network for network architecture search.

[0084] Phase 3 network architecture search: The “1st sub-network and 2nd sub-network” trained in phase 2 are added to the 3rd sub-network for network architecture search; ....

[0086] Network architecture search in the Nth stage: The “1st sub-network, 2nd sub-network..., N-1th sub-network” trained in the N-1th stage are added to the Nth sub-network for network architecture search.

[0087] In a specific implementation scenario, the initial neural search network can be represented as consisting of n sub-(search) networks that increase in depth with the network, such as Figure 6 As shown, Figure 6 is a schematic diagram of the architecture of an initial neural search network, such as Figure 6 In some implementations, each subnetwork is associated with each other (can be regarded as two-by-two subnetworks connected), that is, the i-th subnetwork is associated with the i-1-th subnetwork and the i+1-th subnetwork. The network depth of the i-th subnetwork is greater than the network depth of the i-1-th subnetwork and less than the network depth of the i+1-th subnetwork. Each subnetwork is composed of a class of search model units of corresponding network depth.

[0088] Optionally, the number of subnetworks included in the initial neural search network is pre-set by the user based on actual business conditions. Model application services of different business types can set different numbers of subnetworks based on experience. In some embodiments, a specification mapping relationship between the business type of the model application business and the subnetwork specification corresponding to the initial neural search network can be pre-established. The subnetwork specification can be understood as the number of subnetworks, the cell architecture contained in the subnetwork, the number of cells in the subnetwork, the subnetwork type, etc. In actual applications, the subnetwork specification corresponding to the initial neural search network can be determined in the aforementioned specification mapping relationship according to the current model application business, and then the initial neural search network can be built, that is, the initial neural search network containing each subnetwork is generated.

[0089] Optionally, in the subnetwork corresponding to the initial neural search network, a subnetwork may be composed of a stack of search model units with the same architectural parameters, which can be understood as a subnetwork may be composed of a stack of NCs and / or RCs. In some embodiments, the number of cells in some or all of the subnetworks in the initial neural network is usually the same, and the difference is the architectural parameters of the cells; in some embodiments, the first subnetwork or the last subnetwork usually has a different number of cells from other subnetworks in the initial neural search network.

[0090] For example, the first subnetwork to the N-1th subnetwork are composed of a number of NCs and b number of RCs, and the last subnetwork, that is, the Nth subnetwork, is composed of c number of NCs, where a, b, and c are positive integers.

[0091] In a specific implementation scenario, the following explanation is given by taking an initial neural search network constructed by search model units with three different architecture parameters as an example:

[0092] 1. The electronic device determines three types of search model units with different architecture parameters, and the architecture parameters corresponding to the search model units of the same architecture type in each type of search model unit are the same.

[0093] According to some embodiments, the electronic device can use the intervention of the expert end to analyze and process the reference neural network model developed for the same type of model application task based on the model development experience, and set it based on the commonality of the unit architecture of the search model unit of the reference neural network model for processing such model application tasks in a large number of reference neural network models; the commonality of the unit architecture is mainly reflected in the weight characteristics of the candidate operations of each node at the same network depth, the unit model parameter characteristics of the search model unit, the final candidate operation type characteristics corresponding to the nodes, etc. The construction of the entire initial neural search network involved in this application does not adopt the method of simply stacking cells with the same architectural parameters to avoid the single architecture being unsuitable for unit architecture requirements at different network depths. In some embodiments, different types of "model application task types" and "architecture parameter types corresponding to search model units" can be pre-created. Based on this type mapping relationship, the electronic device can determine the types of architecture parameters of at least two search model units currently required to be adopted according to the scenario of the model application task. Schematically, based on the above method in an application scenario, the electronic device uses three types of search model units with different architecture parameters, which can be represented as cell-1, cell-2, and cell-3. In some implementations, the architecture parameter type may be a parameter of the aforementioned sub-network specification, that is, the architecture mapping relationship may be a sub-relationship in the aforementioned specification mapping relationship.

[0094] It can be understood that the three types of search model units with different architectural parameters are respectively used in different sub-networks, such as "cell-1, cell-2, cell-3" with three different architectural parameters are respectively used in sub-networks of three different network depth types to construct an initial neural search network.

[0095] The three types of search model units with different architectural parameters are embodied in the following architectural parameters, including but not limited to: the candidate operation types corresponding to each node in the search model unit, the operation parameters of the candidate operations corresponding to each node in the search model unit, the weights (also understood as weights) of the candidate operations corresponding to each node in the search model unit, and the unit model parameters of the search model unit (a single search model unit can be regarded as a micro neural network). In some embodiments, the architectural parameters of the search model unit can also be the activation function of the initial state (in some implementation scenarios, the activation function can be further neurally evolved with training).

[0096] Three types of search model units with different architecture parameters can be represented as cell-1, cell-2, and cell-3, where cell-1 is the first type of search model unit, cell-2 is the second type of search model unit, and cell-3 is the third type of search model unit.

[0097] It should be noted that the specific parameter composition of the search model units of various different architecture parameters is not specifically limited here and is determined based on the actual application environment.

[0098] 2. The electronic device constructs a first sub-network by stacking the first type of search model units among the three types of search model units; constructs a second sub-network by stacking the second type of search model units among the three types of search model units; and constructs a third sub-network by stacking the third type of search model units among the three types of search model units.

[0099] The network depth of the second sub-network is greater than the network depth of the first sub-network and less than the network depth of the third sub-network.

[0100] In a specific implementation scenario, the first type of search model unit used to construct the first sub-network may include the first type of common unit and the first type of compression unit, which can be recorded as: NC-1, RC-1. The second type of search model unit used to construct the second sub-network may include the second type of common unit and the second type of compression unit, which can be recorded as: NC-2, RC-2. The third type of search model unit used to construct the third sub-network may include the third type of common unit, which can be recorded as: NC-3.

[0101] Further, the first type of search model unit includes a first type of common unit and a first type of compression unit, the second type of search model unit includes a second type of common unit and a second type of compression unit, and the third type of search model unit includes a third type of common unit;

[0102] like Figure 7 As shown, the first sub-network is composed of two of the first-class common units NC-1 and one of the first-class compression units RC-1 stacked; the second sub-network is composed of two of the second-class common units NC-2 and one of the second-class compression units RC-2 stacked; the third sub-network is composed of two of the third-class common units NC-3 stacked. The above structure can meet different task requirements and adapt to unit architecture requirements at different network depths in some scenarios, greatly improve architecture search efficiency, and save network model search time.

[0103] Step S203: Based on the business sample data, the model weight parameters and model architecture parameters of at least two sub-networks included in the initial neural search network are trained and optimized according to the network depth to obtain a first neural network model after training and optimization.

[0104] In this application, in order to achieve a better model effect and avoid the direct architecture search of a single-structure network architecture in related technologies, at least two sub-networks are deployed for the initial neural search network based on at least two types of search model units. In this structure, the search unit structure of each sub-network is different to meet different task requirements. The search unit structure of each sub-network is optimized by the automatic update of the network. Through multiple rounds of iterative training, a better model effect can be achieved.

[0105] Furthermore, when the present application performs architecture search training on the initial neural search network, sample business data can be used to directly perform architecture search training on the complete initial neural search network to obtain a first neural network model after training and optimization.

[0106] Furthermore, considering the search cost and search efficiency of the network, the search architecture of each sub-network is gradually trained by partial training (which can be understood as molecular network training) based on the number of sub-networks of the initial neural search network. It can be understood that this application does not train each sub-network individually but is carried out in a partial and progressive manner, as follows:

[0107] S2031: The electronic device determines a current target sub-network from at least two sub-networks included in the initial neural search network according to the network depth;

[0108] It can be understood that the initial neural search network is a super network composed of multiple sub-networks stacked and connected, and each sub-network is composed of at least one type of search model unit stacked.

[0109] S2032: training and optimizing the model weight parameters and model architecture parameters of the target sub-network based on the business sample data to obtain an optimized neural network model;

[0110] The architecture search process of the initial neural search network is usually trained from the first sub-network to the last sub-network one by one. Therefore, in each round of sub-network training optimization process, the target sub-network to be trained is first determined, and then the target sub-network is trained and optimized using the business sample data.

[0111] Because in the initial neural search network stage, for each search model unit in each sub-network contained in the initial neural search network: a set of candidate operations that constitute the network is set between two nodes in each search model unit (including multiple candidate operations), and the corresponding weight parameters of these candidate operations in the initial neural search network are specifically assigned weights by the softmax function and relaxed to the continuous space corresponding to the initial neural search network.

[0112] In the specific search stage, business sample data is input into the current target sub-network corresponding to the initial neural search network for network training and network architecture search. During the search process, the softmax function is used to weight the candidate operations at the nodes of each search model unit of the target sub-network. The neural network architecture of the weighted target sub-network is evaluated using a performance evaluation strategy. The search strategy is continuously adjusted based on the evaluation results using a back propagation algorithm (BP algorithm). The model weight parameters and model architecture parameters of the target sub-network are jointly optimized. Then, the softmax function is used for weighting during the optimization process, and the parameters of the entire target sub-network are updated through back propagation. In some implementation methods, only the operation with the largest weight on each connection at the node corresponding to the cell of the target sub-network is retained until the initial neural search network converges to obtain the optimized neural network model.

[0113] S2033: If there is a next sub-network corresponding to the next network depth of the target sub-network, and the next sub-network is connected to the neural network model, the neural network model is updated to the target sub-network and the steps of training and optimizing the model weight parameters and model architecture parameters of the target sub-network are performed based on the business sample data;

[0114] Connecting the next sub-network to the neural network model can be understood as: the electronic device determines the first search model unit corresponding to the tail end of the neural network model and determines the second search model unit corresponding to the head end of the next sub-network; and then performs unit connection processing on the first search model unit and the second search model unit.

[0115] It can be understood that the network architecture search method for the initial neural search network is:

[0116] Phase 1 network architecture search: Perform network architecture search on the first sub-network;

[0117] Phase 2 network architecture search: Perform network architecture search on the first sub-network and the second sub-network trained in the first phase. This can be understood as adding the first sub-network trained in the first phase to the second sub-network for network architecture search.

[0118] Phase 3 network architecture search: The “1st sub-network and 2nd sub-network” trained in phase 2 are added to the 3rd sub-network for network architecture search;

[0119] Network architecture search in the Nth stage: The “1st sub-network, 2nd sub-network..., N-1th sub-network” trained in the N-1th stage are added to the Nth sub-network for network architecture search.

[0120] Schematically, the electronic device first performs a first-stage network structure search, at which time the target sub-network is also the first sub-network, and performs a network architecture search process on the first sub-network. After the target sub-network completes the first-stage network structure search, the optimized first sub-network is used as a neural network model;

[0121] Perform the second-stage network architecture search to obtain the "next subnetwork corresponding to the next network depth of the target subnetwork", that is, the "second subnetwork", connect the next subnetwork (second subnetwork) to the neural network model, and then update the neural network model to the network architecture search processing object of the next stage, that is, update the neural network model to the target subnetwork, and use business sample data to execute the network architecture search processing process for the target subnetwork.

[0122] ...and so on.

[0123] Perform the i-th (i is a positive integer less than N) stage network architecture search to obtain the "next subnetwork of the target subnetwork corresponding to the next network depth", that is, the "i-th subnetwork". If there is a next subnetwork of the target subnetwork corresponding to the next network depth, connect the next subnetwork (i-th subnetwork) to the neural network model, and then update the neural network model to the network architecture search processing object of the next stage, that is, update the neural network model to the target subnetwork, and use business sample data to perform the network architecture search processing process for the target subnetwork.

[0124] ...and so on

[0125] S2034: If there is no next sub-network corresponding to the next network depth of the target sub-network, the neural network model is used as the first neural network model after training and optimization.

[0126] “There is no next subnetwork corresponding to the next network depth of the target subnetwork” can be understood as the network architecture search process for the last subnetwork has been completed.

[0127] It should be noted that in each of the aforementioned stages of the network architecture search and processing process for the target subnetwork, the sample business data in each stage may be the same, partially the same, or completely different. For example, the total business data may be evenly distributed according to the number of subnetworks, with each portion serving as sample business data.

[0128] Furthermore, in each stage of the training process of the target sub-network, the number of training rounds y can be set, and the target sub-network is trained according to the number of training rounds y.

[0129] If the total number of rounds for the initial neural search network is set to X, and the number of sub-networks is m, then y = X / m;

[0130] In a specific implementation scenario, the initial neural search network includes the first sub-network, the second sub-network and the third sub-network as an example. Figure 6 As shown in the sub-networks, the first sub-network is composed of 2 NC-1s and 1 RC-1, the second sub-network is composed of 2 NC-2s and 1 RC-2, and the third sub-network is composed of 2 NC-3s. Figure 8 As shown, Figure 8 It is a schematic diagram of the network architecture search scenario of an initial neural search network.

[0131] Specifically, yes Figure 8 The way of showing the network architecture search performed by the initial neural search network is:

[0132] Phase 1 network architecture search: Input sample business data into the first sub-network for network architecture search processing. Assuming that the total number of training rounds of the initial neural search network is N, we set the number of training rounds of the first sub-network (equivalent to the shallow network) in Phase 1 to 1 / 3*N.

[0133] Phase 2 network architecture search: Perform network architecture search on the first subnetwork and the second subnetwork trained in the first phase, which can be understood as adding the first subnetwork trained in the first phase to the second subnetwork for network architecture search processing; which is equivalent to training the shallow network (the first subnetwork) to a certain extent to obtain a neural network model containing the first subnetwork, adding the middle network (the second subnetwork) to the neural network model for training, and setting the number of training rounds to 1 / 3*N. In this way, the unit structures in the shallow and middle networks are set to be different, which can adapt to the needs of networks at different depths.

[0134] The third stage is network architecture search: the "1st sub-network and 2nd sub-network" trained in the second stage are added to the 3rd sub-network for network architecture search; after the neural network models corresponding to the shallow network and the middle network (1st sub-network and 2nd sub-network) are trained, the two common units NC in the deep network are added to the neural network model for training, and the number of training rounds is set to 1 / 3*N. At this point, the complete proxy model, that is, the first neural network model, can be obtained. This step mainly completes the training and optimization of the entire network. Subsequently, the initial second neural network model is generated based on the model parameters of the first neural network model, and then the initial second neural network model is finally trained using sample business data to obtain the final model - the "second neural network model".

[0135] In a feasible implementation, at least one first search model unit in the initial neural search network is associated with a second search model unit and a third search model unit located before the first search model unit. It can be understood that starting from the third search model unit: the input of each search model unit (first search model unit) is the output of the first two units (second search model unit and third search model unit) of the search model unit. Furthermore, within each search model unit, after calculation through the forward node, the feature graphs on the four intermediate nodes are aggregated to become the output of the search model unit.

[0136] It is understandable that, in the above: connecting the next sub-network to the neural network model, updating the neural network model to the target sub-network, in the process of "connecting the next sub-network to the neural network model", one way may be to connect directly without changing the network parameters of the next sub-network, that is, to retain the initial network parameters of the next sub-network (including model weight parameters and model architecture parameters) in the connected neural network model.

[0137] Optionally, considering the search cost and effect of the network, the network diversity of the initial neural search network in this application is greatly improved compared with the prior art. One method can also be after the process of "connecting the next sub-network to the neural network model": the initial model parameters of the next sub-network newly connected to the neural network model can be optimized, thereby reducing the training cycle and improving the search efficiency. It can also ensure that the model parameters such as the model weights of the neural network model are fully trained. For details, please refer to the following method involving the sub-network layer inheritance initialization method. As follows:

[0138] S1: Determine the current target sub-network from at least two sub-networks included in the initial neural search network according to the network depth;

[0139] S2: Based on the business sample data, the model weight parameters and model architecture parameters of the target sub-network are trained and optimized to obtain the optimized neural network model;

[0140] S3: If there is a next sub-network corresponding to the next network depth of the target sub-network, and the next sub-network is connected to the neural network model, the model parameters of the next sub-network are updated based on the neural network model.

[0141] It can be understood that: the electronic device obtains the target model weight parameters corresponding to each reference search model unit in the neural network model and determines the target search model unit of the reference search model unit in the next sub-network; and then performs inter-layer parameter inheritance processing on the target search model unit based on the target model weight parameters.

[0142] In some embodiments, the reference search model unit can be understood as a unit selected from the neural network model for parameter inheritance of the search model unit in the next sub-network.

[0143] Determination of the reference search model unit: based on the number and type (type refers to RC and NC) of the search model units of the next sub-network. If the number of search model units of the next sub-network is 3, specifically two NCs and one RC, the basis for selecting the reference search model unit is to traverse one by one from the last cell in the trained and optimized neural network model, and select the corresponding number of NCs and RCs as the reference search model units.

[0144] by Figure 7 For example, assuming that the next sub-network is the second sub-network, and the second sub-network includes 2 NCs and 1 RC, the reference search model unit is the two NC-1s and one RC-2 of the first sub-network in the neural network model, that is, the two NC-2s in the next sub-network inherit the target model weight parameters of "the two NC-1s of the first sub-network", and the one RC-2 inherits the target model weight parameters of "the one RC-1 of the first sub-network".

[0145] That is to say, when the second sub-network is trained in the second stage, when the common unit NC-2 and the compression unit RC-2 of the second sub-network are introduced into the neural network model, the target model weight parameter of the common unit NC-1 in the previous stage is used as the initialization value of the common unit NC-2, and the target model weight parameter of the compression unit RC-1 is used as the initialization value of the compression unit RC-2. Similarly, when the third sub-network is trained in the third stage, when the common unit NC-3 of the third sub-network is introduced into the neural network model, the target model weight parameter of the common unit NC-2 is used as the initialization value of the common unit NC-3. Training and optimizing the network on the basis of inter-layer inheritance initialization can ensure the effectiveness of the network and, to a certain extent, solve the problems of large amount of calculation and insufficient training of deep units.

[0146] It should be noted that the inter-layer parameter inheritance process can be understood as inheriting the model weight parameters of the previous stage rather than the model architecture parameters. Inheriting the model weight parameters can enable the neural network model to be quickly trained to a more optimal state. The architecture parameters do not need to be inherited, and different architecture parameters need to be trained for the neural network model at each stage. Furthermore, if the aforementioned "in the process of "connecting the next sub-network to the neural network model", one method can be to directly connect without changing the network parameters of the next sub-network, that is, to retain the initial network parameters of the next sub-network in the connected neural network model (including model weight parameters and model architecture parameters)" is adopted, then the model weight parameters and model architecture parameters of the next sub-network in the neural network model of each stage need to be trained.

[0147] S4: Update the neural network model to the target sub-network, and perform the step of training and optimizing the model architecture parameters of the target sub-network based on the business sample data.

[0148] S5: If there is no next sub-network corresponding to the next network depth of the target sub-network, the neural network model is used as the first neural network model after training and optimization.

[0149] Step S204: generating a second neural network model based on the model parameters corresponding to the first neural network model.

[0150] The number of search model units corresponding to the second neural network model is greater than or equal to the number of search model units corresponding to the first neural network model.

[0151] For details, please refer to step s103, which will not be repeated here.

[0152] In an embodiment of the present application, an electronic device obtains an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type, and then performs architecture search training processing on the initial neural search network based on business sample data to obtain a trained first neural network model, and finally generates a second neural network model based on the model parameters corresponding to the first neural network model. By avoiding using the same type of search model units with the same architecture parameters to build the initial neural search network, the architecture search requirements of different neural network depths can be adapted, and the low efficiency and long time of architecture search caused by a single network structure can be avoided, thereby greatly improving the efficiency of the neural network search architecture and saving network model search time; and, compared with the traditional method of using repeatedly stacked identical units (such as DARTS, FairDARTS and other search spaces are all based on repeatedly stacked identical search units), the flexibility and variability of the network are significantly increased. The flexible model structure enables the model to adapt to the needs of different network depths, significantly improving the overall performance of the model; and, by improving the neural network architecture search method, instead of directly performing architecture search on the entire network, the network architecture search training is performed in segments, saving network architecture search time; and, based on the model parameters of the previous stage, the parameters of the model search unit in the next stage are inherited, which can improve network performance and save time and efficiency in network architecture search training.

[0153] The following are device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0154] See also Fig. 9 , which shows a schematic diagram of the structure of a neural network architecture search device provided by an exemplary embodiment of the present application. The neural network architecture search device can be implemented as all or part of the device through software, hardware or a combination of both. The device 1 includes a network acquisition module 11, a search training module 12 and a model determination module 13.

[0155] A network acquisition module 11 is used to acquire an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type;

[0156] A search training module 12, used to perform architecture search training processing on the initial neural search network based on the business sample data to obtain a trained first neural network model;

[0157] The model determination module 13 is used to generate a second neural network model based on the model parameters corresponding to the first neural network model.

[0158] Optionally, the network acquisition module 11 is specifically used to:

[0159] At least two types of search model units are stacked to construct at least two sub-networks of different network depth types, and an initial neural search network including each of the sub-networks is generated;

[0160] Among them, the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of the same network depth type are the same; and the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of different network depth types are different.

[0161] Optional, such as Fig.10 As shown, the network acquisition module 11 includes:

[0162] A parameter determination unit 111, used to determine three types of search model units with different architecture parameters, wherein the architecture parameters corresponding to the search model units of the same architecture type in each type of search model unit are the same;

[0163] The sub-network construction unit 112 is used to construct a first sub-network based on stacking the first type of search model units among the three types of search model units; to construct a second sub-network based on stacking the second type of search model units among the three types of search model units; and to construct a third sub-network based on stacking the third type of search model units among the three types of search model units;

[0164] The network depth of the second sub-network is greater than the network depth of the first sub-network and less than the network depth of the third sub-network.

[0165] Optionally, the first type of search model unit includes a first type of common unit and a first type of compression unit, the second type of search model unit includes a second type of common unit and a second type of compression unit, and the third type of search model unit includes a third type of common unit and a third type of compression unit;

[0166] The first sub-network is composed of two stacked first-type common units and one stacked first-type compression unit; the second sub-network is composed of two stacked second-type common units and one stacked second-type compression unit; and the third sub-network is composed of two stacked third-type common units.

[0167] Optionally, the search training module 12 is specifically used for:

[0168] Based on the business sample data, the model weight parameters and model architecture parameters of at least two sub-networks included in the initial neural search network are trained and optimized according to the network depth to obtain a first neural network model after training and optimization.

[0169] Optionally, the search training module 12 is specifically used for:

[0170] Determine a current target sub-network from at least two sub-networks included in the initial neural search network according to the network depth;

[0171] Based on the business sample data, the model weight parameters and model architecture parameters of the target sub-network are trained and optimized to obtain the optimized neural network model;

[0172] If there is a next sub-network corresponding to the next network depth of the target sub-network, and the next sub-network is connected to the neural network model, the neural network model is updated to the target sub-network and the steps of training and optimizing the model weight parameters and model architecture parameters of the target sub-network are performed based on the business sample data;

[0173] If there is no next subnetwork corresponding to the next network depth of the target subnetwork, the neural network model is used as the first neural network model after training and optimization.

[0174] Optionally, the search training module 12 is specifically used for:

[0175] Determine a first search model unit corresponding to the tail end of the neural network model and determine a second search model unit corresponding to the head end of the next sub-network;

[0176] Perform unit connection processing on the first search model unit and the second search model unit.

[0177] Optionally, the search training module 12 is specifically used for:

[0178] Based on the neural network model, model parameter update processing is performed on the next sub-network.

[0179] Optionally, the search training module 12 is specifically used for:

[0180] Obtaining target model weight parameters corresponding to each reference search model unit in the neural network model and determining a target search model unit of the reference search model unit in the next sub-network;

[0181] The target search model unit is subjected to inter-layer parameter inheritance processing based on the target model weight parameters.

[0182] Optionally, the search training module 12 is specifically used for:

[0183] The neural network model is updated to the target sub-network, and the step of training and optimizing the model architecture parameters of the target sub-network is performed based on the business sample data.

[0184] Optionally, at least one first search model unit in the initial neural search network is associated with a second search model unit and a third search model unit located before the first search model unit.

[0185] It should be noted that the neural network architecture search device provided in the above embodiment only uses the division of the above functional modules as an example when executing the neural network architecture search method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the neural network architecture search device provided in the above embodiment and the neural network architecture search method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.

[0186] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0187] The present application also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figure 1-Figure 8 The neural network architecture search method of the embodiment shown in the figure can be specifically executed by referring to Figure 1-Figure 8 The specific description of the illustrated embodiment will not be repeated here.

[0188] The present application also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figure 1-Figure 8 The neural network architecture search method of the embodiment shown in the figure can be specifically executed by referring to Figure 1-Figure 8 The specific description of the illustrated embodiment will not be repeated here.

[0189] Please refer to Fig.11 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. The electronic device in the present application may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.

[0190] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions and processes data of the electronic device 100 by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Optionally, the processor 110 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 110 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 110, but may be implemented separately through a communication chip.

[0191] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems. The data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat record data, etc.

[0192] See also Fig.12As shown, the memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good operating results, the operating system allocates corresponding system resources to different third-party applications. However, different application scenarios in the same third-party application also have different requirements for system resources. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and third-party applications are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0193] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0194] Taking the Android operating system as an example, the programs and data stored in the memory 120 are as follows: Fig.13As shown, the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360 and an application layer 380, wherein the Linux kernel layer 320, the system runtime library layer 340 and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides the main feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support, etc. The Android runtime library (Android runtime) is also provided in the system runtime library layer 340, which mainly provides some core libraries that allow developers to use the Java language to write Android applications. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application is running in the application layer 380. These applications can be native applications that come with the operating system, such as contact applications, SMS applications, clock applications, camera applications, etc.; they can also be third-party applications developed by third-party developers, such as game applications, instant messaging applications, photo beautification applications, neural network architecture search applications, etc.

[0195] Taking the operating system as an IOS system as an example, the programs and data stored in the memory 120 are as follows: Fig.14As shown, the IOS system includes: a core operating system layer 420 (Core OS layer), a core service layer 440 (Core Services layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch Layer). The core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks, which provide functions closer to the hardware for use by the program framework located in the core service layer 440. The core service layer 440 provides system services and / or program frameworks required by the application, such as the foundation framework, account framework, advertising framework, data storage framework, network connection framework, geographic location framework, motion framework, etc. The media layer 460 provides audio-visual interfaces for the application, such as graphics and image related interfaces, audio technology related interfaces, video technology related interfaces, and wireless playback (AirPlay) interfaces for audio and video transmission technologies. The touchable layer 480 provides various commonly used interface-related frameworks for application development, and the touchable layer 480 is responsible for the user's touch interaction operations on the electronic device. For example, local notification service, remote push service, advertising framework, game tool framework, message user interface (UI) framework, user interface UIKit framework, map framework, etc.

[0196] exist Fig.14 Among the frameworks shown, the frameworks related to most applications include but are not limited to: the basic framework in the core service layer 440 and the UIKit framework in the touchable layer 480. The basic framework provides many basic object classes and data types, provides the most basic system services for all applications, and has nothing to do with UI. The classes provided by the UIKit framework are basic UI class libraries for creating touch-based user interfaces. iOS applications can provide UIs based on the UIKit framework, so it provides the basic architecture of applications for building user interfaces, drawing, processing and user interaction events, responding to gestures, etc.

[0197] Among them, the method and principle of implementing data communication between third-party applications and the operating system in the IOS system can be referred to the Android system, and this application will not go into details here.

[0198] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays, which are used to receive touch operations on or near the user using any suitable object such as a finger or a touch pen, and to display the user interface of each application. The touch screen display is usually set on the front panel of the electronic device. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of the present application.

[0199] In addition, those skilled in the art will appreciate that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange the components differently. For example, the electronic device also includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module and other components, which will not be described in detail here.

[0200] In the embodiment of the present application, the execution subject of each step may be the electronic device described above. Optionally, the execution subject of each step is the operating system of the electronic device. The operating system may be an Android system, an IOS system, or other operating systems, which is not limited in the embodiment of the present application.

[0201] The electronic device of the embodiment of the present application may also be equipped with a display device, which may be any device capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. The user may use the display device on the electronic device 101 to view displayed text, images, videos and other information. The electronic device may be a smart phone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, an electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, electronic clothing and other devices.

[0202] exist Fig.11 In the electronic device shown, the processor 110 can be used to call the neural network architecture search application stored in the memory 120, and specifically perform the following operations:

[0203] Acquire an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type;

[0204] Performing architecture search training processing on the initial neural search network to obtain a trained first neural network model;

[0205] Generate a second neural network model based on the model architecture parameters corresponding to the first neural network model.

[0206] In one embodiment, when the processor 110 executes the step of obtaining a neural network search space composed of at least two types of search model units, the processor 110 specifically performs the following steps:

[0207] At least two types of search model units are stacked to construct at least two sub-networks of different network depth types, and an initial neural search network including each of the sub-networks is generated;

[0208] Among them, the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of the same network depth type are the same; and the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of different network depth types are different.

[0209] In one embodiment, when the processor 110 executes the stacking of at least two types of search model units to construct at least two sub-networks of different network depth types, the processor 110 specifically performs the following steps:

[0210] Determine three types of search model units with different architecture parameters, wherein the architecture parameters corresponding to the search model units of the same architecture type in each type of search model unit are the same;

[0211] Based on the stacking of the first type of search model units among the three types of search model units, a first sub-network is constructed; based on the stacking of the second type of search model units among the three types of search model units, a second sub-network is constructed; based on the stacking of the third type of search model units among the three types of search model units, a third sub-network is constructed;

[0212] The network depth of the second sub-network is greater than the network depth of the first sub-network and less than the network depth of the third sub-network.

[0213] In one embodiment, the first type of search model unit includes a first type of common unit and a first type of compressed unit, the second type of search model unit includes a second type of common unit and a second type of compressed unit, and the third type of search model unit includes a third type of common unit and a third type of compressed unit;

[0214] The first sub-network is composed of two stacked first-type common units and one stacked first-type compression unit; the second sub-network is composed of two stacked second-type common units and one stacked second-type compression unit; and the third sub-network is composed of two stacked third-type common units.

[0215] In one embodiment, the processor 110 specifically performs the following steps when executing the architecture search training process for the initial neural search network:

[0216] Based on the business sample data, the model weight parameters and model architecture parameters of at least two sub-networks included in the initial neural search network are trained and optimized according to the network depth to obtain a first neural network model after training and optimization.

[0217] In one embodiment, when the processor 110 performs the training and optimization of the model weight parameters and model architecture parameters of at least two sub-networks included in the initial neural search network according to the network depth based on the business sample data to obtain the first neural network model after training and optimization, the processor 110 specifically performs the following steps:

[0218] Determine a current target sub-network from at least two sub-networks included in the initial neural search network according to the network depth;

[0219] Based on the business sample data, the model weight parameters and model architecture parameters of the target sub-network are trained and optimized to obtain the optimized neural network model;

[0220] If there is a next sub-network corresponding to the next network depth of the target sub-network, and the next sub-network is connected to the neural network model, the neural network model is updated to the target sub-network and the steps of training and optimizing the model weight parameters and model architecture parameters of the target sub-network are performed based on the business sample data;

[0221] If there is no next subnetwork corresponding to the next network depth of the target subnetwork, the neural network model is used as the first neural network model after training and optimization.

[0222] In one embodiment, when the processor 110 connects the next sub-network to the neural network model, the processor 110 specifically performs the following steps:

[0223] Determine a first search model unit corresponding to the tail end of the neural network model and determine a second search model unit corresponding to the head end of the next sub-network;

[0224] Perform unit connection processing on the first search model unit and the second search model unit.

[0225] In one embodiment, after executing the step of connecting the next sub-network to the neural network model, the processor 110 further includes:

[0226] Based on the neural network model, model parameter update processing is performed on the next sub-network.

[0227] In one embodiment, the processor 110 performs the model parameter updating process on the next sub-network based on the neural network model, including:

[0228] Obtaining target model weight parameters corresponding to each reference search model unit in the neural network model and determining a target search model unit of the reference search model unit in the next sub-network;

[0229] The target search model unit is subjected to inter-layer parameter inheritance processing based on the target model weight parameters.

[0230] In one embodiment, the processor 110 performs the steps of updating the neural network model to the target sub-network and performing training and optimization on the model weight parameters and model architecture parameters of the target sub-network based on the business sample data, including:

[0231] The neural network model is updated to the target sub-network, and the step of training and optimizing the model architecture parameters of the target sub-network is performed based on the business sample data.

[0232] Those skilled in the art can clearly understand that the technical solution of the present application can be implemented with the help of software and / or hardware. The "unit" and "module" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, where the hardware can be, for example, a field programmable gate array (FPGA), an integrated circuit (IC), etc.

[0233] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0234] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0235] In the several embodiments provided in the present application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0236] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0237] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0238] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk and other media that can store program code.

[0239] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by entering a program to instruct the relevant hardware. The program may be stored in a computer-readable memory, and the memory may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0240] The above is only an exemplary embodiment of the present disclosure, and the scope of the present disclosure cannot be limited thereto. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure here, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A neural network architecture search method, characterized in that: The method comprises: For a model application service, an initial neural search network composed of at least two types of search model units is obtained, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type, and one type of search model units of the same architecture type includes search model units with different architecture parameters, and the search model units are trained based on a data set corresponding to service sample data, and the model application services include image classification services, semantic segmentation services, and speech recognition services; Based on the business sample data, the initial neural search network is subjected to architecture search training processing to obtain a trained first neural network model; For the model application business, generating a second neural network model based on the model parameters corresponding to the first neural network model, wherein the second neural network model is used for calculation processing of the model application business; The obtaining of an initial neural search network composed of at least two types of search model units comprises: At least two types of search model units are stacked to construct at least two sub-networks of different network depth types, and an initial neural search network including each of the sub-networks is generated; Among them, the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of the same network depth type are the same; and the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of different network depth types are different.

2. The method according to claim 1, characterized in that: The stacking of at least two types of search model units to construct at least two sub-networks of different network depth types includes: Determine three types of search model units with different architecture parameters, wherein the architecture parameters corresponding to the search model units of the same architecture type in each type of search model unit are the same; Based on the stacking of the first type of search model units among the three types of search model units, a first sub-network is constructed; based on the stacking of the second type of search model units among the three types of search model units, a second sub-network is constructed; based on the stacking of the third type of search model units among the three types of search model units, a third sub-network is constructed; The network depth of the second sub-network is greater than the network depth of the first sub-network and less than the network depth of the third sub-network.

3. The method according to claim 2, characterized in that The first type of search model unit includes a first type of common unit and a first type of compressed unit, the second type of search model unit includes a second type of common unit and a second type of compressed unit, and the third type of search model unit includes a third type of common unit and a third type of compressed unit; The first sub-network is composed of two stacked first-type common units and one stacked first-type compression unit; the second sub-network is composed of two stacked second-type common units and one stacked second-type compression unit; and the third sub-network is composed of two stacked third-type common units.

4. The method according to any one of claims 1 to 3, wherein the performing architecture search training processing on the initial neural search network comprises: Based on the business sample data, the model weight parameters and model architecture parameters of at least two sub-networks included in the initial neural search network are trained and optimized according to the network depth to obtain a first neural network model after training and optimization.

5. The method according to claim 4, wherein the training and optimization of the model weight parameters and model architecture parameters of at least two sub-networks included in the initial neural search network are performed according to the network depth based on the business sample data to obtain the first neural network model after training and optimization, including: Determine a current target sub-network from at least two sub-networks included in the initial neural search network according to the network depth; Based on the business sample data, the model weight parameters and model architecture parameters of the target sub-network are trained and optimized to obtain the optimized neural network model; If there is a next sub-network corresponding to the next network depth of the target sub-network, and the next sub-network is connected to the neural network model, the neural network model is updated to the target sub-network and the steps of training and optimizing the model weight parameters and model architecture parameters of the target sub-network are performed based on the business sample data; If there is no next subnetwork corresponding to the next network depth of the target subnetwork, the neural network model is used as the first neural network model after training and optimization.

6. The method according to claim 5, wherein connecting the next sub-network to the neural network model comprises: Determine a first search model unit corresponding to the tail end of the neural network model and determine a second search model unit corresponding to the head end of the next sub-network; Perform unit connection processing on the first search model unit and the second search model unit.

7. The method according to claim 5, after connecting the next sub-network to the neural network model, further comprising: Based on the neural network model, model parameter update processing is performed on the next sub-network.

8. The method according to claim 5, wherein the updating of the model parameters of the next sub-network based on the neural network model comprises: Obtaining target model weight parameters corresponding to each reference search model unit in the neural network model and determining a target search model unit of the reference search model unit in the next sub-network; The target search model unit is subjected to inter-layer parameter inheritance processing based on the target model weight parameters.

9. According to the method of claim 8, the step of updating the neural network model to the target sub-network and performing training and optimization on the model weight parameters and model architecture parameters of the target sub-network based on the business sample data comprises: The neural network model is updated to the target sub-network, and the step of training and optimizing the model architecture parameters of the target sub-network is performed based on the business sample data.

10. According to the method according to any one of claims 1-9, at least one first search model unit in the initial neural search network is associated with a second search model unit and a third search model unit located before the first search model unit.

11. A neural network architecture search device, characterized in that: The device comprises: A network acquisition module, for acquiring, for a model application service, an initial neural search network composed of at least two types of search model units, wherein the at least two types of search model units include at least two types of search model units corresponding to different architecture parameters belonging to the same architecture type, and one type of search model units of the same architecture type includes search model units with different architecture parameters, wherein the search model units are trained based on a data set corresponding to service sample data, and the model application service includes an image classification service, a semantic segmentation service, and a speech recognition service; A search training module, used to perform architecture search training processing on the initial neural search network based on business sample data to obtain a trained first neural network model; A model determination module, used to generate a second neural network model based on the model parameters corresponding to the first neural network model, wherein the second neural network model is used for calculation processing of the model application business; The obtaining of an initial neural search network composed of at least two types of search model units comprises: At least two types of search model units are stacked to construct at least two sub-networks of different network depth types, and an initial neural search network including each of the sub-networks is generated; Among them, the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of the same network depth type are the same; and the architecture parameters corresponding to the search model units of the same architecture type in the sub-networks of different network depth types are different.

12. A computer storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps as claimed in any one of claims 1 to 10.

13. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Neural network architecture search method and device, image processing method and device and storage medium

    CN112561027A