System and Method for Machine Learning Evaluation Pipeline

The method optimizes neural architecture search by sequentially using a first search space without training, followed by gradient-based and sampling method search, efficiently discovering high-quality architectures in less time.

JP7702533B2Active Publication Date: 2025-07-03WOVEN BY TOYOTA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024082835
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-07-17
Filing Date
2024-05-21
Publication Date
2025-07-03
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

Existing neural architecture search (NAS) methods are time-consuming and may fail to find optimal architectures, especially when the search space is large, or result in sub-optimal architectures when the search space is small, necessitating a more efficient and high-quality architecture discovery process.

Method used

A method and system that performs NAS by sequentially using a first search space without training, followed by gradient-based search and sampling method search, reducing the number of architectures that need training, and ensuring high-quality architectures are found in minimal time.

Benefits of technology

This approach enables the discovery of high-quality neural network architectures in a significantly reduced time frame by coarsely setting an initial search space and narrowing it down, combining supernet NAS and iterative sampling method search to achieve optimal performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702533000001
    Figure 0007702533000001
  • Figure 0007702533000002
    Figure 0007702533000002
  • Figure 0007702533000003
    Figure 0007702533000003
Patent Text Reader

Abstract

To provide a method, or the like, for a neural architecture search (NAS) pipeline.SOLUTION: Provided are a method, system, and device for a neural architecture search (NAS) pipeline for performing an optimized neural architecture search (NAS). The method may include: obtaining a first search space comprising a plurality of candidate layers for a neural network architecture; performing a training-free NAS in the first search space to obtain a first set of architectures; performing the training-free NAS in the first search space to obtain a second set of architectures; performing a gradient-based search in the second search space to obtain a second set of architectures; performing a sampling method search utilizing the second set of architectures as an initial sample; and obtaining an output architecture as an output of the sampling method search.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Systems and methods according to exemplary embodiments of the present disclosure relate to providing a pipeline for evaluating a machine learning model.

Background Art

[0002] In related art, a neural network can generally be characterized by two main parameters: (1) the architecture of the neural network, and (2) the weights applied to the inputs transmitted between neurons. Typically, the architecture is designed manually, i.e., by hand, by the user, and the weights are optimized by training the network using a training set and a neural network algorithm. Thus, architecture design is an important consideration for optimizing the performance of a neural network, as it is generally static, especially after the neural network has been deployed for use. FIG. 1 is a diagram showing an example of a related art neural network architecture.

[0003] Referring to FIG. 1, an example of a related art neural network architecture is defined by connections, layer types (or operations), the number of layers, and the number of channels in each layer. Layer types can include, for example, convolution (in the case of a Convolutional Neural Network (CNN)), activation (e.g., ReLU (Rectified Linear Unit)), pooling, fully connected, batch normalization, dropout, etc., and the architecture design of the network is embodied by a combination of layers corresponding to at least some of these layer types.

[0004] In the related art, as a technique for automatically designing the architecture of a neural network, there is NAS (Neural Architecture Search). Related art methods for performing NAS may include zero-shot NAS, NAS based on a super network, and simple iterative search.

[0005] In the related art, zero-shot NAS predicts network performance without training network parameters. This method is fast, but may not have good accuracy. For example, this method can be done by evaluating metrics or architectures within the search space without actually performing training.

[0006] FIG. 2 is a diagram showing NAS based on a super network (gradient-based) according to the related art. FIG. 3 shows an example of a super block in NAS based on a super network, and FIG. 4 shows an example of the flow of the NAS method based on the super network according to the related art.

[0007] Gradient-based NAS based on a supernet can use the supernet. A supernet is a network consisting of all candidate architectures. Referring to FIG. 2, supernet-based search explores a neural network architecture through a process of training a plurality of connected superblocks only once, where each superblock corresponds to one layer (and / or set of layers) and includes many choices of candidate layers (and / or sets of layers). For example, Option 1 may be a 3×3 CNN, Option 2 may be a 5×5 CNN, Option 3 may be a max pooling or residual block, and so on. In supernet-based search, the best layer is explored and selected in each superblock. As shown in FIGS. 3 and 4, each candidate has a θ value (or architecture parameter), and that value is updated through the training process. That is, when a certain candidate is determined to have significantly contributed to the reduction of the loss, the θ value of that candidate (representing the importance of the candidate layer) is increased through backpropagation, and vice versa. After the architecture parameter θ is updated in one iteration of supernet-based search, the weight parameters are updated using training images (e.g., images of cars). This process is repeated until a certain degree of network stability is reached (i.e., a certain degree of latency (inference time) or loss is reached), or until a predetermined number of iterations is completed. When the training is completed, the best layer (i.e., the layer with the highest θ value) is selected for each superblock to generate the final (optimal) network architecture.

[0008] FIG. 5 is a functional block diagram of a sampling method according to the related art. Referring to FIG. 5, this sampling method is an iterative process by which a controller generates sample architecture candidates (from a search space consisting of, for example, a set of layers). The sample architecture is optimized through training and evaluated to obtain metric values (such as accuracy, latency, model size, etc.). The controller then generates the next sample architecture candidate using methods such as reinforcement learning, Bayesian optimization, evolutionary algorithms, etc., based on the evaluation results of the previous sample. This process is repeated until the target performance is reached (for example, until the target accuracy, target latency, target mode size, etc. are obtained).

[0009] The sampling method illustrated in FIG. 5 can result in the most accurate / optimal architecture, but due to involving a lot of training, it is very time-consuming to perform over a large number of candidates.

[0010] Performing NAS manually (as is common in related art systems) may ensure a higher-quality architecture in some cases, but it is very time-consuming and burdensome to the user. Therefore, it is desirable to automate NAS.

[0011] However, the above-mentioned existing methods for automating NAS in the related art may be time-consuming and may also fail to obtain the optimal network architecture. In particular, when the search space is large, it takes too much resources and time to complete. On the other hand, if the search space is too small, there is a very high possibility of getting a sub-optimal architecture with low performance. Therefore, there is a need for a method for NAS that can find a network with optimal performance in the shortest time. SUMMARY OF THE INVENTION

[0012] According to an embodiment, a method, a system, and a device of a neural architecture search (NAS) pipeline for performing optimized neural architecture search (NAS) are provided. Specifically, the apparatus and method according to an exemplary embodiment perform NAS without training, supernet / gradient-based search, and sampling method search in sequence, and can ensure the high quality of the architecture provided at the end of the sampling method search while reducing the number of architectures that need to be trained at the end of the sampling method search. Therefore, a high-quality / optimal architecture can be found in a minimum amount of time overall.

[0013] According to an embodiment, a method for performing neural architecture search (NAS) may be provided. The method includes obtaining a first search space including a plurality of candidate layers for a neural network architecture, performing NAS without training within the first search space to obtain a first architecture set, obtaining a second search space based on the first architecture set, performing gradient-based search within the second search space to obtain a second architecture set, performing sampling method search using the second architecture set as an initial sample, and obtaining an output architecture as an output of the sampling method search.

[0014] According to some embodiments, the first search space may be obtained based on a set of architecture parameters. The second search space may be obtained based on a subspace including one or more architectures within the first architecture set.

[0015] According to some embodiments, the sampling method search may be performed repeatedly. The sampling method search may include an evolutionary search algorithm. The number of repetitions of the sampling method search may be based on a predetermined threshold. NAS without training may be performed based on one or more metrics.

[0016] Additional aspects are described in part in the following description, become apparent in part from the description, or can be realized by practicing the presented embodiments of the disclosure.

Brief Description of the Drawings

[0017] The features, advantages, and significance of exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which like reference numerals indicate like elements.

[0018]

Figure 1

[0019]

Figure 2

[0020]

Figure 3

[0021]

Figure 4

[0022]

Figure 5

[0023]

Figure 6

[0024]

Figure 7

Modes for Carrying Out the Invention

[0025] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementation forms to the exact forms disclosed. Modifications and variations are contemplated in light of the above disclosure, or may be obtained from the implementation of the implementation forms. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). In addition, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed (at least partially) simultaneously, and the order of one or more operations may be interchanged.

[0026] It is clear that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation forms. Thus, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code. Therefore, it is understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.

[0027] Certain combinations of features are recited in the claims and / or disclosed herein, but these combinations are not intended to limit the disclosure of possible implementation forms. In fact, many of these features can be combined in ways not specifically recited in the claims and / or specifically disclosed herein. Each of the dependent claims listed below may directly depend on only one claim, but the disclosure of possible implementation forms includes the combination of each dependent claim with every other claim in the claims.

[0028] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly described. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more". When only one item is intended, the term "one" or similar words are used. Also, as used in this specification, the terms "has", "have", "having", "include", "including", etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless explicitly stated otherwise. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0029] Throughout this specification, references to "one embodiment", "an embodiment", "a non-limiting exemplary embodiment", or similar phrases mean that the particular features, structures, or characteristics described in connection with the indicated embodiment are included in at least one embodiment of the present solution. Thus, the phrases "in one embodiment", "in an embodiment", "in one non-limiting exemplary embodiment", and similar phrases throughout this specification may all refer to the same embodiment, but not necessarily so.

[0030] Furthermore, the features, advantages, and characteristics described in this disclosure may be combined in any suitable manner in one or more embodiments. It will be recognized by those skilled in the art, in light of the description herein, that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in a particular embodiment, but they do not exist in all embodiments of the present disclosure.

[0031] Exemplary embodiments of the present disclosure perform NAS without training, supernet / gradient-based search, and sampling method search in sequence, while reducing the number of architectures that need to be trained at the end of the sampling method search, and ensuring the high quality of the architectures provided at the end of the sampling method search, providing a method and system. Thus, a high-quality / optimal architecture as a whole can be found in a minimum amount of time.

[0032] FIG. 6 shows an example of a neural architecture search (NAS) pipeline 600. Referring to FIG. 6, a baseline model and / or a target latency for a desired architecture may be provided to the NAS pipeline 600 (e.g., input by a user). This is useful for providing an initial search space, in particular because an initial search space can be generated based on the baseline model, and architectures that do not match the target latency can also be filtered out using the target latency.

[0033] Referring to FIG. 6, in operation S601, an initial first search space may be obtained (e.g., constructed). The initial first search space may include a plurality of candidate layers for a neural network architecture. According to an embodiment, the initial first search space may be provided by a user or obtained based on a set of architecture parameters.

[0034] Referring to FIG. 6, in step S602, in order to obtain the second search space, the initial first search space may be narrowed down. According to an embodiment, the narrowing down may be performed by first performing NAS without training on the first search space to obtain a first architecture set. In NAS without training, it should be understood that various methods can be used as long as there is no training. For example, a good architecture may be evaluated based on simple metrics. Thus, a first architecture set can be obtained. Since operation S602 does not include training, the completion speed should be relatively fast.

[0035] After that, the second search space may be obtained using the first architecture set. The second search space can also be considered a "reduced" or "shrunk" search space compared to the first search space. According to an embodiment, in order to determine the second search space, a subspace that includes and / or surrounds each architecture in the first architecture set may be used. For example, if the first architecture set includes architectures 1, 2, and 3, subspaces 1, 2, and 3 may correspond to the search spaces (subspaces) that surround architectures 1, 2, and 3, respectively. Thus, in one example, the second search space may be the union of subspaces 1, 2, and 3.

[0036] Referring to FIG. 6 again, in step S603, NAS based on the supernet (gradient-based search) may be performed using the second (reduced) search space. From this method, a second architecture set based on the second search space is obtained. The gradient-based search involves only one training, and moreover, the second search space is already a reduced search space, so it should still be relatively fast.

[0037] Referring back to FIG. 6, in operation S604, a sampling method may be performed using each architecture within the second architecture set as an initial sample (full training search). According to an embodiment, this may be an iterative full training search. Since the second architecture set has already been relatively narrowed down compared to the first architecture set, even if full training is performed, the time is still relatively short. Rather, since full training is performed, the accuracy can be considered optimal / high. Thus, the NAS pipeline 600 can result in obtaining an optimal architecture.

[0038] According to one embodiment, the sampling method search may be an evolutionary search algorithm. The number of repetitions may also be based on a predetermined threshold. Specifically, the evolutionary search is performed repeatedly and may include using a performance metric or fitness / verification score to obtain an architecture optimized based on a second search space. According to one embodiment, in each repetition of the evolutionary search, a candidate architecture is mutated (e.g., changing a 5×5 convolutional layer to a 3×3 convolutional layer) using the entire search space (e.g., any possible layer, layer type, number of channels, number of layers), trained (e.g., for several epochs), and evaluated. Then, the architecture with the lowest score in the candidate pool is replaced with a new (mutated) architecture that is determined to function better. This is repeated over multiple iterations until a predetermined number of iterations are performed or some predetermined threshold (e.g., a predetermined optimal performance) is reached, and an optimized architecture within the second space is output. It should be understood that several possible criteria may be used for the threshold to select the optimized architecture. The simplest method is to select the architecture with the highest score, but if multiple metrics (e.g., accuracy and latency (runtime speed)) are considered, one exemplary process would output all the architectures in the Pareto front set. Another possible method would be to use crossover during the process of the evolutionary search. Specifically, an architecture with a low score can be replaced with one obtained by crossing over two architectures with higher scores.

[0039] It is understood that the above-described embodiment utilizes an evolutionary search, but one or more other embodiments are not limited thereto. In another embodiment, another sampling method or algorithm may be utilized (e.g., reinforcement learning).

[0040] FIG. 7 shows a flowchart of an exemplary method 700 for performing an optimized NAS process. Referring to FIG. 7, in operation S710, a first search space having candidate layers for a neural network architecture is obtained. This may be similar to operation S601 described above with respect to FIG. 6.

[0041] In operation S720, NAS without training may be performed on the first search space to obtain a first set of architectures. This may be similar to multiple parts of operation S602 described above with respect to FIG. 6.

[0042] In operation S730, a second (reduced) search space may be obtained based on the first set of architectures. According to some embodiments, this may be obtained based on a subspace that includes and / or surrounds each architecture within the first set of architectures to determine the second search space. This may be similar to multiple parts of operation S602 described above with respect to FIG. 6.

[0043] In operation S740, a (gradient-based) search based on the supernet may be performed in the second search space obtained in operation S730 to output a second set of architectures. This may be similar to multiple parts of the portion of operation S603 described above with respect to FIG. 6.

[0044] In operation S750, a sampling method may be performed using the second set of architectures obtained from operation S740 as an initial sample. According to one embodiment, the sampling method may be an evolutionary search algorithm. The evolutionary search algorithm may be repeated over several iterations, and the number of these iterations may be based on a predetermined threshold. This may be similar to operation S604 described above with respect to FIG. 6.

[0045] In operation S760, an optimal architecture is output as a result of the sampling method performed in operation S750. Although this embodiment has been described as outputting one optimal architecture, it should be understood that some embodiments may output multiple optimal architectures.

[0046] In view of the above, exemplary embodiments of the present disclosure provide a method and system for performing neural architecture search (NAS) optimized by coarsely setting an initial search space and then narrowing down the search space so as to substantially speed up the search, and by using the narrowed-down search space in combination with supernet NAS and iterative sampling method search, an architecture having optimal performance can be obtained. Therefore, by sequentially combining and using these search methods, an architecture with optimal performance can be found in substantially shorter search time and calculation time.

[0047] It should be understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed herein is an illustration of an example of the technique. Based on design preferences, it should be understood that the specific order or hierarchy of blocks in a process / flowchart can be rearranged. Further, some blocks may be combined or omitted. The appended method claims present the elements of the various blocks as an example of an order and are not intended to be limited to the specific order or hierarchy presented.

[0048] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible technical detail level. Further, one or more of the above-described components may be implemented as instructions executable by at least one processor stored on a computer-readable medium (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium (or multiple media) having computer-readable program instructions for causing a processor to perform operations.

[0049] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, devices such as mechanically encoded punch cards, or raised structures in grooves in which instructions are recorded, and any suitable combination thereof. A computer-readable storage medium, as used herein, should not be construed to be a transient signal such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0050] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network such as, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0051] The computer-readable program code / instructions for performing the operations may be in the form of assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or either source code or object code, and the source code or object code may be written in any combination of one or more programming languages including object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer as a stand-alone software package, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by personalizing the electronic circuit using the state information of the computer-readable program instructions to perform the aspects or operations.

[0052] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to create a machine that causes the instructions executed via the processor of the computer or other programmable data processing apparatus to implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium storing the instructions comprises an article of manufacture including instructions implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0053] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0054] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function(s). The methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in the figures. In some alternative implementations, the functions described in the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may actually be performed simultaneously or substantially simultaneously, or the blocks may sometimes be performed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.

[0055] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting of the implementations. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.

Claims

1. A method for performing neural architecture search (NAS), comprising: obtaining a first search space including a plurality of candidate layers for a neural network architecture; performing NAS without training within the first search space to obtain a first architecture set; obtaining a second search space based on the first architecture set; performing gradient-based search within the second search space to obtain a second architecture set; performing sampling method search using the second architecture set as an initial sample; obtaining an output architecture as an output of the sampling method search; wherein the sampling method search is performed iteratively; the sampling method search includes an evolutionary search algorithm; each iteration of the evolutionary search algorithm of the sampling method search includes mutating, training, and evaluating a candidate architecture, and replacing the architecture with the lowest score in the candidate pool with a new candidate architecture.

2. The method according to claim 1, wherein the first search space is obtained based on a set of architecture parameters.

3. The method according to claim 1 or 2, wherein the second search space is obtained based on a subspace including one or more architectures within the first architecture set.

4. The method according to claim 1, wherein the number of iterations of the sampling method search is based on a predetermined threshold.

5. The method according to claim 1 or 2, wherein the NAS without training is performed based on one or more metrics.

6. An apparatus for performing neural architecture search (NAS), comprising: at least one memory storing computer-executable instructions; and at least one processor, wherein the processor executes the computer-executable instructions to obtain a first search space including a plurality of candidate layers for a neural network architecture; perform NAS without training within the first search space to obtain a first architecture set; obtain a second search space based on the first architecture set; perform gradient-based search within the second search space to obtain a second architecture set; Performing sampling method search using the second architecture set as an initial sample, configured to obtain an output architecture as an output of the sampling method search, the sampling method search is performed repeatedly, the sampling method search includes an evolutionary search algorithm, each iteration of the evolutionary search algorithm of the sampling method search includes mutating, training, and evaluating a candidate architecture, and replacing the architecture with the lowest score in the candidate pool with a new candidate architecture, the device.

7. The device according to claim 6, wherein the first search space is obtained based on a set of architecture parameters.

8. The device according to claim 6 or 7, wherein the second search space is obtained based on a subspace including one or more architectures in the first architecture set.

9. The device according to claim 6, wherein the number of repetitions of the sampling method search is based on a predetermined threshold.

10. The device according to claim 6 or 7, wherein the NAS without training is performed based on one or more metrics.

11. A non-transitory computer-readable recording medium having instructions executable by at least one processor, the instructions causing the at least one processor to, obtain a second search space based on a first architecture set, perform gradient-based search within the second search space to obtain a second architecture set, perform sampling method search using the second architecture set as an initial sample, obtain an output architecture as an output of the sampling method search, including, causing to execute a method, the sampling method search is performed repeatedly, the sampling method search includes an evolutionary search algorithm, each iteration of the evolutionary search algorithm of the sampling method search includes mutating, training, and evaluating a candidate architecture, and replacing the architecture with the lowest score in the candidate pool with a new candidate architecture, the non-transitory computer-readable recording medium.

12. The non-transitory computer-readable recording medium according to claim 11, wherein the first search space is obtained based on a set of architecture parameters.

13. The non-transitory computer-readable recording medium according to claim 11 or 12, wherein the second search space is obtained based on a subspace including one or more architectures within the first architecture set.

14. The non-transitory computer-readable recording medium according to claim 11, wherein the number of repetitions of the sampling method search is based on a predetermined threshold value.

Citation Information

Patent Citations

  • Deep learning network determination method, image classification method and equipment

    CN116051964A

  • Weak neural architecture search (NAS) predictor

    US20220188599A1

  • Design space reduction apparatus, control method, and computer-readable storage medium

    WO2022137393A1