Method and system for obtaining a performance prediction model of a neural network architecture
By dividing the channel sparsity sampling interval of the neural network into multiple sub-sampling intervals, the reconstructed neural network architecture is obtained, which solves the problem of low efficiency in neural network architecture search and achieves more efficient architecture evaluation and ranking accuracy.
Patent Information
- Application Number
- CN202310569947.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-19
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-05-19
AI Technical Summary
Existing technologies have low search efficiency in neural network architecture search and struggle to maintain accuracy in architecture evaluation and ranking, mainly due to the insufficient generalization ability of neural network performance prediction models.
By dividing the sampling interval of the channel sparsity of the neural network into multiple sub-sampling intervals, and performing sparsity sampling in these intervals, a reconstructed neural network architecture can be obtained, so as to obtain a smaller or larger neural network architecture and improve the generalization ability of the performance prediction model.
It accelerates the search process for neural network architectures, improves the accuracy of architecture evaluation and the reliability of ranking, and enhances the generalization ability of performance prediction models.
Smart Images

Figure CN116562340B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of neural network architecture search, and in particular to a method and system for obtaining a performance prediction model of a neural network architecture. BACKGROUND
[0002] The real world is full of various random events. In probability theory and statistics, a random variable is used to represent each possible result after a random event occurs. Random variables are divided into discrete type (a finite number of results, such as the heads and tails of a coin, and the winning numbers of a lottery) and continuous type (an infinite number of results, such as the temperature tomorrow and the life of an electronic component).
[0003] The statistical law of a random variable can be described by a probability distribution function. Common discrete probability distributions include Bernoulli distribution, binomial distribution, geometric distribution, and Poisson distribution. Common continuous distributions include uniform distribution, Gaussian distribution, and power-law distribution.
[0004] Random sampling refers to sampling according to a certain probability distribution. For example, flipping a coin is sampling according to a Bernoulli distribution, and each flip of the coin produces a sample (“heads” or “tails”). The height of a person in a crowd follows a Gaussian distribution, and each time a person is randomly selected from the crowd to measure their height, it is a sampling. Random number generation in a computer is random sampling in a uniform distribution.
[0005] Random sampling is very valuable in applications. For example, the field of neural network architecture search (NAS) has attracted attention in academia and industry in recent years. NAS refers to using algorithms to automatically search for a locally optimal network architecture, which can achieve better results than manual design.
[0006] The problem with NAS at present is that the search efficiency is low. To solve this problem, the architecture evaluation can be accelerated and the accuracy in the sense of architecture ranking can be ensured, or the search can be performed in architectures with more potential, reducing the number of architectures that need to be evaluated, thereby accelerating the search process.
[0007] In the case of accelerating architecture evaluation and ensuring accuracy in the sense of architecture ranking to improve the search efficiency of neural network architecture, there are various methods for accelerating architecture evaluation. For example, a performance prediction model of a neural network architecture can be trained, that is, the neural network architecture is encoded and input into the prediction model, and the predicted value of its performance (such as precision or latency) is output.
[0008] However, to train the prediction model, a sufficient number of samples need to be collected to form a training set and a test set required for training. In this case, a sample includes a neural network architecture and the actual performance of the architecture. Therefore, in order to make the trained prediction model have stronger generalization ability, the actual performance distribution of the collected samples needs to be wide.
[0009] The background description is provided only for the purpose of understanding the relevant technology, and is not considered as an admission that the prior art is known. SUMMARY
[0010] Therefore, the embodiments of the present application aim to provide a method and system for obtaining a performance prediction model of a neural network architecture, which can construct a larger or smaller neural network architecture to improve the generalization ability of the neural network performance prediction model, and thus can accelerate architecture evaluation and ensure accuracy in the sense of architecture ranking, to improve the search efficiency of the neural network architecture.
[0011] In a first aspect, the embodiments of the present application provide a method for obtaining a performance prediction model of a neural network architecture, comprising: obtaining an original architecture of a neural network. According to the original architecture of the neural network, obtaining an original channel number of the neural network; the original channel number of the neural network is the original number of convolution kernels of the 2D convolution layer of the original architecture of the neural network. According to the sampling interval of the sparsity of the original channel number, obtaining a first sub-sampling interval, wherein the first sub-sampling interval is any interval in the sampling interval of the sparsity of the original channel number, and the first sub-sampling interval is truly contained in the sampling interval of the sparsity of the original channel number. According to the first sub-sampling interval, obtaining the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network, wherein the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network is obtained by sampling in the first sub-sampling interval. According to the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network, obtaining an updated value of the channel number of the 2D convolution layer in the reconstructed neural network. According to the updated value of the channel number of the 2D convolution layer in the reconstructed neural network, obtaining the reconstructed neural network. And according to the reconstructed neural network, obtaining the performance prediction model of the neural network architecture.
[0012] In the embodiments of the present application, in the second aspect, the embodiments of the present application provide a system for obtaining a performance prediction model of a neural network architecture, the system comprising a data acquisition module and a data processing module. The data acquisition module is configured to: acquire an original architecture of a neural network, and acquire an original channel number of the neural network according to the original architecture of the neural network; the original channel number being a number of convolution kernels of a 2D convolution layer in the original architecture of the neural network. The data processing module is connected to the data acquisition module and is configured to: acquire a first sub-sampling interval according to a sampling interval of a sparsity of the original channel number, wherein the first sub-sampling interval is any interval in the sampling interval of the sparsity of the original channel number, and the first sub-sampling interval is truly contained in the sampling interval of the sparsity of the original channel number; acquire a sparsity of a channel number of a 2D convolution layer in a reconstructed neural network according to the first sub-sampling interval, wherein the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network is obtained by sampling in the first sub-sampling interval; acquire an updated value of the channel number of the 2D convolution layer in the reconstructed neural network according to the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network; acquire the reconstructed neural network according to the updated value of the channel number of the 2D convolution layer in the reconstructed neural network; and acquire a performance prediction model of the neural network architecture according to the reconstructed neural network.
[0013] In the method for obtaining the performance prediction model of the neural network architecture used in the embodiments of the present application, the sampling interval of the sparsity of the original channel number is divided into a plurality of sub-sampling intervals, and one interval is selected from the plurality of sub-sampling intervals for sampling of the sparsity, so that a smaller or larger neural network architecture can be obtained, so that the actual performance distribution of the obtained neural network architecture is in a wider range, the generalization ability of the neural network performance prediction model can be improved, and then the architecture evaluation can be accelerated and the accuracy in the sense of architecture ranking can be ensured, so as to improve the search efficiency of the neural network architecture.
[0014] Some of the other optional features and technical effects of the embodiments of the present application are described below, and some can be understood by reading this document. BRIEF DESCRIPTION OF DRAWINGS
[0015] In the following, the embodiments of the present application will be described in detail with reference to the accompanying drawings, wherein the elements shown are not limited by the proportions shown in the drawings, and the same or similar reference numerals in the drawings represent the same or similar elements, wherein:
[0016] Figure 1 An exemplary flowchart of a method for obtaining a performance prediction model of a neural network architecture according to an embodiment of the present application is shown;
[0017] Figure 2 An exemplary flowchart of another method for obtaining a performance prediction model of a neural network architecture according to an embodiment of the present application is shown;
[0018] Figure 3A A result diagram showing the latency performance distribution of samples of the performance prediction model of the neural network architecture acquired according to the related art is shown;
[0019] Figure 3B A result diagram showing the latency performance distribution of samples of the performance prediction model of the neural network architecture acquired according to an embodiment of the present application is shown;
[0020] Figure 4 An exemplary structural diagram of a performance prediction model system of a neural network architecture acquired according to an embodiment of the present application is shown;
[0021] Figure 5 An exemplary structural diagram of an electronic device capable of implementing the method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions, and advantages of the present application clearer, further detailed description will be given to the present application in combination with specific embodiments and drawings. Herein, the exemplary embodiments of the present application and the description thereof are used to explain the present application, but not as a limitation to the present application.
[0023] The term "comprising" and variations thereof as used in the present document are used to mean "including, but not limited to". The term "or" as used in the present document is used to mean "and / or". The term "based on" is used to mean "based, at least in part, on". The terms "one example embodiment" and "an embodiment" are used to mean "at least one example embodiment". The term "another embodiment" is used to mean "at least one additional embodiment". The terms "first", "second", etc. can refer to different or same objects. Other explicit or implicit definitions can also be included below.
[0024] In order to make the trained prediction model have stronger generalization ability, it is necessary to make the actual performance distribution of the collected samples wider. In the related art, a sampling interval needs to be specified first, and then the sampling of the neural network architecture is performed. The search interval can be artificially designed. Taking a convolutional neural network as an example, the size, step, and number of each layer of convolution kernel, as well as the connection relationship between layers (multi-branch, skip connection, etc.) all belong to the degrees of freedom of the search. The following takes the number of convolution kernels, i.e., the number of channels, as an example to illustrate the degrees of freedom of the search.
[0025] Generally, the sampling interval of the number of channels is designed as a discrete interval, for example, the number of convolution kernels (i.e., the number of channels) of a certain layer is randomly selected from a sequence of finite integers (for example, the sequence {4, 8, 12, 16, 32, 64}). If only considering the number of channels as one degree of freedom, the sampling of the number of channels for each layer of convolution can obtain a neural network architecture.
[0026] However, the related art uses a uniform sampling method when sampling the number of channels, so the joint probability distribution of the number of channels of all convolution layers is likely to be concentrated in the middle area (for example, neural network A includes 10 convolution layers, for the discretized sequence of the number of channels {4, 8, 12, 16, 20, 24}, we consider that for the subnetwork architecture B sampled according to neural network A, in the case that the number of channels of each layer of convolution of subnetwork architecture B is 4, subnetwork architecture B is the smallest neural network architecture collected according to neural network A, and in the case that the number of channels of each layer of convolution of subnetwork architecture B is 24, subnetwork architecture B is the largest neural network architecture collected according to neural network A. The probability of sampling the smallest neural network architecture or the largest neural network architecture according to neural network A is It can be seen that it is difficult to sample the smallest neural network architecture or the largest neural network architecture according to neural network A, so smaller or larger neural network architectures cannot be sampled, thereby causing the generalization ability of the neural network performance prediction model to be low.
[0027] It can be understood that for the discretized sequence of the number of channels of the convolution layer {4, 8, 12, 16, 20, 24}, if the probability of sampling the number of channels of the convolution layer is greater for 4 or 8, a smaller neural network architecture can be obtained, and if the probability of sampling the number of channels of the convolution layer is greater for 20 or 24, a larger neural network architecture can be obtained.
[0028] Therefore, an embodiment of the present application provides a method for obtaining a performance prediction model of a neural network architecture, by dividing the sampling interval of the sparsity of the number of channels into a plurality of sub-sampling intervals, and selecting an interval in the plurality of sub-sampling intervals to sample the sparsity of the number of channels, so that the number of channels is concentrated in a smaller interval of the discretized sequence, thereby obtaining a smaller or larger neural network architecture.
[0029] As shown in FIG. 1, Figure 1 An embodiment of the present application provides a method for obtaining a performance prediction model of a neural network architecture, which includes steps 101 to 107.
[0030] Step 101, obtaining an original architecture of a neural network.
[0031] Step 102, obtaining an original number of channels of the neural network according to the original architecture of the neural network.
[0032] It can be understood that the original number of channels of the neural network is the original number of convolution kernels of the 2D convolution layer of the original architecture of the neural network.
[0033] Step 103, obtaining a first sub-sampling interval according to the sampling interval of the sparsity of the original number of channels.
[0034] It can be understood that the first sub-sampling interval is any one of the sampling intervals of the sparsity of the original channel number, and the first sub-sampling interval is truly contained in the sampling interval of the sparsity of the original channel number. Then, the sparsity is collected in the first sub-sampling interval, and the channel number of the 2D convolution layer is obtained according to the collected sparsity, so that the channel number of the 2D convolution layer is distributed in a more concentrated range, and a smaller or larger neural network architecture is more easily obtained. For example, in the case where the sampling interval of the sparsity of the original channel number is [0, 1], if the first sub-sampling interval is [0, 0.1], a larger neural network architecture can be obtained, and if the first sub-sampling interval is [0.9, 1], a smaller neural network architecture can be obtained.
[0035] In some embodiments, the implementation method of step 103 can include: obtaining a plurality of sub-sampling intervals, and the first sub-sampling interval is any one of the plurality of sub-sampling intervals. For example, the plurality of sub-sampling intervals are obtained by sampling the sampling interval of the sparsity of the original channel number.
[0036] In some embodiments, the plurality of sub-sampling intervals are obtained by uniformly sampling or Gaussian sampling the sampling interval of the sparsity of the original channel number. It can be understood that uniformly sampling or Gaussian sampling the sampling interval of the sparsity of the original channel number can obtain a plurality of sub-sampling intervals, and uniformly sampling or Gaussian sampling in the plurality of sub-sampling intervals can select one sampling interval for sampling, and then sampling the sparsity in the sampling interval, so as to increase the probability of obtaining a smaller or larger neural network architecture.
[0037] In some embodiments, the first sub-sampling interval is preset or obtained by randomly sampling the plurality of sub-sampling intervals. It can be understood that the first sub-sampling interval can be artificially preset to avoid excessive concentration in a certain area within the sampling interval of the sparsity of the original channel number, and to facilitate obtaining a smaller or larger neural network architecture.
[0038] Step 104: obtaining the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network according to the first sub-sampling interval.
[0039] It can be understood that the sparsity of the 2D convolution layer in the original architecture of the neural network is obtained by sampling in the first sub-sampling interval.
[0040] Step 105: obtaining an updated value of the channel number of the 2D convolution layer in the reconstructed neural network according to the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network.
[0041] Step 106, obtaining the reconstructed neural network according to the updated value of the channel number of the 2D convolution layer in the reconstructed neural network.
[0042] The following embodiments are exemplarily described taking the original architecture of the neural network yolov5 as an example. The network structure of yolov5 is shown in Table 1. In Table 1, conv2d is a 2D convolution layer, C3 includes multiple 2D convolutions and skip connections, SPPF includes 2D convolution and pooling, Upsample+Concat is up-sampling and stacking, and Concat is stacking.
[0043] Table 1 Network structure of yolov5
[0044] Type Number of channels conv2d 48 conv2d 96 C3 96 conv2d 192 C3 192 conv2d 384 C3 384 conv2d 768 C3 768 SPPF 768 conv2d 384 Upsample + Concat - C3 384 conv2d 192 Upsample + Concat - C3 192 conv2d 192 Concat 384 C3 384 conv2d 384 Concat 768 C3 768
[0045] According to Table 1, the original channel number of the network structure of yolov5 can be obtained.
[0046] It can be understood that the sampling interval of the sparsity of the original channel number can be [0, 1]. According to steps 103 to 106, a random seed seed can be set (for example, a number is randomly selected from [0, 1, 2, 3, 4, 5, 6, 7, 8] as a random seed seed). For each 2D convolution layer, a floating-point number is randomly selected from the range of (seed*0.1, seed*0.1+0.1) as the sparsity of the 2D convolution layer. According to the sparsity of the 2D convolution layer, the updated value of the channel number of the 2D convolution layer in the reconstructed neural network can be calculated. For example, the original channel number of a certain 2D convolution layer is 98, and if the sparsity of the certain 2D convolution layer is 0.1, then the updated value of the channel number of the certain 2D convolution layer is: 98*(1-0.1) = 88. The same operation is performed for each 2D convolution layer, and the updated value of the channel number of each 2D convolution layer in the reconstructed neural network can be obtained.
[0047] It can be understood that, since there can be connections or addition operations between different convolution layers, there is a mutual dependence relationship between their channel numbers. After obtaining the updated value of the channel number, fine tuning can be performed according to the actual situation to form a correctly reconstructed neural network architecture.
[0048] Step 107, obtaining a performance prediction model of the neural network architecture according to the reconstructed neural network.
[0049] In some embodiments, step 107 can include steps 201 to 203.
[0050] In step 201, a plurality of reconstructed neural networks is obtained, and actual performance of each of the plurality of reconstructed neural networks is obtained according to the plurality of reconstructed neural networks.
[0051] In step 201, the actual performance of each of the plurality of reconstructed neural networks can include accuracy of each of the plurality of reconstructed neural networks, or / and latency performance of each of the plurality of reconstructed neural networks. For example, for a network structure of a neural network original architecture yolov5, the actual performance includes latency performance.
[0052] Exemplarily, the performance of each of the plurality of reconstructed neural networks can be measured to obtain the actual performance of each of the plurality of reconstructed neural networks. For a network structure of a neural network original architecture yolov5.
[0053] In step 202, a plurality of sample pairs of a performance prediction model of a neural network architecture is obtained according to the plurality of reconstructed neural networks and the actual performance of each of the plurality of reconstructed neural networks.
[0054] In step 203, the performance prediction model of the neural network architecture is obtained according to the plurality of sample pairs of the performance prediction model of the neural network architecture.
[0055] According to the related art, a result graph of a latency performance distribution of a plurality of neural network architectures obtained by a uniform sampling method is as shown in FIG. 1. Figure 3A According to step 201, the latency performance of each of the plurality of reconstructed neural networks obtained is as shown in FIG. 2. Figure 3B Figure 3A Compared with FIG. 1, the latency performance of each of the plurality of reconstructed neural networks obtained according to step 201 is distributed in a wider range. Figure 3B In FIG. 1, the abscissa represents latency, and the ordinate represents probability density. It can be known that the latency performance distribution of each of the plurality of reconstructed neural networks obtained according to step 201 is distributed in a wider range, and the performance prediction model of the neural network architecture obtained according to step 203 has stronger generalization ability.
[0056] The above embodiment divides the sampling interval of the sparsity into a plurality of sub-sampling intervals, selects an interval in the plurality of sub-sampling intervals for sampling of the sparsity, and updates the degree of freedom, so that a smaller or larger neural network architecture can be obtained, the actual performance of the plurality of reconstructed neural networks is distributed in a wider range, and the performance prediction model of the neural network architecture has stronger generalization ability, so as to accelerate the architecture evaluation and ensure the accuracy in the sense of architecture ranking, and improve the search efficiency of the neural network architecture.
[0057] Embodiments of the present application also provide a system for obtaining a performance prediction model of a neural network architecture, as shown in FIG. 3. Figure 4 As shown, the system comprises a data acquisition module 401 and a data processing module 402.
[0058] The data acquisition module 401 is configured to acquire an original architecture of a neural network, and acquire an original channel number of the neural network according to the original architecture of the neural network. The original channel number is the number of convolution kernels of a 2D convolution layer in the original architecture of the neural network.
[0059] The data processing module 402 is connected with the data acquisition module 401, and is configured to: acquire a first sub-sampling interval according to a sampling interval of a sparsity of the original channel number; acquire a sparsity of a channel number of a 2D convolution layer in a reconstructed neural network according to the first sub-sampling interval; acquire an updated value of the channel number of the 2D convolution layer in the reconstructed neural network according to the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network; acquire a performance prediction model of the neural network architecture according to the updated value of the channel number of the 2D convolution layer in the reconstructed neural network, and according to the reconstructed neural network. The first sub-sampling interval is any interval in the sampling interval of the sparsity of the original channel number, the first sub-sampling interval is truly contained in the sampling interval of the sparsity of the original channel number, and the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network is obtained by sampling in the first sub-sampling interval.
[0060] In some embodiments, the data processing module 402 is further configured to acquire a plurality of sub-sampling intervals. The plurality of sub-sampling intervals are obtained by sampling the sampling interval of the sparsity. The first sub-sampling interval is any sub-sampling interval in the plurality of sub-sampling intervals.
[0061] In some embodiments, the plurality of sub-sampling intervals are obtained by uniformly sampling or Gaussian sampling the sampling interval of the sparsity of the original channel number.
[0062] In some embodiments, the first sub-sampling interval is preset, or is obtained by randomly sampling the plurality of sub-sampling intervals.
[0063] In some embodiments, the data processing module 402 is configured to: acquire a plurality of reconstructed neural networks; acquire an actual performance of each of the plurality of reconstructed neural networks according to the plurality of reconstructed neural networks; acquire a plurality of sample pairs of the performance prediction model of the neural network architecture according to the plurality of reconstructed neural networks and the actual performance of each of the plurality of reconstructed neural networks; and acquire the performance prediction model of the neural network architecture according to the plurality of sample pairs of the performance prediction model of the neural network architecture. The actual performance of each of the plurality of reconstructed neural networks includes the accuracy of each of the plurality of reconstructed neural networks, or / and the latency performance of each of the plurality of reconstructed neural networks.
[0064] The beneficial effects of the system for obtaining the performance prediction model of the neural network architecture described in the above embodiments can refer to the related descriptions of the embodiments of the method for obtaining the performance prediction model of the neural network architecture, which will not be repeated here
[0065] In some embodiments, the system for obtaining the performance prediction model of the neural network architecture can be combined with the method for obtaining the performance prediction model of the neural network architecture of any embodiment, and vice versa, which will not be repeated here.
[0066] In the embodiments of the present application, an electronic device is provided, comprising a processor and a memory storing a computer program, the processor is configured to execute the method for obtaining the performance prediction model of the neural network architecture of any embodiment of the present application when running the computer program.
[0067] Figure 5 A schematic diagram of an electronic device 1000 that can implement the method or realize the embodiments of the present application is shown, which can include more or less electronic devices than shown in some embodiments. In some embodiments, it can be implemented with a single or multiple electronic devices. In some embodiments, it can be implemented with a cloud or distributed electronic device.
[0068] As shown in Figure 5 The electronic device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to programs and / or data stored in a read-only memory (ROM) 1002 or loaded from a storage portion 1008 into a random access memory (RAM) 1003. The processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 1001 can include a general-purpose main processor and one or more special-purpose coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), etc. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0069] The above processor and memory are used together to execute programs stored in the memory, which can realize the methods, steps or functions described in the above embodiments when executed by the computer.
[0070] The following components are connected to the I / O interface 1005: an input part 1006 including a keyboard, a mouse, a touch screen, and the like; an output part 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage part 1008 including a hard disk, and the like; and a communication part 1009 including a network interface card such as a LAN card, a modem, and the like. The communication part 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as necessary. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 1010 as necessary, so that a computer program read out therefrom is installed in the storage part 1008 as necessary. Figure 5 Only some of the components of the computer system 1000 are shown, and are not meant to imply architectural limitations of the computer system 1000. Figure 5 The components shown are merely exemplary.
[0071] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by a computer or its associated components. The computer may, for example, be a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart television, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.
[0072] Although not shown, in an embodiment of the present application, a storage medium is provided, which stores a computer program configured to be executed to perform the file difference based compiling method of any embodiment of the present application.
[0073] The storage medium of an embodiment of the present application includes a permanent and non-permanent, removable and non-removable article that can store information by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device.
[0074] The methods, programs, systems, apparatuses, etc., in embodiments of the present invention can be executed or implemented in one or more networked computers, or practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be performed by remote processing devices connected via a communication network.
[0075] Those skilled in the art will understand that the embodiments described in this specification can be provided as methods, systems, or computer program products. Therefore, those skilled in the art will realize that the functional modules / units or controllers and related method steps described in the above embodiments can be implemented in software, hardware, or a combination of both.
[0076] Unless explicitly stated otherwise, the actions or steps of the methods and procedures described in the embodiments of the present invention do not necessarily have to be performed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0077] This document describes several embodiments of the present invention; however, for the sake of brevity, the descriptions of the embodiments are not exhaustive, and identical or similar features or parts between the embodiments may be omitted. In this document, "one embodiment," "some embodiments," "example," "specific example," or "some examples" refers to embodiments applicable to at least one, but not all, of the present invention. The above terms do not necessarily refer to the same embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of the different embodiments or examples.
[0078] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above embodiments, which are merely examples of the best mode for implementing the systems and methods. Those skilled in the art will understand that various changes can be made to the embodiments of the systems and methods described herein without departing from the spirit and scope of the invention as defined in the appended claims when implementing the systems and / or methods.
Claims
1. A method of obtaining a performance prediction model for a neural network architecture, characterized by, The method comprises: obtaining an original architecture of the neural network; obtaining an original channel number of the neural network according to the original architecture of the neural network, wherein the original channel number of the neural network is an original number of convolution kernels of a 2D convolution layer of the original architecture of the neural network; obtaining a first sub-sampling interval according to a sampling interval of a sparsity of the original channel number, wherein the first sub-sampling interval is any interval in the sampling interval of the sparsity of the original channel number, and the first sub-sampling interval is truly contained in the sampling interval of the sparsity of the original channel number; obtaining a sparsity of a channel number of a 2D convolution layer in a reconstructed neural network according to the first sub-sampling interval, wherein the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network is obtained by sampling in the first sub-sampling interval; obtaining an updated value of the channel number of the 2D convolution layer in the reconstructed neural network according to the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network; and obtaining the reconstructed neural network according to the updated value of the channel number of the 2D convolution layer in the reconstructed neural network; and obtaining a performance prediction model of a neural network architecture according to the reconstructed neural network. The obtaining of the first sub-sampling interval according to the sampling interval of the sparsity of the original channel number comprises:
2. The method of claim 1, wherein, obtaining a plurality of sub-sampling intervals, wherein the plurality of sub-sampling intervals are obtained by sampling the sampling interval of the sparsity of the original channel number, and the first sub-sampling interval is any sub-sampling interval in the plurality of sub-sampling intervals. The plurality of sub-sampling intervals are obtained by uniformly sampling or Gaussian sampling the sampling interval of the sparsity of the original channel number.
3. The method of claim 2, wherein, The first sub-sampling interval is preset or obtained by randomly sampling the plurality of sub-sampling intervals.
4. The method of claim 2, wherein, The obtaining of the performance prediction model of the neural network architecture according to the reconstructed neural network comprises:
5. The method of obtaining a performance prediction model of a neural network architecture according to any one of claims 1-4, wherein, obtaining a plurality of reconstructed neural networks, and obtaining actual performances of each of the plurality of reconstructed neural networks according to the plurality of reconstructed neural networks, wherein the actual performance of each of the plurality of reconstructed neural networks comprises an accuracy of the each of the plurality of reconstructed neural networks or / and a latency performance of the each of the plurality of reconstructed neural networks; obtaining a plurality of sample pairs of the performance prediction model of the neural network architecture according to the plurality of reconstructed neural networks and the actual performances of each of the plurality of reconstructed neural networks; and obtaining the performance prediction model of the neural network architecture according to the plurality of sample pairs of the performance prediction model of the neural network architecture. The method comprises:
6. A system for obtaining a performance prediction model for a neural network architecture, characterized in that, a data obtaining module configured to: obtain an original architecture of the neural network, and obtain an original channel number of the neural network according to the original architecture of the neural network, wherein the original channel number is a number of convolution kernels of a 2D convolution layer in the original architecture of the neural network; and a data processing module connected with the data obtaining module and configured to: According to the sampling interval of the sparsity of the original channel number, a first sub-sampling interval is obtained, wherein the first sub-sampling interval is any interval in the sampling interval of the sparsity of the original channel number, and the first sub-sampling interval is truly contained in the sampling interval of the sparsity of the original channel number; According to the first sub-sampling interval, the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network is obtained, wherein the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network is obtained by sampling in the first sub-sampling interval; According to the sparsity of the channel number of the 2D convolution layer in the reconstructed neural network, an updated value of the channel number of the 2D convolution layer in the reconstructed neural network is obtained; According to the updated value of the channel number of the 2D convolution layer in the reconstructed neural network, the performance prediction model of the neural network architecture is obtained. The data processing module is further configured to: obtain a plurality of sub-sampling intervals; the plurality of sub-sampling intervals are obtained by sampling the sampling interval of the sparsity of the original channel number; and the first sub-sampling interval is any sub-sampling interval in the plurality of sub-sampling intervals.
7. The system for obtaining a performance prediction model of a neural network architecture according to claim 6, wherein, The plurality of sub-sampling intervals are obtained by uniformly sampling or Gaussian sampling the sampling interval of the sparsity of the original channel number. The first sub-sampling interval is preset, or is obtained by randomly sampling the plurality of sub-sampling intervals.
8. The system for obtaining a performance prediction model of a neural network architecture according to claim 7, wherein, The data processing module is configured to:
9. The system for obtaining a performance prediction model of a neural network architecture according to claim 7, wherein, obtain a plurality of reconstructed neural networks; and according to the plurality of reconstructed neural networks, obtain the actual performance of each reconstructed neural network in the plurality of reconstructed neural networks; the actual performance of each reconstructed neural network includes the accuracy of each reconstructed neural network, or / and the latency performance of each reconstructed neural network; 10. The system for obtaining a performance prediction model of a neural network architecture according to any one of claims 6-9, wherein, According to the plurality of reconstructed neural networks and the actual performance of each reconstructed neural network in the plurality of reconstructed neural networks, a plurality of sample pairs of the performance prediction model of the neural network architecture are obtained; and According to the plurality of sample pairs of the performance prediction model of the neural network architecture, the performance prediction model of the neural network architecture is obtained.
Citation Information
Patent Citations
Method and device for processing terminal convolutional neural network, storage medium and processor
CN107316079A
A structure searching method and device of a depth neural network
CN109284820A