Neural Network Architecture Search, Model Publishing Method, Electronic Device, and Storage Medium
By obtaining the computing power information and performance parameters of the application platform, adjusting the model calculation amount using conversion rules, and combining hyper-network search, a target model suitable for the hardware platform is generated, which solves the problem of insufficient utilization of expert experience and computing power in neural network structure design, and realizes efficient model design and operation.
Patent Information
- Application Number
- CN202210033810.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-01-12
AI Technical Summary
In the prior art, neural network structure design relies on expert experience, has long time cycles and high cost, and the designed model cannot fully utilize or exceeds the computing power of the hardware platform, resulting in limited model accuracy and operational efficiency.
By obtaining the computing power information and performance parameters of the application platform, adjusting the model calculation volume using conversion rules, and combining hyper-network search, a target model suitable for the hardware platform is generated.
The target model has fully utilized computing power on the hardware platform, improved the operating efficiency and accuracy of the model, and reduced the time and resource consumption of model design.
Smart Images

Figure CN114492742B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a neural network architecture search, model publishing method, electronic device, and storage medium. Background Art
[0002] Deep learning models have achieved very good results in many tasks. However, the design of neural network architectures highly depends on expert experience. The time cycle for manually designing network architectures is long, and the cost of hiring corresponding experts is high. Neural network architecture search algorithms can enable machines to replace experts in designing neural network architectures and greatly improve model performance, with search efficiency higher than that of manual design.
[0003] Currently, network models designed manually or models obtained by neural network architecture search algorithms often cannot fully utilize the computing power of the hardware platform or exceed the computing power of the hardware platform. When the computing power of the hardware platform cannot be fully utilized, the size of the model will limit the accuracy of the model, and when it exceeds the computing power of the hardware platform, the operating efficiency of the model cannot meet the requirements. Summary of the Invention
[0004] This application provides a neural network architecture search, model publishing method, electronic device, and storage medium, aiming to enable the target model to fully utilize the computing power of the application platform and have high operating efficiency.
[0005] In a first aspect, an embodiment of this application provides a neural network architecture search method, including:
[0006] Obtain first information of a first application platform, where the first information is used to indicate a first computing power value of the first application platform;
[0007] Obtain second information, where the second information is used to indicate performance parameters of a second application platform when running models with different amounts of computation;
[0008] Based on a preset conversion rule, convert the second information according to the first information of the first application platform to obtain third information;
[0009] Obtain target performance parameters of a target model, and determine the amount of computation of the target model from the third information according to the target performance parameters;
[0010] Search for the target model in a preset super network according to the amount of computation of the target model.
[0011] Exemplarily, the obtaining of the second information includes: obtaining performance parameters of the second application platform when running multiple models with different computational amounts; fitting, according to the computational amounts of the models and the corresponding performance parameters, performance parameters of the second application platform when running models with any computational amount within a continuous range, and using the performance parameters of the second application platform when running models with any computational amount within the continuous range as the second information. Performance parameters of a model with a computational amount not actually measured can be obtained by actually measuring performance parameters when running several models with different computational amounts.
[0012] Exemplarily, after the fitting to obtain the performance parameters of the second application platform when running models with any computational amount within a continuous range, the method further includes: obtaining performance parameters of the second application platform when running a test model; correcting, according to the computational amount of the test model and the corresponding performance parameters, the performance parameters of the second application platform when running models with any computational amount within the continuous range obtained by fitting; the using the performance parameters of the second application platform when running models with any computational amount within the continuous range as the second information includes: using the corrected performance parameters of the second application platform when running models with any computational amount within the continuous range as the second information. To improve the accuracy of the second information.
[0013] Exemplarily, the converting, based on a preset conversion rule, the second information according to the first information of the first application platform to obtain third information includes: converting the second information according to a first ratio to obtain third information, where the first ratio is the ratio of the first computing power value of the first application platform indicated by the first information to the second computing power value of the second application platform in the second information. The third information can accurately indicate the performance parameters of the first application platform when running models with different computational amounts.
[0014] Exemplarily, the method further includes: determining the first computing power value of the first application platform according to the product of the nominal computing power value and the computing power utilization rate of the first application platform, where the nominal computing power value is the nominal computing power value of the first application platform indicated by the first information, and the computing power utilization rate is the computing power utilization rate corresponding to the first application platform. When the first computing power value of the first application platform is not actually measured, the first computing power value of the first application platform is determined according to a preset computing power utilization rate.
[0015] Exemplarily, the obtaining of the second information includes: determining, based on a preset correspondence between the second information and the application platform, the second information corresponding to the first application platform according to the first information of the first application platform. It can make the determined second information more accurately reflect the correspondence between the computational amount and the performance parameters of the first application platform.
[0016] Exemplarily, the performance parameter includes processing time, and the conversion of the second information according to the first ratio includes: dividing the processing time corresponding to different amounts of calculation in the second information by the first ratio; or the performance parameter includes processing frame rate, and the conversion of the second information according to the first ratio includes: multiplying the processing frame rate corresponding to different amounts of calculation in the second information by the first ratio.
[0017] Exemplarily, determining the amount of calculation of the target model in the third information according to the target performance parameter includes: based on the correspondence between the amount of calculation and the performance parameter in the third information, determining the amount of calculation corresponding to the target performance parameter as the amount of calculation of the target model. Based on the third information, determining the amount of calculation of the target model according to the target performance parameter of the target model, so as to search for the target model in a preset super network.
[0018] Exemplarily, searching for the target model in a preset super network according to the amount of calculation of the target model includes: searching for a number of models in a pre-trained super network according to the amount of calculation of the target model, where the models include pre-trained weights; training each of the models according to a training data set to obtain a number of trained models; determining the model performance of each of the trained models according to a validation data set; determining the target model among the number of models or determining the target model among the number of trained models according to the model performance of each of the trained models. By searching for the target model in a pre-trained super network, a more suitable network model and corresponding pre-trained weights can be provided for downstream training tasks, and the pre-trained weights can be directly obtained, which can reduce the consumption of time and resources.
[0019] In a second aspect, an embodiment of the present application provides a model publishing method, including:
[0020] Determining a target model according to the foregoing neural network structure search method;
[0021] Sending the target model to a target device according to an instruction of a terminal device. Model distribution can be achieved.
[0022] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor;
[0023] The memory is used to store a computer program;
[0024] The processor is used to execute the computer program and when executing the computer program, implement the steps of the foregoing neural network structure search method or the steps of the foregoing model publishing method.
[0025] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to implement the steps of the above neural network structure search method or the steps of the above model release method.
[0026] An embodiment of the present application provides a neural network structure search, model release method, electronic device, and storage medium. The method includes: obtaining first information of a first application platform, where the first information is used to indicate a first computing power value of the first application platform; obtaining second information, where the second information is used to indicate performance parameters of a second application platform when running models with different amounts of computation; based on a preset conversion rule, converting the second information according to the first information of the first application platform to obtain third information; determining the amount of computation of a target model according to the target performance parameters of the target model in the third information; and searching for the target model in a preset super network according to the amount of computation of the target model. By converting the correspondence between the amount of computation and performance parameters of the model indicated by the second information according to the first information of the first application platform, the third information can accurately indicate the performance parameters of the first application platform when running models with different amounts of computation; and based on the third information, determining the amount of computation of the target model according to the target performance parameters of the target model, so as to search for the target model in a preset super network; when the obtained target model is applied to the first application platform, the computing power of the first application platform can be fully utilized, and the computing power of the first application platform can ensure that the target model has a high operation efficiency.
[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the disclosure of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 is a flowchart of a neural network structure search method provided by an embodiment of the present application;
[0030] Figure 2 is a schematic diagram of the architecture of an AutoML platform in an embodiment;
[0031] Figure 3 is a schematic diagram of the architecture of a neural network in an embodiment;
[0032] Figure 4It is a schematic diagram of the application scenario of the neural network architecture search method in an embodiment;
[0033] Figure 5 It is a schematic diagram of the second information in an embodiment;
[0034] Figure 6 It is a schematic flowchart of a model publishing method provided by an embodiment of the present application;
[0035] Figure 7 It is a schematic block diagram of an electronic device provided by an embodiment of the present application;
[0036] Figure 8 It is a schematic diagram of data interaction between an electronic device and a terminal device in an embodiment. Specific embodiments
[0037] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0038] The flowcharts shown in the accompanying drawings are only illustrative, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.
[0039] Next, some embodiments of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0040] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a neural network architecture search method provided by an embodiment of the present application. The neural network architecture search method can be applied to an electronic device, such as a terminal device or a server, for processes such as models; the model can be called an AI (Artificial Intelligence) model; among them, the terminal device can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device; the server can be an independent server or a server cluster.
[0041] Neural architecture search (NAS) can automatically generate better or even optimal neural network architectures. Generally speaking, the neural network architectures obtained through NAS will run on some resource-constrained devices. If the resource requirements of the neural network architecture obtained by NAS are too large, the target device will not be able to run this neural network architecture properly, which will lead to the need to re-perform NAS, and thus result in low NAS efficiency; if the resource requirements of the neural network architecture obtained by NAS are too small, the computing power of the device cannot be fully utilized, which will limit the accuracy of the model.
[0042] In some embodiments, the neural architecture search method is applied to an automatic machine learning (AutoML) platform.
[0043] Figure 2 It is a schematic diagram of the architecture of an AutoML platform. As Figure 2 shown, the AutoML platform can provide users with data processing and archiving services, model training (train) services, model evaluation (evaluate) services, model management services, model hyperparameter tuning services, and model automatic generation (AutoNet) services. Among them, the AutoNet service can automatically generate AI models for users.
[0044] The AutoNet service can be implemented through an AI model generation system. The AI model generation system can use AI model search algorithms to search for AI models in order to generate a target AI model that meets the user's task objectives. For example: neural architecture search (NAS) technology or efficient neural architecture search (ENAS) technology can be used to generate the target neural network model. NAS and ENAS are emerging technologies in the field of neural networks. They can apply technologies such as reinforcement learning and genetic algorithms to search from scratch for a target neural network model that meets the task objectives for a specific scenario. That is, it can automatically search for and complete the design of the target neural network model using the sample data of a specific scenario. Users do not need to understand the principles of the neural network model to complete the feature extraction of the sample data, the creation, optimization, and performance evaluation of the neural network model, and obtain a target neural network model that meets the requirements of its application scenario (i.e., the task objective), reducing the usage threshold of neural network technology.
[0045] In some embodiments, for the AI model generation system to generate an AI model that meets the user's needs, it needs to include the following three aspects:
[0046] Aspect 1: Search Space - including all modules or combinations of modules for generating an AI model. Each module is used to implement an operation. Multiple modules are connected in a certain combination to form an initial candidate AI model, which may include multiple nodes, and each node includes at least one module (block).
[0047] Taking the AI model as a neural network model as an example, a neural network model is a mathematical model or computational model that mimics the structure and function of a biological neural network (i.e., the central nervous system of an animal, especially the brain). Figure 3 is a schematic diagram of the architecture of a neural network. As Figure 3 shown, the architecture of a neural network is mainly composed of neurons and the connections between neurons. Neurons are the most basic units of a neural network, Figure 3 each ○ in represents a neuron. Each neuron can be connected to one or more other neurons to form a network. It can also be said that a neural network is a complex network formed by a large number of simple neurons widely interconnected.
[0048] Exemplarily, a neural network may include an input layer, an output layer, and a hidden layer. Among them, each neuron in the input layer can receive various different feature information in the sample data. That is, the input layer only receives information from the external environment. Each neuron in the input layer is equivalent to an independent variable, without performing any calculations, and only transmits information to the next layer. There can be at least one hidden layer between the input layer and the output layer. That is, a neural network can have multiple hidden layers, Figure 3 is a neural network with one hidden layer as an example. The hidden layer is used to analyze information. The function used by each neuron in the hidden layer when calculating links the variables of the previous layer and the variables of the next layer to make it more suitable for the data. Finally, the output layer generates the final result. For example, in a classification neural network, each neuron in the output layer corresponds to a specific classification.
[0049] Currently, common operations of the hidden layer include: convolution, pooling, fully connected, etc. The hidden layer used to implement the convolution operation can be called a convolutional layer, and the convolutional layer is mainly used for feature extraction. Common convolution operations include 3*3 convolution, 5*5 convolution, etc. The hidden layer used to implement the pooling operation can be called a pooling layer, and the pooling layer is mainly used for compressing features and simplifying the complexity of neural network calculations. The hidden layer used to implement the fully connected operation can be called a fully connected layer, and the fully connected layer is mainly used to connect all features.
[0050] In this example, the nodes included in the above-mentioned initial candidate AI model can be equivalent to a group of neurons in a neural network. Each node includes at least one module, and each module is used to implement an operation, such as 3*3 convolution, 5*5 convolution, pooling, fully connected, etc.
[0051] Aspect 2: Search strategy - an algorithm for generating an initial candidate AI model.
[0052] Aspect 3: Performance evaluation method - a method for evaluating the performance of a candidate AI model. The candidate AI model mentioned here refers to the AI model obtained after training the initial candidate AI model. It should be understood that the candidate AI model has the same structure as the initial candidate AI model, that is, the candidate AI model may include multiple nodes, and each node includes at least one module.
[0053] The neural network architecture search method according to the embodiments of the present application can ensure that the searched model meets the resource constraint conditions of the target device, thereby improving the search efficiency of the neural network architecture, and the obtained model can make full use of the computing power of the device and has a high operation efficiency.
[0054] Exemplarily, as Figure 4 shown, a schematic diagram of the scenario when the neural network architecture search method is applied to a server. The server can obtain the first information of the first application platform and the target performance parameters of the target model from the terminal device. The server executes the neural network architecture search method to generate the target model, and can also send the generated target model to the terminal device for operations such as model testing or deployment by the terminal device.
[0055] As Figure 1 shown, the neural network architecture search method according to the embodiments of the present application includes steps S110 to S150.
[0056] S110. Obtain the first information of the first application platform, where the first information is used to indicate the first computing power value of the first application platform.
[0057] Exemplarily, the first application platform refers to the platform on which the target model obtained by the neural network architecture search method will be actually applied. Exemplarily, the target model can be deployed on the first application platform to process at least one of images, texts, and voices, and of course, it is not limited thereto; for example, the first application platform can be called the target application platform.
[0058] Exemplarily, the first information of the first application platform includes at least one of the model, architecture, number of cores, and nominal computing power of the first application platform, and / or at least one of the model, architecture, number of cores, and nominal computing power of the processor of the first application platform; the first information of the first application platform can also be referred to as the description information of the first application platform. The processor includes at least one of the following: CPU (Central Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), NPU (Neural-network Processing Units), and of course, it is not limited thereto.
[0059] In some embodiments, the first information includes the first computing power value. For example, the first computing power value can be the nominal computing power; or the corresponding first computing power value can be determined according to the first information. For example, if the first application platform includes n processor cores, and the first computing power value of a platform with one same or similar processor core is obtained in advance, the first computing power value of the first application platform can be obtained by multiplying the first computing power value of this platform by n, where n is a natural number not equal to 0.
[0060] Exemplarily, the method further includes: determining the first computing power value of the first application platform according to the product of the nominal computing power value of the first application platform and the computing power utilization rate, where the nominal computing power value is the nominal computing power value of the first application platform indicated by the first information, and the computing power utilization rate is the computing power utilization rate corresponding to the first application platform. That is, according to the product of the nominal computing power value of the first application platform indicated by the first information and the computing power utilization rate corresponding to the first application platform, the first computing power value of the first application platform indicated by the first information is determined. It can be understood that the first computing power value can be the computing power that the first application platform can actually utilize when running the model. The computing power utilization rate of the first application platform is a preset value or an empirical value.
[0061] For example, the distribution law of the computing power utilization rate of several platforms when running the model is obtained in advance, and the computing power utilization rate corresponding to the first application platform is determined according to this distribution law, such as 20%, 25% or 30%.
[0062] Exemplarily, the computing power utilization rate corresponding to the first application platform can be determined by obtaining the computing power utilization rate of platforms of the same type or architecture when running the model according to the first information of the first application platform.
[0063] For example, the unit of the first computing power value is FLOPS (floating point operations per second), which is an indicator for measuring the performance of hardware.
[0064] S120. Obtain second information, where the second information is used to indicate performance parameters of the second application platform when running models with different computational amounts.
[0065] In some embodiments, the second information is obtained based on the actual operating capabilities measured by running a specific model on a specific device, and can be used to indicate the actual operating efficiency and operating characteristics of the hardware platform, providing a computing power basis for model selection and search.
[0066] In some embodiments, refer to Figure 4 , the second information can be directly obtained from the terminal device. The terminal device providing the second information and the terminal device providing the first information can be the same or different. Exemplarily, prepare several models with different computational amounts, run them on the second application platform respectively, obtain the performance parameters during operation, such as processing time and / or processing frame rate, and associate the performance parameters with the corresponding computational amount, then the second information is obtained. The second information obtained from the terminal device can be stored locally on the server, for example, for subsequent use.
[0067] Exemplarily, the several different computational amounts are distributed according to a preset flops (floating point operations) gradient, and the computational amount can be used to measure the complexity of the algorithm / model.
[0068] Exemplarily, prepare several models with different computational amounts, run them on the second application platform respectively, obtain the actual running time for processing several data, obtain the processing time for each frame of data, and the processing frame rate can be obtained based on the processing time for each frame of data.
[0069] Exemplarily, the second information can be stored in the form of a computing power table, as shown in Table 1:
[0070] Table 1 Second Information
[0071]
[0072] Exemplarily, prepare several models with different computational amounts, run them on different second application platforms respectively, and obtain the second information corresponding to different second application platforms.
[0073] In some other embodiments, the second information can also be obtained through preset processing such as fitting based on the measured data obtained from the terminal device. The obtained second information can be stored locally on the server, for example, for subsequent use.
[0074] Exemplarily, refer to Figure 5 , where curves a (or curve a'), curve b, and curve c in the coordinate system represent second information corresponding to different second application platforms.
[0075] In some embodiments, obtaining the second information includes: based on a preset correspondence between the second information and the application platform, determining the second information corresponding to the first application platform according to the first information of the first application platform.
[0076] Exemplarily, the second application platform corresponding to the second information is the same as the first application platform or has a preset association relationship.
[0077] Exemplarily, the association relationship between the second application platform and the first application platform can be determined according to at least one of the model, architecture, core count, and nominal computing power of the platform, and / or at least one of the model, architecture, core count, and nominal computing power of the processor of the platform. For example, an application platform classification table can be obtained to indicate the association relationship between application platforms, and there is an association relationship between application platforms in the same category. For example, if the processors of the second application platform corresponding to the second information and the first application platform both belong to the Da Vinci architecture, but only the core counts are different, then it can be determined that the second application platform and the first application platform have a preset association relationship.
[0078] For example, the association relationship between application platforms can be represented by Table 2:
[0079] Table 2 Association Relationship between Application Platforms
[0080] Application platform category The processor is a CPU The processor is a GPU The processor is an NPU Application platform A, B C, D, E F
[0081] In some embodiments, prepare several models of the same model type but different computing amounts, and run them on the second application platform respectively to obtain the second information corresponding to this model type. Exemplarily, obtaining the second information includes: determining the second information corresponding to the target model according to the model type of the target model, such as the task type and / or the structure type; wherein, the model type corresponding to the second information is the same as the target model or has a preset association relationship.
[0082] Exemplarily, the task types of the models with the preset association relationship are the same, and / or the structure types are the same. For example, the task types of the model type corresponding to the second information and the target model are both classification or detection, and / or the structure types are both SSD structure or RetinaNet structure.
[0083] Exemplarily, the obtaining of the second information includes: determining the second information according to the first information of the first application platform and the model type of the target model.
[0084] For example, the correspondence between the second information, the first information of the first application platform, and the model type can be represented by Table 3:
[0085] Table 3 Correspondence between model type and second information
[0086]
[0087] In some embodiments, the obtaining of the second information includes: obtaining the performance parameters of the second application platform when running multiple models with different computational amounts; fitting, according to the computational amount of the model and the corresponding performance parameters, to obtain the performance parameters of the second application platform when running models with any computational amount within a continuous range, and using the performance parameters of the second application platform when running models with any computational amount within the continuous range as the second information.
[0088] Exemplarily, the fitting includes, but is not limited to, at least one of regression, interpolation, and approximation, which can make the data of the second information more and cover the performance parameters corresponding to models with more computational amounts. Therefore, the performance parameters of models with computational amounts not actually measured can be obtained by actually measuring the performance parameters when running several models with different computational amounts.
[0089] Exemplarily, a computational amount - performance parameter curve is fitted, and according to this computational amount - performance parameter curve, the performance parameters corresponding to models with any computational amount within the continuous range can be determined. For example, as Figure 5 shown, the circles represent the performance parameters when running several models with different computational amounts obtained by actually measuring a certain second application platform, and the solid curve a represents the performance parameters of the second application platform when running models with any computational amount within the continuous range obtained by fitting.
[0090] Exemplarily, the computational amount - performance parameter curve can be fitted by the least squares method or the like, so that the sum / sum of squares of the distances between the actually measured points and the curve obtained by fitting is minimized. For example, by the least squares method, a straight line or a non - straight curve can be fitted and determined based on more than two points.
[0091] In some embodiments, after fitting to obtain the performance parameters of the second application platform when running models with any amount of computation within a continuous range, the method further includes: obtaining the performance parameters of the second application platform when running a test model; correcting the performance parameters of the second application platform when running models with any amount of computation within a continuous range obtained by fitting according to the amount of computation of the test model and the corresponding performance parameters; and using the performance parameters of the second application platform when running models with any amount of computation within a continuous range as the second information includes: using the corrected performance parameters of the second application platform when running models with any amount of computation within a continuous range as the second information to improve the accuracy of the second information.
[0092] Please refer to Figure 5 , after fitting to obtain the solid curve a, measure the processing frame rate when the newly obtained model runs on the second application platform to obtain new measured points, and correct the solid curve a according to the new measured points to obtain the dashed curve a'. Exemplarily, the amount of computation of the newly obtained model can be determined according to the amount of computation of the most widely used model in actual applications, so that the corrected result can more accurately reflect the relationship between the performance parameters of the model and the amount of computation.
[0093] Exemplarily, the curve to be corrected can be translated and / or rotated so that the translated and / or rotated curve covers at least one measured point, and / or the sum / sum of squares of the distances between all measured points and the corrected curve is minimized. Of course, this is not limited thereto.
[0094] S130. Based on a preset conversion rule, convert the second information according to the first information of the first application platform to obtain third information.
[0095] Limited by the actual measurement conditions, the second application platforms for obtaining the second information are limited, that is, some second application platforms cannot be actually measured due to conditions. By converting the second information of the second application platform based on the conversion rule, the conversion result can be used as the second information of the first application platform.
[0096] In some embodiments, the converting the second information according to the first information of the first application platform based on a preset conversion rule to obtain third information includes: converting the second information according to a first ratio to obtain third information, where the first ratio is the ratio of the first computing power value of the first application platform indicated by the first information to the second computing power value of the second application platform in the second information; that is, converting the second information according to the ratio of the first computing power value of the first application platform indicated by the first information to the second computing power value of the second application platform in the second information to obtain third information.
[0097] Exemplarily, when the first application platform is the same as the second application platform corresponding to the second information, or there is a preset association relationship, such as the same architecture, the second information is converted according to the ratio of the first computing power value to the second computing power value. For example, both the first computing power value and the second computing power value are nominal computing power values; or the first computing power value is the product of the nominal computing power value and the computing power utilization rate, and the second computing power value is the actually measured and actually used computing power value.
[0098] Exemplarily, the performance parameter includes the processing frame rate. The converting the second information according to the first ratio includes: multiplying the processing frame rate corresponding to different amounts of calculation in the second information by the first ratio. For example, the first application platform has four processor cores, and the second application platform corresponding to the second information is a single identical processor core, then the ratio is 4. According to the second information before and after conversion, it can be determined that when the same model runs on the two platforms, the processing frame rate of the first application platform can reach 4 times that of the single-core platform.
[0099] Exemplarily, the performance parameter includes the processing time. The converting the second information according to the first ratio includes: dividing the processing time corresponding to different amounts of calculation in the second information by the first ratio. For example, the first application platform has four processor cores, and the second application platform corresponding to the second information is a single identical processor core, then the ratio is 4. According to the second information before and after conversion, it can be determined that when the same model runs on the two platforms, the processing time for the first application platform to process a single frame of data can reach one-fourth of that of the single-core platform.
[0100] S140. Obtain the target performance parameter of the target model, and determine the amount of calculation of the target model in the third information according to the target performance parameter.
[0101] Exemplarily, the amount of calculation of the target model can be determined in the third information according to the target performance parameter. For example, on the display device of the terminal device, a requirement acquisition interface is displayed, and the user can input the target performance parameter of the target model in the requirement acquisition interface. It can be understood that the target model represents the model required by the user.
[0102] In some embodiments, the service requirements that can be set in the requirement acquisition interface include at least one of the following: application scenario, task type, first application platform, target performance parameter. Thus, the user of the terminal device can set service requirements as needed, so that the server executes the neural network structure search method to obtain a target model that meets the user's needs, and the generated target model is more targeted.
[0103] Exemplarily, determining the computational complexity of the target model from the third piece of information according to the target performance parameter includes: determining, based on the correspondence between the computational complexity and the performance parameter in the third piece of information, that the computational complexity corresponding to the target performance parameter is the computational complexity of the target model.
[0104] Exemplarily, based on the third piece of information, the computational complexity corresponding to the target performance parameter can be determined, and based on this computational complexity, the computational complexity of the target model can be determined. For example, the computational complexity of the target model is the computational complexity corresponding to the target processing frame rate in the third piece of information, or the computational complexity of the target model is ±10% or ±5% of the computational complexity corresponding to the target processing frame rate in the third piece of information. Of course, this is not limited to this. It can be understood that the computational complexity of the target model determined in step S140 can be a value or a range.
[0105] Exemplarily, based on the performance parameter of the second application platform when running a model with any computational complexity within a continuous range after conversion, determine the computational complexity corresponding to the target performance parameter. Please refer to Figure 5 , it is possible to determine the computational complexity of the model when the target processing frame rate is 25.
[0106] Exemplarily, based on the correspondence between the discrete computational complexity and the performance parameter after conversion, determine the computational complexity corresponding to the target performance parameter by means of interpolation or the like. For example, among the performance parameters corresponding to multiple measured computational complexities, determine the computational complexity corresponding to the performance parameter adjacent to the target performance; and determine the computational complexity corresponding to the target performance parameter according to the computational complexity corresponding to the performance parameter adjacent to the target performance.
[0107] S150. Search for the target model in a preset super network according to the computational complexity of the target model.
[0108] Exemplarily, the target model can be a neural network for classifying images, or a neural network for segmenting images, or a neural network for detecting images, or a neural network for recognizing images, or a neural network for generating specified images, or, it can be a neural network for translating text, or, it can be a neural network for paraphrasing text, or a neural network for generating specified text, or a neural network for recognizing speech, or a neural network for translating speech, or a neural network for generating specified speech, etc. From another dimension, the target model can include but is not limited to convolutional target models or recurrent target models, etc.
[0109] Exemplarily, based on an AutoML (Automated Machine Learning technology) system, the target model can be searched and obtained in a preset supernet according to the computational amount of the target model. For example, based on a search control network, a target network structure is searched and obtained in the network structure vector space corresponding to the supernet, and the target model is obtained according to the searched network structure. Among them, the search control network can also be called a controller, and for example, it can be set using an LSTM neural network, and of course, it is not limited to this, and it can also be set using an RNN neural network. The search control network is used to determine a number of network units and the connections between the number of network units in the network structure vector space to obtain the network structure.
[0110] Among them, a supernet is a network model composed of multiple network models through weight sharing. Weight sharing is to share the weights of all network structures in the network structure vector space to achieve the effect of accelerating the search. Using the weight sharing strategy to share and reuse the weights of different network structures in the network structure vector space can improve the search efficiency.
[0111] In some embodiments, the searching for the target model in the preset supernet according to the computational amount of the target model includes: searching for a number of models in the pre-trained supernet according to the computational amount of the target model, where the models include pre-trained weights; training each of the models according to the training dataset to obtain a number of trained models; and determining the target model among the number of trained models.
[0112] Exemplarily, the supernet is trained on datasets such as the ImageNet dataset and the COCO dataset, and the supernet and the trained weights are saved to obtain a pre-trained supernet; among them, the ImageNet dataset is mainly used for classification tasks, and the COCO dataset is mainly used for detection tasks. The models searched and obtained in the pre-trained supernet include pre-trained weights, which can save the training time of the models. For example, when a new model needs to be deployed on a new hardware platform, the supernet can directly load the saved pre-trained weights without additional training time and resources.
[0113] Exemplarily, the search control network searches for a number of network structures in the pre-trained supernet according to the computational amount of the target model, and the number of models can be obtained according to each of the network structures and the weights of each of the network structures in the supernet.
[0114] By training each of the models based on a training data set, the trained models can better adapt to downstream tasks, reducing or avoiding large deviations in the performance of the network structures obtained by search in downstream tasks.
[0115] Optionally, several network structures obtained by search are quickly trained on the training data set of the downstream task. For example, when normal training requires training for n epochs (cycles), quick training only trains for n / m epochs. Other training parameters and training methods of quick training are the same as those of normal training, where n and m are non-zero natural numbers and n is greater than m. By sharing the pre-trained weights of the super network, the training computational amount of the network model obtained by search can be reduced, saving time and resources.
[0116] Exemplarily, the model performance of each of the trained models can be determined according to the validation data set; according to the model performance of each of the trained models, the target model can be determined among the several models, or the target model can be determined among the several trained models.
[0117] For example, by testing the performance of k models after quick training using the validation data set of the downstream task, the model with the best performance can be determined as the target model. The weight of the target model can be the pre-trained weight of the super network or the weight after training based on the training data set.
[0118] The neural network structure search method provided by the embodiments of the present application includes: obtaining first information of a first application platform, where the first information is used to indicate a first computing power value of the first application platform; obtaining second information, where the second information is used to indicate performance parameters of a second application platform when running models with different computational amounts; based on a preset conversion rule, converting the second information according to the first information of the first application platform to obtain third information; determining the computational amount of the target model according to the target performance parameters of the target model in the third information; and searching for the target model in a preset super network according to the computational amount of the target model. By converting the correspondence between the computational amount and performance parameters of the model indicated by the second information according to the first information of the first application platform, the third information can accurately indicate the performance parameters of the first application platform when running models with different computational amounts; and based on the third information, determining the computational amount of the target model according to the target performance parameters of the target model, so as to search for the target model in the preset super network; when the obtained target model is applied to the first application platform, the computing power of the first application platform can be fully utilized, and the computing power of the first application platform can ensure that the target model has a high operation efficiency.
[0119] In some embodiments, the neural network architecture search method according to the embodiments of the present application can obtain a target model by searching in a pre-trained super network, which can provide a more suitable network model and corresponding pre-trained weights for downstream training tasks. Moreover, the pre-trained weights can be directly obtained, which can reduce the consumption of time and resources.
[0120] In some embodiments, the neural network architecture search method according to the embodiments of the present application fits the performance parameters of the second application platform when running models with any computational amount within a continuous range according to the computational amount of the model and the corresponding performance parameters, and uses the performance parameters of the second application platform when running models with any computational amount within a continuous range as the second information. Based on the second information, the computational amount of the target model can be determined more accurately according to the target performance parameters of the target model, avoiding the defect that the existing network models usually only cover individual computational amounts, resulting in the determined target model not being well applicable to the first application platform.
[0121] Exemplarily, the neural network architecture search method includes: obtaining the performance parameters of the second application platform when running models with multiple different computational amounts; fitting the performance parameters of the second application platform when running models with any computational amount within a continuous range according to the computational amount of the model and the corresponding performance parameters; and the performance parameters of the second application platform when running the test model and the computational amount of the test model can also be used to correct the performance parameters of the second application platform when running models with any computational amount within a continuous range obtained by fitting, and using the corrected performance parameters of the second application platform when running models with any computational amount within a continuous range as the second information. Obtain the first information of the first application platform, convert the second information according to the first information to obtain the third information; and determine the computational amount of the target model in the third information according to the target performance parameters of the target model; then the target model can be obtained by searching in a preset super network according to the computational amount of the target model.
[0122] Exemplarily, the neural network architecture search method includes the following steps:
[0123] Step 1: Prepare several models with different computational amounts, such as models with model computational amounts of 1,050,000, 210,000, 105,000,..., 35,000 flops respectively, and run them on the second application platform, such as platform C, platform D, and platform E respectively, to obtain at least one set of performance parameters of the second application platform when running these models, as shown in Table 1. Limited by the actual measurement conditions, usually only the performance parameters of a small number of models can be actually measured.
[0124] Step 2: According to the corresponding relationship between the computational amount and the performance parameters shown in Table 1, it can be fitted to obtain, for example Figure 5The shown computational load - performance parameter curve can indicate the performance parameters of the second application platform when running models with any computational load within a continuous range, and the performance parameters of the model with a computational load that has not been actually measured can be obtained.
[0125] Step 3: Based on the test results of the test model, the computational load - performance parameter curve can be corrected, which can improve the relationship between the performance parameters and the computational load of the model on the second application platform. The second information includes, for example, the corrected computational load - performance parameter curve.
[0126] By conducting actual measurements, fitting, and correction on different types of models on different second application platforms, the second information corresponding to different second application platforms and / or different types of models can be obtained, as shown in Table 2 or Table 3.
[0127] Step 4: After obtaining the first information of the first application platform and the target performance parameters of the target model from the terminal device, the second information corresponding to the first application platform can be determined according to the first information of the first application platform and / or the type of the target model. Referring to Table 3, when the first information indicates that the first application platform is Platform C and the target model is a detection model with an SSD structure, the second information Ⅳ corresponding to the first application platform is obtained; for example, the second information Ⅳ is obtained from the performance parameters of the actual measurement model on the second application platform, and the second application platform includes at least one of Platform C, Platform D, and Platform E. The second information Ⅳ includes, for example, Figure 5 the dotted curve a' in
[0128] Step 5: The second information is converted according to the first information to obtain the third information. For example, the second information Ⅳ is obtained from the performance parameters of the actual measurement model on Platform E, and Platform E has one processor core, and the first information indicates that the first application platform has four identical processor cores, then the second information Ⅳ can be converted; for example, multiply Figure 5 the ordinate of the dotted solid line in
[0129] Step 6: According to the target performance parameters of the target model, the computational load of the target model is determined in the third information. For example, referring to Figure 5 the dotted curve a' in
[0130] Step 7: Search for the target model in a preset supernetwork according to the computational amount of the target model. When performing neural network architecture search in step S150, a model with this computational amount can be searched. For example, multiple searched models can be quickly trained, and based on the validation dataset, the model performance of each model after quick training can be determined, and the model with higher model performance can be selected as the target model.
[0131] Please refer to the above embodiments in conjunction with Figure 6 , Figure 6 which is a schematic flowchart of a model publishing method provided by another embodiment of the present application. The model publishing method can be applied to an electronic device, such as a terminal device or a server, for processes such as generating a model and distributing the model; among them, the terminal device can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device; the server can be an independent server or a server cluster.
[0132] Exemplarily, as Figure 4 shown is a schematic scenario diagram when the model publishing method is applied to a server. The server can interact with the terminal device, execute the neural network architecture search method to generate a target model, and can also send the generated target model to the target device according to the instruction of the terminal device. The target device can be the terminal device or other electronic devices outside the terminal device.
[0133] As Figure 6 shown, the model publishing method of the embodiment of the present application includes step S210 to step S220.
[0134] S210: Determine the target model according to the aforementioned neural network architecture search method.
[0135] Specifically, determining the target model according to the neural network architecture search method includes: obtaining first information of a first application platform, where the first information is used to indicate a first computing power value of the first application platform; obtaining second information, where the second information is used to indicate performance parameters of a second application platform when running models with different computational amounts; based on a preset conversion rule, converting the second information according to the first information of the first application platform to obtain third information; obtaining target performance parameters of the target model, and determining the computational amount of the target model in the third information according to the target performance parameters; and searching for the target model in a preset supernetwork according to the computational amount of the target model.
[0136] S220: Send the target model to the target device according to the instruction of the terminal device.
[0137] The target device can be the terminal device or other electronic devices outside the terminal device. Exemplarily, the server can obtain the target device specified by the terminal device through interaction with the terminal device. The target device is an electronic device for deploying the target model, such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, and a wearable device, etc. Thus, the server can deploy the target model to the target device so that the target device can apply the target model to perform a preset task, such as image classification.
[0138] Please refer to the above embodiments in conjunction with Figure 7 , Figure 7 which is a schematic block diagram of the electronic device 600 provided by the embodiments of the present application.
[0139] Exemplarily, the electronic device can include a terminal device or a server; among them, the terminal device can be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, and a wearable device; the server can be an independent server or a server cluster.
[0140] Please refer to Figure 8 , the electronic device 600, as an AutoML platform, can interact with the terminal device 700 for data, execute a neural network architecture search method to generate a target model, and can also send the generated target model to the target device according to the instruction of the terminal device 700. The target device can be the terminal device 700 or other electronic devices outside the terminal device 700. Exemplarily, the electronic device 600 can obtain the first information of the first application platform and the target performance parameters of the target model from the terminal device 700. The electronic device 600 executes a neural network architecture search method to generate a target model, and can also send the generated target model to the terminal device for operations such as model testing or deployment by the terminal device.
[0141] The electronic device 600 includes a processor 601 and a memory 602.
[0142] Exemplarily, the processor 601 and the memory 602 are connected through a bus 603, and this bus 603 is, for example, an I2C (Inter - integrated Circuit) bus.
[0143] Specifically, the processor 601 can be a micro - control unit (MCU), a central processing unit (CPU), or a digital signal processor (DSP), etc.
[0144] Specifically, the memory 602 can be a Flash chip, a read-only memory (ROM), a magnetic disk, an optical disc, a USB flash drive, a mobile hard disk, etc.
[0145] The processor 601 is configured to run a computer program stored in the memory 602, and implement the steps of the aforementioned neural network architecture search method when executing the computer program.
[0146] Exemplarily, the processor 601 is configured to run a computer program stored in the memory 602, and implement the following steps when executing the computer program:
[0147] Obtain first information of a first application platform, where the first information is used to indicate a first computing power value of the first application platform;
[0148] Obtain second information, where the second information is used to indicate performance parameters of a second application platform when running models with different amounts of computation;
[0149] Based on a preset conversion rule, convert the second information according to the first information of the first application platform to obtain third information;
[0150] Obtain target performance parameters of a target model, and determine the amount of computation of the target model from the third information according to the target performance parameters;
[0151] Search for the target model in a preset super network according to the amount of computation of the target model.
[0152] In some embodiments, the processor 601 is configured to run a computer program stored in the memory 602, and implement the steps of the aforementioned model publishing method when executing the computer program.
[0153] Exemplarily, the processor 601 is configured to run a computer program stored in the memory 602, and implement the following steps when executing the computer program:
[0154] Determine a target model according to the aforementioned neural network architecture search method;
[0155] Send the target model to a target device according to an instruction of a terminal device.
[0156] The specific principles and implementation manners of the electronic device provided in the embodiments of the present application are similar to those of the neural network architecture search method in the foregoing embodiments, and will not be elaborated herein.
[0157] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to implement the steps of the neural network architecture search method provided in the above embodiment.
[0158] Among them, the computer-readable storage medium may be an internal storage unit of the electronic device described in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device.
[0159] It should be understood that the terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0160] It should also be understood that the term "and / or" used in the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0161] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A neural network architecture search method, characterized in that, Including: Obtain first information of a first application platform, where the first information is used to indicate a first computing power value of the first application platform; Obtain second information, where the second information is used to indicate performance parameters of a second application platform when running models with different amounts of computation; Based on a preset conversion rule, convert the second information according to the first information of the first application platform to obtain third information; Obtain target performance parameters of a target model, and determine the amount of computation of the target model from the third information according to the target performance parameters; Search for the target model in a preset super network according to the amount of computation of the target model; Wherein, the performance parameter includes processing time, and the converting the second information according to the first information of the first application platform based on the preset conversion rule to obtain third information includes: dividing the processing time corresponding to different amounts of computation in the second information by a first ratio to obtain the third information, where the first ratio is a ratio of the first computing power value to a second computing power value of the second application platform; or The performance parameter includes processing frame rate, and the converting the second information according to the first information of the first application platform based on the preset conversion rule to obtain third information includes: multiplying the processing frame rate corresponding to different amounts of computation in the second information by a first ratio to obtain the third information, where the first ratio is a ratio of the first computing power value to a second computing power value of the second application platform.
2. The neural network structure search method according to claim 1, wherein The obtaining the second information includes: Obtain performance parameters of the second application platform when running multiple models with different amounts of computation; According to the amount of computation of the model and the corresponding performance parameters, fit to obtain performance parameters of the second application platform when running models with any amount of computation within a continuous range, and use the performance parameters of the second application platform when running models with any amount of computation within the continuous range as the second information.
3. The neural network structure search method according to claim 2, wherein After the fitting to obtain the performance parameters of the second application platform when running models with any amount of computation within a continuous range, the method further includes: Obtain performance parameters of the second application platform when running a test model; According to the amount of computation of the test model and the corresponding performance parameters, correct the performance parameters of the second application platform when running models with any amount of computation within the continuous range obtained by fitting; The using the performance parameters of the second application platform when running models with any amount of computation within a continuous range as the second information includes: using the corrected performance parameters of the second application platform when running models with any amount of computation within the continuous range as the second information.
4. The neural network structure search method according to claim 1, wherein The second computing power value of the second application platform is the second computing power value of the second application platform in the second information.
5. The neural network structure search method according to claim 1, wherein The method further includes: Determine the first computing power value of the first application platform according to the product of the nominal computing power value and the computing power utilization rate of the first application platform, where the nominal computing power value is the nominal computing power value of the first application platform indicated by the first information, and the computing power utilization rate is the computing power utilization rate corresponding to the first application platform.
6. The neural network structure search method according to claim 1, characterized in that The obtaining the second information includes: Based on the correspondence between the preset second information and the application platform, determine the second information corresponding to the first application platform according to the first information of the first application platform.
7. The neural network structure search method according to any one of claims 1-6, characterized in that The determining the computational amount of the target model from the third information according to the target performance parameter includes: Based on the correspondence between the computational amount and the performance parameter in the third information, determine that the computational amount corresponding to the target performance parameter is the computational amount of the target model.
8. The neural network structure search method according to any one of claims 1-6, characterized in that, The searching for the target model in the preset super network according to the computational amount of the target model includes: Search for several models in the pre-trained super network according to the computational amount of the target model, where the models include pre-trained weights; Train each of the models according to the training data set to obtain several trained models; Determine the model performance of each of the trained models according to the validation data set; Determine the target model from the several models or determine the target model from the several trained models according to the model performance of each of the trained models.
9. A model release method, characterized in that, including: Determine the target model according to the neural network structure search method according to any one of claims 1-8; Send the target model to the target device according to the instruction of the terminal device.
10. An electronic device, characterized in that, including a memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program and when executing the computer program, implement: The steps of the neural network structure search method according to any one of claims 1-8; or The steps of the model publishing method according to claim 9.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor is caused to implement: The steps of the neural network structure search method according to any one of claims 1-8; or The steps of the model publishing method according to claim 9.
Citation Information
Patent Citations
Neural network architecture searching method and device
CN110276442A
Neural network structure searching method and neural network structure searching device
CN111382868A