Neural network model processing method, reasoning method and device thereof and electronic device
By dividing and combining the nodes of the neural network model into sub-models suitable for operation of different processors and generating executable sub-programs, the problem that the processing performance of neural network models in the prior art cannot be optimal, and more efficient processing performance is achieved.
Patent Information
- Application Number
- CN202010209744.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-03-23
AI Technical Summary
In the prior art, the processing performance of neural network models on processors cannot be optimal, especially with the development of neural network models, GPUs are gradually unable to meet actual needs in terms of performance and power consumption.
By obtaining the graph model of the neural network model, the nodes are divided into nodes suitable for running on the neural network dedicated processor and nodes that are not suitable for running, and they are merged to form sub-models respectively. These submodels are then compiled to generate executable subprograms, which are run on neural network dedicated processors and other types of processors respectively.
By making full use of the characteristics of multiple processors, the overall processing performance of neural network models is improved, the time of data transmission between processors is reduced, and the utilization of processor resources is maximized.
Smart Images

Figure CN113435565B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and more specifically, to a processing method and reasoning method of a neural network model, and a device and electronic device thereof. Background Art
[0002] In recent years, deep learning has become one of the hottest research directions in the field of artificial intelligence (AI). It has developed rapidly in application fields such as vision, speech, and natural language, and has empowered various industries.
[0003] At present, the vast majority of researchers or companies engaged in deep learning research and application use graphics processing units (GPUs) for training and reasoning of neural network models. However, with the continuous development of neural network models, GPUs are increasingly unable to meet actual needs in terms of performance and power consumption. Therefore, academia and industry have begun to vigorously carry out research on dedicated neural network processors, such as neural network processors (Neural network Processing Unit, NPU) or tensor processors (Tensor Processing Unit, TPU), in order to improve the processor's processing performance for neural network models. However, not all operations and data in neural network models are suitable for processing by dedicated neural network processors. In some cases, running neural network models through dedicated neural network processors cannot achieve the optimal processing performance of the processor for the neural network.
[0004] Therefore, how to improve the processor's processing performance for neural network models is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The embodiments of the present application provide a processing method, an inference method, and an apparatus and electronic device for a neural network model, which can improve the processing performance of a processor for a neural network model.
[0006] In a first aspect, a method for processing a neural network model is provided, comprising: obtaining a graphical model of the neural network model, the graphical model comprising a plurality of nodes, each of the plurality of nodes comprising an operator; dividing the plurality of nodes into at least two categories according to characteristics of the plurality of nodes; merging the plurality of nodes according to the division result to form at least two categories of sub-models; compiling the at least two categories of sub-models to obtain at least two categories of executable sub-programs, the at least two categories of executable sub-programs being used to run on at least two categories of processors.
[0007] In the solution of the present application, according to the characteristics of the nodes in the graph model of the neural network model, the graph model is split into at least two types of sub-models, and after the at least two types of sub-models are compiled respectively, the at least two types of sub-models can be respectively run in at least two types of processors. Compared with a single type of processor for neural network processing, the method of the present application uses multiple types of processors for neural network processing, which can make full use of the characteristics of multiple processors and comprehensively improve the processing performance of multiple processors for the overall neural network.
[0008] In one possible implementation, the multiple nodes are divided into at least two categories based on the characteristics of the multiple nodes, including: based on the characteristics of the multiple nodes, the multiple nodes are divided into two categories, one of which is nodes suitable for running on a neural network dedicated processor, and the other is nodes not suitable for running on a neural network dedicated processor.
[0009] Through the scheme of this embodiment, the nodes of the graph model of the neural network model are divided into nodes suitable for running on a neural network dedicated processor and nodes not suitable for running on a neural network dedicated processor, fully utilizing the characteristics of the neural network dedicated processor, running only part of the operations in the neural network model on the neural network dedicated processor, and setting the operations not suitable for running on the neural network dedicated processor on other processors, thereby improving the processing capability of the neural network dedicated processor for the neural network model, and further comprehensively improving the processing capability of multiple processors for the neural network model.
[0010] In a possible implementation, the multiple nodes are merged according to the division results to form at least two types of sub-models, including: merging the nodes among the multiple nodes that are suitable for running on a neural network dedicated processor to form at least one first sub-model; merging the nodes among the multiple nodes that are not suitable for running on a neural network dedicated processor to form at least one second sub-model; and dividing the at least one first sub-model and the at least one second sub-model to form two types of sub-models.
[0011] In a possible implementation, the at least one first sub-model and the at least one second sub-model are divided into two types of sub-models, including: using the at least one first sub-model as the first type of sub-model in the two types of sub-models, and using the at least one second sub-model as the second type of sub-model in the two types of sub-models.
[0012] In a possible implementation, the at least one first sub-model and the at least one second sub-model are divided into two categories of sub-models, including: calculating the computational amount of each first sub-model in the at least one first sub-model, dividing the first target sub-model with the largest computational amount into the first category of sub-models in the two categories of sub-models, and dividing the other first sub-models except the first target sub-model in the at least one first sub-model and the at least one second sub-model into the second category of sub-models in the two categories of sub-models.
[0013] In a possible implementation, the at least one first sub-model and the at least one second sub-model are divided into two categories of sub-models, including: calculating the computational amount of each first sub-model in the at least one first sub-model, dividing the first target sub-model whose computational amount is greater than a preset threshold into the first category of sub-models in the two categories of sub-models, and dividing the other first sub-models except the first target sub-model in the at least one first sub-model and the at least one second sub-model into the second category of sub-models in the two categories of sub-models.
[0014] In a possible implementation, the at least one first sub-model and the at least one second sub-model are divided into two categories of sub-models, including: calculating the computational amount of each first sub-model in the at least one first sub-model, dividing the first target sub-model with the largest computational amount and the second target sub-model with a computational amount greater than a preset threshold into the first category of sub-models in the two categories of sub-models, and dividing the other first sub-models except the first target sub-model and the second target sub-model in the at least one first sub-model and the at least one second sub-model into the second category of sub-models in the two categories of sub-models.
[0015] Through this implementation, part of the first sub-model is divided into the second type of sub-model, and runs on a non-neural network dedicated processor. Only the first sub-model with the largest amount of computation or the first sub-model with a computation amount greater than a preset threshold is run on the neural network dedicated processor, which can reduce the time of data transmission between processors, maximize the use of the operating resources of the neural network dedicated processor and the non-neural network dedicated processor, and comprehensively improve the processor's processing ability for the neural network.
[0016] In one possible implementation, the calculation of the computational amount of each first sub-model in the at least one first sub-model includes: calculating the computational amount of each first sub-model according to the number of target nodes in each first sub-model, wherein the target node includes an operator of a computational type.
[0017] In one possible implementation, the at least two types of sub-models are compiled to obtain at least two types of executable sub-programs, including: compiling the first type of sub-model to form a first type of executable sub-program, and the first type of executable sub-program is used to run on a neural network dedicated processor; compiling the second type of sub-model to form a second type of executable sub-program, and the second type of executable sub-program is used to run on a non-neural network dedicated processor.
[0018] In a possible implementation, the multiple nodes are divided into at least two categories according to the characteristics of the multiple nodes, including: setting different labels on the multiple nodes to distinguish different categories; and the multiple nodes are merged according to the division results to form at least two types of sub-models, including: according to the labels on the multiple nodes, the multiple nodes are merged to form at least two types of sub-models.
[0019] In a possible implementation, different tags are set on the multiple nodes to distinguish different categories, including: setting a first tag on a node among the multiple nodes that is suitable for running on a neural network dedicated processor, and setting a second tag on a node among the multiple nodes that is not suitable for running on a neural network dedicated processor; merging the multiple nodes according to the tags on the multiple nodes to form at least two types of sub-models, including: merging adjacent nodes among the multiple nodes with the first tag to form at least one first initial sub-model; merging adjacent nodes among the multiple nodes with the second tag to form at least one second initial sub-model; merging the at least one first initial sub-model and the at least one second initial sub-model to form at least one first sub-model and at least one second sub-model; dividing the at least one first sub-model and the at least one second sub-model to form two types of sub-models.
[0020] In a possible implementation, the adjacent nodes with the first mark include a first parent node and a first child node, and the adjacent nodes with the first mark among the multiple nodes are merged, including: if the first parent node and the first child node form a ring structure with each other, the first parent node and the first child node are merged; if the first parent node and the first child node do not form a ring structure, the first parent node and the first child node are merged; if the first parent node, the first child node and other nodes form a ring structure together, the first parent node, the first child node and other nodes are merged in sequence according to preset rules.
[0021] In a possible implementation, no ring structure is formed in each of the at least one first initial sub-model; and no ring structure is formed in each of the at least one second initial sub-model.
[0022] In a possible implementation, the at least one first initial sub-model and the at least one second initial sub-model are merged and judged to form at least one first sub-model and at least one second sub-model, including: if the output node of the first specific initial sub-model in the at least one first initial sub-model is a specific node, the first specific initial sub-model is merged with the second specific initial sub-model, the second specific initial sub-model is the second initial sub-model where the child node of the output node of the first specific initial sub-model is located, and the merged first specific sub-model and the second specific sub-model are used as a second sub-model; if the output node of the second specific initial sub-model in the at least one second initial sub-model is a specific node, the second specific initial sub-model is merged with the first specific initial sub-model, the first specific initial sub-model is the first initial sub-model where the child node of the output node of the second specific initial sub-model is located, and the merged first specific sub-model and the second specific sub-model are used as a second sub-model.
[0023] In a possible implementation, the at least one first initial sub-model and the at least one second initial sub-model are merged and judged to form at least one first sub-model and at least one second sub-model, and further includes: if the output node of one of the at least one first initial sub-model is a non-specific node, the one first initial sub-model is used as a first sub-model; if the output node of one of the at least one second initial sub-model is a non-specific node, the one second initial sub-model is used as a second sub-model.
[0024] In one possible implementation, the processing method further includes: connecting in series the at least two types of executable subroutines to form an executable program of the neural network model; establishing an interface between the executable program and a user; acquiring data through the interface; and executing the executable program to infer the data to obtain an inference result.
[0025] Through the solution of this embodiment, the user can directly input commands and data into the executable program of the neural network model based on a specific user language. At least two types of executable subroutines in the executable program are respectively run in different processors. While improving the processing capability of the neural network model, the neural network model can quickly obtain the data inference results, which is convenient for user operation and improves the user experience.
[0026] In one possible implementation, the first type of executable subprogram in the executable program runs on a neural network dedicated processor, and the second type of executable subprogram in the executable subprogram runs on a non-neural network dedicated processor.
[0027] In a possible implementation, the data is image, voice or text data, and executing the executable program to infer the data to obtain an inference result includes: executing the executable program to perform target detection on the image, voice or text data to obtain a detection result.
[0028] In a second aspect, a method for reasoning a neural network model is provided, comprising: obtaining data input by a user; executing at least two types of executable subprograms in an executable program to reason about the data and obtain reasoning results; wherein the at least two types of executable subprograms are used to run on at least two types of processors, and the at least two types of executable subprograms are obtained by classifying and merging multiple nodes according to the characteristics of multiple nodes in a graph model of the neural network model, and the graph model of the neural network model is used to process the data.
[0029] Through the solution of the present application, users can directly input data into the executable program of the neural network model based on a specific user language. At least two types of executable subroutines in the executable program are respectively run in different processors. While improving the processing capability of the neural network model, the neural network model can quickly obtain the data inference results, which is convenient for users to operate and improves the user experience.
[0030] In one possible implementation, a first type of executable subprogram among the at least two types of executable subprograms runs on a neural network dedicated processor, and a second type of executable subprogram among the at least two types of executable subprograms runs on a non-neural network dedicated processor.
[0031] Through the solution of this embodiment, the processing capability of the neural network model can be further improved by using a neural network dedicated processor to execute the first type of executable subroutine in the executable program.
[0032] In one possible implementation, the first type of executable subprogram is compiled according to the first type of submodel in the graph model of the neural network model, and the second type of executable subprogram is compiled according to the second type of submodel in the graph model of the neural network model; wherein the first type of submodel and the second type of submodel are obtained by merging two types of nodes from the multiple nodes, respectively, the first type of nodes from the two types of nodes are nodes suitable for running on a neural network dedicated processor, and the second type of nodes from the two types of nodes are nodes that are not suitable for running on a neural network dedicated processor.
[0033] In a possible implementation, the first type of sub-model is obtained by merging the first type of nodes; and the second type of sub-model is obtained by merging the second type of nodes.
[0034] In a possible implementation, the first type of sub-model is the first target sub-model with the largest computational complexity in at least one first sub-model formed after the first type of nodes are merged; the second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
[0035] In a possible implementation, the first type of sub-model includes a first target sub-model whose computational complexity is greater than a preset threshold in at least one first sub-model formed after the first type of nodes are merged; the second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
[0036] In a possible implementation, the first type of sub-model includes a first target sub-model with the largest computational complexity in at least one first sub-model formed after the first type of nodes are merged, and a second target sub-model with a computational complexity greater than a preset threshold; the second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model and the second target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
[0037] In a possible implementation, the amount of computation of each first sub-model in the at least one first sub-model is proportional to the number of target nodes in each first sub-model, wherein the target nodes include operators of computational type.
[0038] In a possible implementation, the executable program is obtained by serially connecting each executable subprogram of the at least two types of executable subprograms according to the topological structure of the graph model.
[0039] In a possible implementation, the data is image, voice or text data, and reasoning is performed on the data to obtain a reasoning result, including: performing target detection on the image, voice or text data to obtain a detection result.
[0040] In a third aspect, a processing device for a neural network model is provided, comprising: a first acquisition unit, used to acquire a graphical model of the neural network model, the graphical model comprising a plurality of nodes, each of the plurality of nodes comprising an operator; a first processing unit, used to divide the plurality of nodes into at least two categories according to characteristics of the plurality of nodes; merging the plurality of nodes according to the division result to form at least two categories of sub-models; and compiling the at least two categories of sub-models to obtain at least two categories of executable sub-programs, the at least two categories of executable sub-programs being used to run on at least two types of processors.
[0041] In a fourth aspect, an inference device for a neural network model is provided, comprising: an acquisition unit for acquiring data input by a user; at least two processing units for respectively executing at least two types of executable subroutines in an executable program to infer the data and obtain inference results; wherein the at least two processing units are processing units in at least two types of processors, and the at least two types of executable subroutines are obtained by classifying and merging multiple nodes according to the characteristics of the multiple nodes in a graph model of the neural network model, and the graph model of the neural network model is used to process the data.
[0042] In a fifth aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory is used to store program code, and the processor is used to call the program code to execute the processing method of the neural network model in the first aspect or any possible implementation manner of the first aspect.
[0043] In a sixth aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory is used to store program code, and the processor is used to call the program code to execute the reasoning method of the neural network model in the second aspect or any possible implementation manner of the second aspect.
[0044] In a seventh aspect, a computer-readable storage medium is provided for storing a program code, wherein the program code is used to execute the processing method of the neural network model in the first aspect or any possible implementation manner of the first aspect.
[0045] In an eighth aspect, a computer-readable storage medium is provided for storing program code, wherein the program code is used to execute the reasoning method of the neural network model in the second aspect or any possible implementation manner of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a schematic diagram of an artificial intelligence technology architecture based on deep learning according to an embodiment of the present application;
[0047] Figure 2 is a schematic diagram of a processing system of a neural network model according to an embodiment of the present application;
[0048] Figure 3 is a schematic diagram of a processing system of another neural network model according to an embodiment of the present application;
[0049] Figure 4 is a schematic flow chart of a processing method of a neural network model according to an embodiment of the present application;
[0050] Figure 5 is a structural schematic diagram of a graph model of a neural network model according to an embodiment of the present application;
[0051] Figure 6 is a schematic flow chart of another processing method of a neural network model according to an embodiment of the present application;
[0052] Figure 7 and Figure 8 It is a schematic structural diagram of a graph model of a neural network model after node division and node merging according to an embodiment of the present application;
[0053] Figures 9 to 11 This is a situation of several combinations of parent nodes and child nodes according to the embodiments of the present application;
[0054] Fig.12 and Fig.13 are two schematic diagrams of merging initial sub-models to form sub-models according to embodiments of the present application;
[0055] Fig.14 is a schematic flow chart of another processing method of a neural network model according to an embodiment of the present application;
[0056] Fig.15 is a schematic flow chart of a reasoning method of a neural network model according to an embodiment of the present application;
[0057] Fig.16 is a schematic flow chart of another inference method of a neural network model according to an embodiment of the present application;
[0058] Fig.17 is a schematic block diagram of a processing device for a neural network model according to an embodiment of the present application;
[0059] Fig.18 It is a schematic block diagram of an inference device of a neural network model according to an embodiment of the present application. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0061] To facilitate understanding, the following first introduces relevant concepts such as deep learning, neural network models and related terms involved in the embodiments of the present application.
[0062] Deep learning is also known as deep neural network (DNN). It is essentially a multi-level artificial neural network (ANN) algorithm, which simulates the operating mechanism of the human brain from the structural point of view and simulates the operating mechanism of the human brain from the most basic unit. Deep learning is divided into two parts: training and inference. Training requires massive data input to train a complex deep neural network model. Inference refers to using the trained model and the data to be judged to "infer" various conclusions.
[0063] Figure 1 A schematic diagram of an artificial intelligence technology architecture based on deep learning is shown.
[0064] like Figure 1 As shown in the figure, the artificial intelligence algorithm based on deep learning is mainly implemented on the computer technology architecture, which includes the hardware layer (Hardware), the neural network model compiler (Compiler), the software framework (SoftwareFramework), and the top-level basic application (Application).
[0065] Specifically, the hardware layer provides basic computing power for the algorithm. In addition to the central processing unit (CPU) and GPU, the hardware layer also includes computing chips customized for specific scenarios, such as application-specific integrated circuits (ASIC) chips, field programmable gate arrays (FPGA) chips, etc. Optionally, in addition to the above-mentioned chips, the hardware layer can also include neural network processors NPU, tensor processors TPU, deep learning processors (DPU), etc., which can be used to process neural networks. Optionally, in addition to computing chips, the hardware layer also includes servers customized based on computing chips, GPU server clusters, or various types of mobile terminal devices and computers, etc.
[0066] The neural network model compiler located above the hardware layer is a bridge between the underlying hardware and software frameworks, as well as between different software frameworks. The neural network model compiler is designed to provide a hardware call interface for upper-layer applications, solving problems such as incompatibility that may exist when different upper-layer applications use different underlying hardware computing chips. Its scope includes deep neural network model compilers optimized for artificial intelligence computing chips, as well as regulations and formats for different neural network model representations. Figure 1 As shown in the figure, the neural network model compiler is a compiler created based on the Low Level Virtual Machine (LLVM) framework. Figure 1 Various existing neural network compilers are shown, such as nGraph compiler, NNVM / TVM compiler, ONNX compiler, NNEF compiler, etc.
[0067] At the upper layer of the neural network model compiler, the deep learning algorithm (neural network model) is encapsulated into the software framework, and massive data sets and training strategies are also input into the software framework to train the parameters in the neural network model. The trained neural network model can be used to predict unknown data attributes. Figure 1 As shown, the current mainstream software frameworks include but are not limited to TensorFlow, MXNet, Caffe or Pytorch, etc., which are used for deep learning.
[0068] At present, the current commercial realization of artificial intelligence is mainly based on basic application technologies such as computer vision, intelligent speech, and natural language processing, and corresponding products or services have been formed. Different application requirements are implemented through software frameworks, and further data calculations and processing are performed through the underlying neural network model compiler and hardware layer.
[0069] In general, in the above Figure 1 In the deep learning architecture, for a specific basic application, the processor chip used to process the neural network in the hardware layer is one of the CPU, GPU, ASIC or NPU, and multiple hardware chips are not used to process the same neural network model at the same time. Therefore, corresponding to the chip type in the hardware layer, the neural network model compiler is usually also a single compiler type, adapted to the chip in the hardware layer. In addition, in addition to the hardware chip and compiler, you can choose one Figure 1 Any software framework in the package, train and infer the neural network model, or the deep learning algorithm.
[0070] In the hardware layer, neural network dedicated processors such as NPU and TPU are usually used to process neural network models to improve the processing efficiency and performance of neural network models. However, neural network models have a variety of different neural network model architectures, such as deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN) or graph neural network, etc. Different neural network models have different model architectures and calculation methods. Therefore, not all neural network models are suitable for neural network dedicated processors.
[0071] Regardless of the neural network model, it includes different types of simple or complex operations, and not all operations and data are suitable for processing by neural network dedicated processors. If only neural network dedicated processors or other single types of processors are used for neural network operations and processing, the processing performance of the neural network cannot be optimized.
[0072] Based on the above problems, the present application proposes a method, device and electronic device for processing a neural network model, which splits the neural network model in the software framework to form multiple sub-models. After compiling the multiple sub-models respectively, the multiple sub-models can be run respectively in the CPU, neural network dedicated processor chip or other processor chips, thereby improving the overall processing performance of the neural network.
[0073] In order to better understand the solution of the embodiment of the present application, Figures 2 to 3 The possible application scenarios of the embodiments of the present application are briefly introduced.
[0074] Figure 2 A processing system of a neural network model is shown, and the processing system of the neural network model includes a user device and a data processing device. The user device includes a smart terminal such as a mobile phone, a personal computer or an information processing center. The user device is the initiator of data processing, and is the initiator of application requirements, such as speech recognition, image recognition and other requests. Usually, the user initiates the request through the user device.
[0075] The above-mentioned data processing device can be a device or server with data processing function such as a cloud server, a network server, an application server and a management server. The data processing device receives voice, text or image data from the intelligent terminal through an interactive interface, and then performs data processing such as machine learning and deep learning through the memory for storing data and the processor for data processing. The memory in the data processing device can be a general term, including local storage and databases for storing historical data. The database can be on the data processing device or on other network servers.
[0076] exist Figure 2 In the processing system of the neural network model shown, the user device can receive instructions from the user. For example, the camera in the user device can capture a video file, and then initiate a request to the data processing device, so that the data processing device performs image recognition or target detection on the image frames in the video file obtained by the user device through the neural network model algorithm, thereby obtaining target detection results (such as face recognition, etc.).
[0077] exist Figure 2 In the data processing device, the data processing device can execute the processing method of the neural network model of the present application.
[0078] Figure 3 Another processing system of a neural network model is shown. Figure 3 In the process, the user device directly serves as a data processing device. The user device can directly receive input from the user and process it directly by the hardware of the user device itself. The specific process is similar to Figure 2 Similarly, please refer to the above description and will not go into details here.
[0079] exist Figure 3 In the present invention, the user device itself can execute the processing method of the neural network model of the present application.
[0080] Figure 2 and Figure 3 The processor in the system can perform data training / machine learning / deep learning through a neural network model or other models (for example, a model based on a support vector machine), and use the model finally trained or learned from the data to perform target detection (for example, image recognition, speech recognition, etc.) on voice, image, text and other data, thereby obtaining corresponding processing results.
[0081] Optionally, Figure 2 and Figure 3 The processor in can be Figure 1 In some embodiments, Figure 2 and Figure 3The processors in may include a CPU and a TPU.
[0082] In addition, it should be noted here that Figure 2 and Figure 3 In addition to the CPU and TPU, the TPU may also be other neural network-specific processors, such as NPU, DPU, etc. The embodiment of the present application does not specifically limit the type of specific neural network-specific processors. In order to distinguish neural network-specific processors from other types of processors, in the present application, other types of processors are written as non-neural network-specific processors, which include but are not limited to general-purpose processors or other types of special processors, such as CPU, GPU, etc. The embodiment of the present application does not specifically limit the type of non-neural network-specific processors.
[0083] Next, combine Figures 4 to 14 , describe in detail the processing method of the neural network model of the present application.
[0084] Figure 4 A schematic flow chart of a method 100 for processing a neural network model is shown. The method may be executed by a processor, for example, Figure 2 and Figure 3 The processor in the above may also be executed by a processing device including a processor, for example, Figure 2 and Figure 3 The data processing device in the computer can be a central processing unit (CPU), a microprocessor / micro controller unit (MPU / MCU), an application-specific integrated circuit (ASIC), or a field programmable gate array (FPGA) or other general-purpose processor modules.
[0085] like Figure 4 As shown, the neural network model processing method 100 may include the following steps.
[0086] S110: Obtain a graph model of a neural network model, where the graph model includes a plurality of nodes, and each of the plurality of nodes includes an operator.
[0087] In some embodiments, the neural network model in this step may be an optimized neural network model after training in a software framework, wherein the network model parameters are optimized through a large amount of data and a training algorithm, and the optimized neural network model after training may be directly used for data reasoning. In other words, data may be directly input into the optimized neural network model to obtain the result of reasoning. For example, if the neural network model is used for target detection, the result of reasoning is the result of target detection.
[0088] In other implementations, the neural network model in this step may also be an initial neural network model, that is, an untrained neural network model, in which the network model parameters are customized initial values or random values.
[0089] Specifically, in a software framework, for example, in the TensorFlow software framework, a neural network model is encapsulated as a graph model, where a graph is a data structure consisting of a set of objects (nodes) and their relationships (edges). Nodes can also be called neurons, operators, or operators (OP), indicating how data is calculated; edges can also be called tensors, indicating the flow of data between operators.
[0090] Figure 5 FIG. 1 shows a schematic diagram of the structure of a graph model of a neural network model. Figure 5 As shown, each node represents an operation, corresponds to an operator, and each node has several input tensors and output tensors. Optionally, the input tensors and output tensors of the node can be defined as constants, variables, or placeholders and other data formats. In some embodiments, the input and output tensors of the node are defined as placeholder data formats. When the neural network model is inferred or trained, the original data is input into the placeholder, and the data after training or inference of the neural network model is output through the calculation and transmission of data between nodes.
[0091] S120: Divide the multiple nodes into at least two categories according to the characteristics of the multiple nodes.
[0092] Optionally, in some embodiments, multiple nodes in the graph model of the neural network model can be divided into two categories according to the types of nodes in the graph model, one category being operators suitable for running on a neural network dedicated processor, such as a TPU or NPU, and the other category being operators not suitable for running on a TPU or NPU.
[0093] Optionally, in other implementations, the nodes in the graph model can be divided into two categories based on Single Instruction Multiple Data (SIMD) technology, where nodes that perform the same operation at the same time are suitable for running on the TPU or NPU, and other nodes that do not meet this feature are not suitable for running on the TPU or NPU.
[0094] It should be understood that in addition to dividing multiple nodes into two categories, multiple nodes can also be divided into multiple categories, and the multiple categories of nodes run on different processors. For example, after the division, one type of node is a node suitable for running on a TPU, one type of node is a node suitable for running on a GPU, and another type of node is a node running on a CPU. The embodiment of the present application does not specifically limit the specific division type.
[0095] It should also be understood that in addition to the above-mentioned method, other judgment criteria in the prior art can also be used to divide the nodes in the graph model of the neural network model. The embodiments of the present application do not specifically limit the specific division rules and methods.
[0096] In addition, in an embodiment of the present application, when multiple nodes are divided into at least two categories, different marks can be set on the multiple nodes to distinguish different categories. The marks can be letters, numbers, symbols or a combination thereof, or different marks in any other form, and the embodiment of the present application does not make any specific limitations on this.
[0097] S130: Merge multiple nodes according to the division result to form at least two types of sub-models.
[0098] Optionally, after the node is divided, adjacent nodes of the same type are merged according to the division result to form a merged new node. The merging of the nodes can also be understood as the merging of operators. In order to distinguish the merged new node from the node before the merger, the merged new node is also called a sub-model. A sub-model includes at least one node and also has input and output tensors.
[0099] In some embodiments, after dividing multiple nodes into two categories, nodes suitable for running on a neural network dedicated processor and nodes not suitable for running on a neural network dedicated processor, the nodes suitable for running on the neural network dedicated processor are merged to form one type of sub-model, and the nodes not suitable for running on the neural network dedicated processor are merged to form another type of sub-model.
[0100] Of course, in addition to the above implementation methods, if multiple nodes are divided into multiple types of nodes suitable for running on multiple types of processors, each type of node can be merged to form multiple types of sub-models.
[0101] Optionally, if different labels are set on nodes of different categories, multiple nodes are merged according to the labels on the nodes to form at least two types of sub-models.
[0102] In the embodiment of the present application, while merging to form at least two types of sub-models, a topological relationship between the multiple sub-models and related information of tensors interacting between the sub-models are also formed.
[0103] S140: Compile at least two types of sub-models to obtain at least two types of executable sub-programs, and the at least two types of executable sub-programs run on at least two types of processors respectively.
[0104] Specifically, in this step, at least two types of sub-models are compiled using at least two compilers to obtain at least two types of executable programs, wherein one compiler corresponds to a type of processor and is used to compile a type of sub-model to obtain a type of executable program that can run on the processor. For example, after the CPU compiler compiles the sub-model, the resulting executable program can run on the CPU, and after the TPU compiler compiles the sub-model, the resulting executable program can run on the TPU.
[0105] It should be understood that in this step, the compilation process of the compiler on the sub-model is the same as the compilation process in the prior art. Those skilled in the art can refer to the compilation process in the prior art to implement this step, which will not be repeated here.
[0106] In the scheme of the embodiment of the present application, according to the characteristics of the nodes in the graph model of the neural network model, the graph model is split into at least two types of sub-models, and after the at least two types of sub-models are compiled respectively, the at least two types of sub-models can be respectively run in the CPU, the neural network dedicated processor or other processors. Compared with a single type of processor for neural network processing, the method of the embodiment of the present application is used to perform neural network processing through multiple types of processors, which can make full use of the characteristics of multiple processors and the resources therein, and comprehensively improve the processing performance of multiple processors for the overall neural network.
[0107] In the following, the processing of the graph model to form two types of sub-models is taken as an example for explanation. The relevant method of forming multiple types of sub-models can be found in the relevant description below and will not be repeated here.
[0108] Figure 6 A schematic flow chart of another processing method 100 of a neural network model is shown.
[0109] like Figure 6 As shown, the processing method 100 of the neural network model may include:
[0110] S110: Obtain a graph model of the neural network model.
[0111] S121: Divide the multiple nodes into two categories according to whether the multiple nodes are suitable for running on a neural network dedicated processor.
[0112] This step may be an implementation of the above step S120.
[0113] Specifically, the nodes among the multiple nodes that are suitable for running on the neural network dedicated processor are classified as the first type of nodes, and the nodes among the multiple nodes that are not suitable for running on the neural network dedicated processor are classified as the second type of nodes.
[0114] Optionally, in the process of dividing the multiple nodes, each of the multiple nodes can be marked to form two types of tags. For example, a node suitable for running on a neural network dedicated processor is marked with a first tag, such as support, and a node not suitable for running on a neural network dedicated processor is marked with a second tag, such as unsupport. Of course, the content of the first tag and the second tag can also be any other characters, numbers or letters used to distinguish between the two types of nodes, and the embodiments of the present application do not specifically limit this.
[0115] S131: Merge nodes among multiple nodes that are suitable for running on a neural network dedicated processor to form at least one first initial sub-model; merge nodes among multiple nodes that are not suitable for running on a neural network dedicated processor to form at least one initial second sub-model.
[0116] Step S131 to the following step S133 may be an implementation of the above-mentioned step S130.
[0117] Optionally, adjacent nodes with the first mark among the multiple nodes are merged to form at least one first initial sub-model; and adjacent nodes with the second mark among the multiple nodes are merged to form at least one second initial sub-model.
[0118] Optionally, after the initial sub-models are merged to form, it is also necessary to record the topological relationship between at least one first initial sub-model and at least one second initial sub-model and relevant information of the tensors interacting between the initial sub-models.
[0119] For example, Figure 7 and Figure 8 A schematic structural diagram of a graph model of a neural network model after node division and node merging is shown.
[0120] like Figure 7 As shown in , the graph model includes seven nodes from A to G. After node division and judgment, except for node D, the other six nodes are suitable for running on a neural network dedicated processor. Node D is marked as unsupported, and the other six nodes are marked as supported. Adjacent nodes with the same mark are merged to form the following Figure 8The three initial sub-models shown, among which nodes A, B, and C form initial sub-model 1, node D forms initial sub-model 2, and nodes E, F, and G form initial sub-model 3. Sub-model 1 and sub-model 3 are both the first initial sub-models, and sub-model 2 is the second initial sub-model.
[0121] Optionally, after dividing the graph model into three initial sub-models, the topological relationship between the three initial sub-models and the relevant information of the tensors interacting between the three initial sub-models are recorded. For example, the output tensors of node B and node C are both used as inputs of node D, and the output of node D is also used as inputs of node E and node F.
[0122] In addition, when merging nodes with the same label, it is also necessary to note that a "ring structure" cannot be formed in the first initial sub-model and the second initial sub-model after merging. Figures 9 to 11 The process of node merging in the embodiment of the present application is described.
[0123] Specifically, in the graph model, two adjacent nodes can be called a parent node and a child node, where the output of the parent node is used as the input of the child node. Parent nodes and child nodes with adjacent labels can be merged, for example, the first parent node and the first child node with the first label can be merged, and similarly, the second parent node and the second child node with the second label can also be merged.
[0124] Figures 9 to 11 Several combinations of parent nodes and child nodes are shown. Figures 9 to 11 In the example, the output of node A is the input of node B. In this case, node A can be called the parent node of node B, and node B is the child node of node A. Node A and node B have the same label.
[0125] like Fig. 9 In the case shown, a "ring structure" is directly formed between node A and node B. Before merging other nodes, the nodes forming the ring structure are merged first, that is, node A and node B are merged to form an initial sub-model.
[0126] like Fig.10 As shown in the several cases, the data output by node A can reach node B directly without passing through other nodes, and no "ring structure" will be generated between node A and node B. Therefore, node A and node B can be directly merged to form an initial sub-model.
[0127] like Fig.11 As shown in the several cases in Figure 1, in addition to directly reaching node B, the data output by node A will also pass through other nodes, forming a "ring structure" among node A, node B and other nodes. For example, Fig.11As shown in the figure above, node A is connected to node B, and node A is also connected to node B through node C. In this case, node A and node B cannot be directly merged. You need to merge node A and node C first, and then merge the merged node AC with node B. Alternatively, you can merge node B and node C first, and then merge the merged node BC with node A to form an initial sub-model.
[0128] S132: Perform a merging judgment on the at least one first initial sub-model and the at least one second initial sub-model to form at least one first sub-model and at least one second sub-model.
[0129] Specifically, in some embodiments, if the output node of a first specific initial submodel in at least one first initial submodel is a specific node, the first specific initial submodel is merged with a second specific initial submodel, where the second specific initial submodel is an initial submodel where a child node of the output node of the first specific initial submodel is located;
[0130] If the second specific initial sub-model is the second initial sub-model, the merged first specific sub-model and the second specific sub-model are used as a second sub-model.
[0131] Similarly, in some other embodiments, if the output node of a second specific initial submodel in at least one second initial submodel is a specific node, the second specific initial submodel is merged with the first specific initial submodel, and the first specific initial submodel is the initial submodel where the child node of the output node of the second specific initial submodel is located;
[0132] If the first specific initial sub-model is the first initial sub-model, the merged second specific sub-model and the first specific sub-model are used as a second sub-model.
[0133] Fig.12 and Fig.13 A schematic diagram showing the merging of initial sub-models to form a sub-model in two cases.
[0134] like Fig.12 and 13 As shown, nodes A and B are merged to form initial sub-model 1, and node C is initial sub-model 2.
[0135] If the output node in the initial sub-model 1, that is, the B node, belongs to a specific node, that is, a node that is not suitable as a model output, for example, the node includes specific types of operators and logical operators, etc. At this time, it is necessary to merge the initial sub-model where the child node of the B node is located with the initial sub-model 1 to form a new sub-model to prevent the B node from being used as the output node of the sub-model. That is, in the embodiment of the present application, the child node of the B node is the C node, and the sub-model where the C node is located is the initial sub-model 2. The initial sub-model 2 is merged with the initial sub-model 1 to form a new sub-model.
[0136] like Fig.12 As shown, initial sub-model 1 is the first initial sub-model, that is, nodes A and B are nodes suitable for running on a neural network dedicated processor, initial sub-model 2 is the second initial sub-model, and node C is a node not suitable for running on a neural network dedicated processor. After initial sub-model 1 and initial sub-model 2 are merged, a second sub-model is formed, which runs on a non-neural network dedicated processor.
[0137] like Fig.13 As shown, initial sub-model 1 is the second initial sub-model, that is, both nodes A and B are not suitable for running on a neural network dedicated processor, initial sub-model 2 is the first initial sub-model, and node C is a node suitable for running on a neural network dedicated processor. After initial sub-model 1 and initial sub-model 2 are merged, a second sub-model is also formed, which runs on a non-neural network dedicated processor.
[0138] Above Fig.12 and Fig.13 The situation where specific nodes are included in the initial sub-model is shown. It should be understood that if the output node of the first initial sub-model is a non-specific node, that is, the output node of the first initial sub-model is a node suitable as an output, then the first initial sub-model is regarded as a first sub-model; similarly, if the output node of the second initial sub-model is a non-specific node, the second initial sub-model is regarded as a second sub-model.
[0139] Optionally, in this step, it is necessary to record the topological relationship between at least one first sub-model and at least one second sub-model and relevant information of tensors interacting between the sub-models according to the update.
[0140] S133: Divide at least one first sub-model and at least one second sub-model into two types of sub-models.
[0141] As a possible implementation, S1331: use the at least one first sub-model as a first-type sub-model, and use the at least one second sub-model as a second-type sub-model.
[0142] As another possible implementation, S1332: calculate the computational amount of each first sub-model in at least one first sub-model, divide the first target sub-model with the largest computational amount into the first type of sub-model among the two types of sub-models, and divide the other first sub-models except the first target sub-model in at least one first sub-model and at least one second sub-model into the second type of sub-model among the two types of sub-models.
[0143] As a third possible implementation, S1333: calculate the computational amount of each first sub-model in at least one first sub-model, divide the first target sub-model whose computational amount is greater than a preset threshold into the first type of sub-model among the two types of sub-models, and divide the other first sub-models except the first target sub-model in at least one first sub-model and at least one second sub-model into the second type of sub-model among the two types of sub-models.
[0144] As a fourth possible implementation, S1334: calculate the computational amount of each first sub-model in at least one first sub-model, and divide the first target sub-model with the largest computational amount and the second target sub-model other than the first target sub-model whose computational amount is greater than a preset threshold into the first type of sub-models among the two types of sub-models, and divide the other first sub-models except the first target sub-model and the second target sub-model in at least one first sub-model and at least one second sub-model into the second type of sub-models among the two types of sub-models.
[0145] In the above process of calculating the computational amount of each first sub-model, the number of nodes in the first sub-model that include computational type operators can be counted. For example, operators such as convolution, addition, subtraction, etc. all belong to computational type operators. The nodes that include operators of this operation type can be called target nodes. The more target nodes in the first sub-model, the greater the computational amount of the first sub-model.
[0146] The preset threshold in the above step S1333 and step S1334 may represent a preset number of target nodes in the first sub-model.
[0147] Optionally, in the above step S121, when dividing and marking multiple nodes in the graphical model, it is possible to further mark whether the first type of node suitable for running on a neural network dedicated processor is a target node. For example, marking computing on the target node indicates that the target node includes a computing type operator, and the target node is also marked with a first mark, such as support.
[0148] In this step, the computational complexity of each first sub-model in at least one first sub-model can be quickly calculated directly according to the label.
[0149] In the above steps S1332, S1333 and S1334, part of the first sub-model is divided into the second type of sub-model and runs on a non-neural network dedicated processor. Only the first sub-model with the largest amount of calculation or the first sub-model with a calculation amount greater than a preset threshold is run on the neural network dedicated processor, which can reduce the time of data transmission between processors, maximize the use of the operating resources of the neural network dedicated processor and the non-neural network dedicated processor, and comprehensively improve the processing capacity of the neural network.
[0150] In addition, in the above case, the amount of calculation of the first sub-model and the time consumption of data transmission can be further weighed. If the sub-model runs on a non-neural network dedicated processor, the running speed is slower, but there is no time consumption for data transmission. If the sub-model runs on a neural network dedicated processor, the running speed is faster, but the time consumption for data transmission is increased. Based on comprehensive consideration of the amount of calculation and the time for data transmission, it is determined that the first sub-model runs on a non-neural network dedicated processor or a neural network dedicated processor.
[0151] Optionally, in addition to the above implementation, step S132 may be omitted, and the at least one first initial sub-model may be used as a first-type sub-model, and the at least one second initial sub-model may be used as a second-type sub-model.
[0152] S141: Compile the two types of sub-models to obtain two types of executable sub-programs, and the two types of executable sub-programs run on two types of processors respectively.
[0153] This step may be an implementation of the above step S140.
[0154] Specifically, for the first type of sub-model, a first compiler is used to compile it to obtain a first type of executable sub-program that can run on a neural network dedicated processor. For the second type of sub-model, a second compiler is used to compile it to obtain a second type of executable sub-program that can run on a non-neural network dedicated processor.
[0155] For example, the first type of sub-model is a sub-model suitable for running on the TPU, and the second type of sub-model is a sub-model not suitable for running on the TPU. The first type of sub-model is compiled on the TPU compiler to obtain a first type of executable sub-program that can run on the TPU. The second type of sub-model that is not suitable for running on the TPU can be run on the CPU or other processor chips. For example, the second type of sub-model can be compiled on the CPU compiler to obtain a second type of executable sub-program that can run on the CPU.
[0156] It should be understood that the second type of sub-model can be compiled on a compiler of other processor chips such as a GPU, ASIC, or FPGA in addition to being compiled on a CPU compiler. The embodiments of the present application do not specifically limit the specific compiler type.
[0157] Fig.14 A schematic flow chart of another processing method 100 of a neural network model is shown.
[0158] like Fig.14 As shown, the neural network model processing method 100 may also include the following steps.
[0159] S150: Connect at least two types of executable subroutines in series to form an executable program of the neural network model.
[0160] Specifically, in step 130, while merging to form at least two types of sub-models, a topological relationship between the multiple sub-models and information about the tensors interacting between the sub-models are also formed. In this step, the input and output of the executable programs corresponding to the multiple sub-models are connected in series through the topological relationship of the multiple sub-models and the information of the interacting tensors to form an executable program of the neural network model.
[0161] In one embodiment, for example, in the above step S132, the topological relationship between at least one first sub-model and at least one second sub-model and the related information of the tensor interacting between the sub-models are recorded. In this step, based on the topological relationship and the related information, the connection relationship between the first type of executable sub-program and the second type of executable sub-program is obtained to connect each of the first type of executable sub-program and the second type of executable sub-program in series to form an executable program of the neural network model.
[0162] S160: Establishing an interface between the executable program and the user.
[0163] Specifically, an input and output interface between the executable program of the neural network model and the user is established, that is, an input and output interface based on a specific programming language, such as C, C++, Python, etc. The user inputs data and commands through a specific programming language, and can directly run the executable program without having to input data and commands into the software framework based on the software framework, and then pass through the compiler in turn to form an executable program, making the data processing process more efficient.
[0164] S170: Acquire data through the interface, execute the executable program to infer the data, and obtain inference results.
[0165] Specifically, in this step, the user inputs data through the above-mentioned interface, and at least two types of executable subroutines in the executable program are respectively run on the corresponding at least two types of processors to perform calculations and reasoning on the data to realize the reasoning function of the neural network on the data.
[0166] In some embodiments, the executable program includes a first type of executable subprogram and a second type of executable subprogram, wherein a neural network dedicated processor executes the first type of executable subprogram, and a non-neural network dedicated processor executes the second type of executable subprogram, and the neural network dedicated processor and the non-neural network dedicated processor directly process data input by the user.
[0167] Optionally, the data input by the user can be any type of data such as text, voice or image. In the embodiment of the present application, the inference process of the neural network model for the data can be a process of target detection of text, voice or image, for example, detecting target text, target voice or target image to obtain the result of target detection.
[0168] By adopting the solution of the embodiment of the present application, the user can directly input commands and data into the executable program of the neural network model based on a specific user language. At least two types of executable subroutines in the executable program are respectively run in different processors. While improving the processing capability of the neural network model, the neural network model can quickly obtain the data inference results, which is convenient for user operation and improves the user experience.
[0169] It should be understood that the executable program of the above neural network model can be used to train the neural network model in addition to reasoning about the data input by the user. At this time, the relevant parameters in the executable program can be preset initial parameters, and the data input by the user is a data set for training. By running the executable program, the training results can be obtained, thereby continuously optimizing the relevant parameters in the executable program.
[0170] Combination of the above Figures 4 to 14 , a processing method of a neural network model of the present application is described in detail.
[0171] The above-mentioned neural network processing method includes dividing the graph model of the neural network model, compiling the divided sub-models into multiple executable sub-programs, and packaging the multiple executable sub-programs into one executable program.
[0172] In addition to the processing method, the present application also provides a neural network model reasoning method, in which the above-mentioned division, compilation and packaging processes are not included, and only the final executable program is executed. In other words, the above-mentioned division, compilation and packaging processes are the pre-processes of the reasoning method and are not reflected in the reasoning method.
[0173] Fig.15 A neural network model reasoning method 200 is shown. The reasoning method 200 can be executed by a processing device, which includes at least two types of processors. The at least two types of processors can be located in the same electronic device or in different electronic devices. The data transmission between the at least two types of processors can be through wired transmission or wireless transmission.
[0174] As an example, the processing device may include the above Figure 2 and Figure 3 The processing device may include a general processor module such as a CPU, MPU / MCU, ASIC or FPGA, and a neural network-specific processor module such as an NPU, TPU or DPU.
[0175] like Fig.15 As shown, the reasoning method 200 of the neural network model may include the following steps.
[0176] S210: receiving data input by a user.
[0177] S220: Execute at least two types of executable subprograms in the executable program to infer the data and obtain inference results.
[0178] Among them, at least two types of executable subroutines are used to run on at least two types of processors. The at least two types of executable subroutines are obtained by classifying and merging multiple nodes according to the characteristics of multiple nodes in the graph model of the neural network model. The graph model of the neural network model is used to process the data input by the above-mentioned user.
[0179] Optionally, step S210 and step S220 may be the same as the above-mentioned step S170.
[0180] Optionally, the executable program is obtained by connecting in series each executable subprogram of at least two types of executable subprograms according to the topological structure of the graph model of the neural network model.
[0181] Specifically, an interface is established between the executable program and the user. The user inputs data and commands through a specific programming language, and can directly run at least two types of executable subroutines in the executable program, without having to input data and commands into the software framework based on the software framework, and then pass through the compiler in turn to form an executable program, making the data processing process more efficient.
[0182] The data input by the user is received through the above interface, and at least two types of executable subroutines in the executable program are respectively run on the corresponding at least two types of processors to perform calculations and reasoning on the data, thereby realizing the reasoning function of the neural network on the data.
[0183] Optionally, the data input by the user can be any type of data such as text, voice or image. In an embodiment of the present application, the process of inferring the data can be a process of performing target detection on the text, voice or image, for example, detecting target text, target voice or target image to obtain the result of target detection.
[0184] Specifically, at least two types of executable subroutines in the embodiment of the present application are obtained by classifying and merging multiple nodes according to the characteristics of multiple nodes in the graph model of the neural network model. The executable subroutines converted from the neural network model can better perform data inference and obtain accurate inference results. And the at least two types of executable subroutines are used to run on at least two types of processors respectively, which can make full use of the resources in the processor and improve data processing performance.
[0185] Through the solution of the embodiments of the present application, the user can directly input commands and data into the executable program of the neural network model based on a specific user language. At least two types of executable subroutines in the executable program are respectively run in different processors. While improving the processing capability of the neural network model, the neural network model can quickly obtain the data inference results, which is convenient for user operation and improves the user experience.
[0186] In some embodiments, Fig.16 Another inference method 200 for a neural network model is shown.
[0187] like Fig.16 As shown, the above step S220 may include:
[0188] S221: The neural network dedicated processor executes a first type of executable subprogram among at least two types of executable subprograms, and the non-neural network dedicated processor executes a second type of executable subprogram among at least two types of executable subprograms to infer data and obtain inference results.
[0189] Specifically, in the embodiment of the present application, the first type of executable subprogram and the second type of executable subprogram can refer to the relevant description in the above-mentioned processing method 100.
[0190] In some embodiments, the first type of executable subprogram is compiled according to the first type of submodel in the graph model of the neural network model, and the second type of executable subprogram is compiled according to the second type of submodel in the graph model of the neural network model;
[0191] Among them, the first type of sub-model and the second type of sub-model are obtained by merging two types of nodes in multiple nodes of the neural network model respectively. The first type of nodes in the two types of nodes are nodes suitable for running on a neural network dedicated processor, and the second type of nodes are nodes that are not suitable for running on a neural network dedicated processor.
[0192] Similarly, refer to the description of the two types of nodes and the two types of sub-models in the above processing method 100.
[0193] In a possible implementation, the first type of sub-model is obtained by merging the first type of nodes, and the second type of sub-model is obtained by merging the second type of nodes.
[0194] In another possible implementation, the first type of sub-model is the first target sub-model with the largest computational complexity in at least one first sub-model formed after the first type of nodes are merged; the second type of sub-model includes other first sub-models in at least one first sub-model except the first target sub-model with the largest computational complexity, and at least one second sub-model formed after the second type of nodes are merged.
[0195] In a third possible implementation, the first type of sub-model includes a first target sub-model whose computational complexity is greater than a preset threshold in at least one first sub-model formed after the first type of nodes are merged; the second type of sub-model includes other first sub-models in at least one first sub-model except the first target sub-model whose computational complexity is greater than the preset threshold, and at least one second sub-model formed after the second type of nodes are merged.
[0196] In a fourth possible implementation, the first type of sub-model includes a first target sub-model with the largest computational complexity in at least one first sub-model formed after the first type of nodes are merged, and a second target sub-model with a computational complexity greater than a preset threshold; the second type of sub-model includes other first sub-models in at least one first sub-model except the first target sub-model and the second target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
[0197] Optionally, in the above embodiment, the amount of calculation of each first sub-model in at least one first sub-model is proportional to the number of target nodes in each first sub-model, wherein the target nodes include operators of calculation type.
[0198] Combination of the above Figures 4 to 16 , describes in detail the processing method of the neural network model of the present application and the method embodiment of the reasoning method, and the following is combined with Figure 17 to Figure 18 , the device embodiments of the processing device and the reasoning device of the neural network model of the present application are described in detail. It should be understood that the device embodiments and the method embodiments correspond to each other, and similar descriptions can refer to the method embodiments.
[0199] Fig.17 A schematic block diagram of a processing device 10 for a neural network model is shown. The processing device 10 can be used to execute the processing method 100 for the neural network model.
[0200] like Fig.17 As shown, the processing device 10 of the neural network model includes: a first acquisition unit 11 and a first processing unit 12.
[0201] Specifically, the first acquisition unit 11 is used to acquire a graph model of a neural network model, where the graph model includes a plurality of nodes, and each of the plurality of nodes includes an operator;
[0202] The first processing unit 12 is used to divide multiple nodes into at least two categories according to the characteristics of the multiple nodes; merge the multiple nodes according to the division results to form at least two types of sub-models; compile the at least two types of sub-models to obtain at least two types of executable sub-programs, and the at least two types of executable sub-programs are used to run on at least two types of processors.
[0203] In a possible implementation, the first processing unit 12 is used to: divide the multiple nodes into two categories according to the characteristics of the multiple nodes, one category is nodes suitable for running on a neural network dedicated processor, and the other category is nodes not suitable for running on a neural network dedicated processor.
[0204] In one possible implementation, the first processing unit 12 is used to: merge nodes among a plurality of nodes that are suitable for running on a neural network dedicated processor to form at least one first sub-model; merge nodes among a plurality of nodes that are not suitable for running on a neural network dedicated processor to form at least one second sub-model; and divide at least one first sub-model and at least one second sub-model into two types of sub-models.
[0205] In a possible implementation, the first processing unit 12 is used to: use at least one first sub-model as the first type of sub-model among the two types of sub-models, and use at least one second sub-model as the second type of sub-model among the two types of sub-models.
[0206] In one possible implementation, the first processing unit 12 is used to: calculate the computational amount of each first sub-model in at least one first sub-model, divide the first target sub-model with the largest computational amount into the first type of sub-model among the two types of sub-models, and divide the other first sub-models except the first target sub-model in at least one first sub-model and at least one second sub-model into the second type of sub-model among the two types of sub-models.
[0207] In one possible implementation, the first processing unit 12 is used to: calculate the computational amount of each first sub-model in at least one first sub-model, divide the first target sub-model whose computational amount is greater than a preset threshold into the first type of sub-model among the two types of sub-models, and divide the other first sub-models except the first target sub-model in at least one first sub-model and at least one second sub-model into the second type of sub-model among the two types of sub-models.
[0208] In one possible implementation, the first processing unit 12 is used to: calculate the computational amount of each first sub-model in at least one first sub-model, divide the first target sub-model with the largest computational amount and the second target sub-model with a computational amount greater than a preset threshold into the first type of sub-models among the two types of sub-models, and divide the other first sub-models except the first target sub-model and the second target sub-model in at least one first sub-model and at least one second sub-model into the second type of sub-models among the two types of sub-models.
[0209] In a possible implementation, the first processing unit 12 is used to calculate the computational amount of each first sub-model according to the number of target nodes in each first sub-model, wherein the target nodes include operators of the computational type.
[0210] In one possible implementation, the first processing unit 12 is used to: compile the first type of sub-model to form a first type of executable sub-program, the first type of executable sub-program is used to run on a neural network dedicated processor; compile the second type of sub-model to form a second type of executable sub-program, the second type of executable sub-program is used to run on a non-neural network dedicated processor.
[0211] In a possible implementation, the first processing unit 12 is used to: set different labels on multiple nodes to distinguish different categories; and merge multiple nodes according to the labels on the multiple nodes to form at least two types of sub-models.
[0212] In one possible implementation, the first processing unit 12 is used to: set a first mark on a node among multiple nodes that is suitable for running on a neural network dedicated processor, and set a second mark on a node among multiple nodes that is not suitable for running on a neural network dedicated processor; merge adjacent nodes among multiple nodes with the first mark to form at least one first initial sub-model; merge adjacent nodes among multiple nodes with the second mark to form at least one second initial sub-model; merge and judge at least one first initial sub-model and at least one second initial sub-model to form at least one first sub-model and at least one second sub-model; divide at least one first sub-model and at least one second sub-model into two types of sub-models.
[0213] In a possible implementation, the adjacent nodes with the first mark include a first parent node and a first child node, and the first processing unit 12 is used to: if the first parent node and the first child node form a ring structure with each other, merge the first parent node and the first child node; if no ring structure is formed between the first parent node and the first child node, merge the first parent node and the first child node; if the first parent node, the first child node and other nodes form a ring structure together, merge the first parent node, the first child node and other nodes in sequence according to preset rules.
[0214] In a possible implementation, no ring structure is formed in each of the at least one first initial sub-model; and no ring structure is formed in each of the at least one second initial sub-model.
[0215] In a possible implementation, the first processing unit 12 is used to: if the output node of a first specific initial submodel in at least one first initial submodel is a specific node, merge the first specific initial submodel with a second specific initial submodel, the second specific initial submodel being a second initial submodel where a child node of the output node of the first specific initial submodel is located, and use the merged first specific submodel and the second specific submodel as a second submodel;
[0216] If the output node of a second specific initial submodel in at least one second initial submodel is a specific node, the second specific initial submodel is merged with the first specific initial submodel, the first specific initial submodel is the first initial submodel where the child node of the output node of the second specific initial submodel is located, and the merged first specific submodel and second specific submodel are used as a second submodel.
[0217] In a possible implementation, the first processing unit 12 is further used to: if the output node of one of the at least one first initial sub-models is a non-specific node, use the first initial sub-model as a first sub-model; if the output node of one of the at least one second initial sub-models is a non-specific node, use the second initial sub-model as a second sub-model.
[0218] Alternatively, if Fig.17 As shown, the processing device 10 further includes: a second acquisition unit 13 and a second processing unit 14;
[0219] The first processing unit 12 is also used to: connect at least two types of executable subprograms in series to form an executable program of the neural network model; and establish an interface between the executable program and the user;
[0220] The second acquisition unit 13 is used to: acquire data through an interface;
[0221] The first processing unit 12 is further used to execute one type of executable subprogram among the at least two types of executable subprograms, and the second processing unit 14 is used to execute another type of executable subprogram among the at least two types of executable subprograms, so as to infer the data and obtain inference results.
[0222] In one possible implementation, the second processing unit 14 is a processing unit on a neural network dedicated processor, and the first processing unit 12 is a processing unit on a non-neural network dedicated processor.
[0223] In a possible implementation, the data is image, voice or text data, and the first processing unit 12 and the second processing unit 14 are used to perform target detection on the image, voice or text data to obtain a detection result.
[0224] Fig.18 A schematic block diagram of an inference device 20 for a neural network model is shown. The inference device 20 can be used to execute the inference method 200 for the neural network model.
[0225] like Fig.18 As shown, the inference device 20 of the neural network model includes: an acquisition unit 21 and at least two processing units, such as Fig.18 The first processing unit 22 and the second processing unit 23 in.
[0226] Specifically, the acquisition unit 21 is used to acquire data input by the user;
[0227] At least two processing units, used to respectively execute at least two types of executable subroutines in the executable program to infer the data and obtain inference results;
[0228] Among them, at least two processing units are processing units in at least two types of processors, and at least two types of executable subroutines are obtained by classifying and merging multiple nodes according to the characteristics of multiple nodes in the graphical model of the neural network model, and the graphical model of the neural network model is used to process data.
[0229] As an example, Fig.18 As shown, both the first processing unit 22 and the second processing unit 23 can receive the data transmitted by the acquisition unit 21. In addition to this method, only one of the first processing unit 22 and the second processing unit 23 can receive the data transmitted by the acquisition unit 21, and then the processing unit transmits the received and processed data to another processing unit. In some embodiments, a general-purpose processor, such as a CPU, can receive the data transmitted by the acquisition unit 21, and then process the data and transmit it to other types of processors.
[0230] In one possible embodiment, the first processing unit 22 of the at least two processing units is a processing unit in a neural network dedicated processor, and the first processing unit 22 is used to execute a first type of executable subroutine in at least two types of executable subroutine; the second processing unit 23 of the at least two processing units is a processing unit in a non-neural network dedicated processor, and the second processing unit is used to execute a second type of executable subroutine in at least two types of executable subroutine.
[0231] In one possible implementation, the first type of executable subprogram is compiled based on the first type of submodel in the graph model, and the second type of executable subprogram is compiled based on the second type of submodel in the graph model; the first type of submodel and the second type of submodel are obtained by merging two types of nodes in multiple nodes, respectively, the first type of nodes in the two types of nodes are nodes suitable for running on a neural network dedicated processor, and the second type of nodes in the two types of nodes are nodes that are not suitable for running on a neural network dedicated processor.
[0232] In a possible implementation, the first type of sub-model is obtained by merging the first type of nodes; and the second type of sub-model is obtained by merging the second type of nodes.
[0233] In one possible implementation, the first type of sub-model is the first sub-model with the largest computational complexity among at least one first sub-model formed after the first type of nodes are merged; the second type of sub-model includes other first sub-models among at least one first sub-model except the first sub-model with the largest computational complexity, and at least one second sub-model formed after the second type of nodes are merged.
[0234] In one possible implementation, the first type of sub-model includes at least one first sub-model whose computational complexity is greater than a preset threshold value in at least one first sub-model formed after the first type of nodes are merged; the second type of sub-model includes at least one first sub-model other than the first sub-model whose computational complexity is greater than the preset threshold in at least one first sub-model, and at least one second sub-model formed after the second type of nodes are merged.
[0235] In one possible implementation, the first type of sub-model includes a first target sub-model with the largest computational complexity in at least one first sub-model formed after the first type of nodes are merged, and a second target sub-model with a computational complexity greater than a preset threshold; the second type of sub-model includes other first sub-models in at least one first sub-model except the first target sub-model and the second target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
[0236] In a possible implementation, the amount of computation of each first sub-model in at least one first sub-model is proportional to the number of target nodes in each first sub-model, wherein the target nodes include operators of computational type.
[0237] In a possible implementation, the executable program is obtained by serially connecting each executable subprogram of at least two types of executable subprograms according to the topological structure of the graph model.
[0238] In a possible implementation, the data is image, voice or text data, and at least two processing units are used to perform target detection on the image, voice or text data to obtain a detection result.
[0239] The present application also provides an electronic device, which includes a processor and a memory, wherein the memory is used to store program code, and the processor is used to call the program code to execute the method described in any method embodiment of the present application.
[0240] The present application also provides a computer-readable storage medium for storing program code, which, when executed by a processor, implements the method described in any method embodiment of the present application. The program code may be a high-level language program or an executable target program.
[0241] The computer-readable storage medium is, for example, a memory. The memory may be a volatile memory or a non-volatile memory, or the memory may include both a volatile memory and a non-volatile memory.
[0242] It should be noted that, under the premise of no conflict, the various embodiments described in this application and / or the technical features in each embodiment can be arbitrarily combined with each other, and the technical solution obtained after the combination should also fall within the protection scope of this application.
[0243] It should be understood that the specific examples in the embodiments of the present application are only intended to help those skilled in the art to better understand the embodiments of the present application, rather than to limit the scope of the embodiments of the present application.
[0244] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0245] It should also be understood that the terms used in the embodiments of the present application and the appended claims are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. For example, the singular forms "a", "above", and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0246] Unless otherwise specified, all technical and scientific terms used in the embodiments of the present application have the same meaning as those commonly understood by those skilled in the art of the present application. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the scope of this application.
[0247] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that contains one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0248] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0249] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0250] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0251] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0252] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A processing method for a neural network model, characterized in that: include: Obtaining a graph model of a neural network model, wherein the graph model includes a plurality of nodes, and each of the plurality of nodes includes an operator; Classifying the multiple nodes into at least two categories according to characteristics of the multiple nodes; Merging the multiple nodes according to the division result to form at least two types of sub-models; Compiling the at least two types of sub-models to obtain at least two types of executable sub-programs, wherein the at least two types of executable sub-programs are used to run on at least two types of processors; The processing method also includes: According to the topological relationship between the at least two types of sub-models and the information of the interacting tensors, the at least two types of executable sub-programs are connected in series to form an executable program of the neural network model; Establishing an interface between the executable program and a user; Acquire data through the interface; The executable program is executed to infer the data to obtain an inference result.
2. The processing method according to claim 1, characterized in that: The dividing the multiple nodes into at least two categories according to the characteristics of the multiple nodes includes: According to the characteristics of the multiple nodes, the multiple nodes are divided into two categories, one of which is nodes suitable for running on a neural network dedicated processor, and the other is nodes not suitable for running on a neural network dedicated processor.
3. The processing method according to claim 2, characterized in that: The merging of the multiple nodes according to the division result to form at least two types of sub-models includes: Merging nodes suitable for running on a neural network dedicated processor among the plurality of nodes to form at least one first sub-model; Merging nodes among the plurality of nodes that are not suitable for running on a neural network dedicated processor to form at least one second sub-model; The at least one first sub-model and the at least one second sub-model are divided into two types of sub-models.
4. The processing method according to claim 3, characterized in that: The dividing the at least one first sub-model and the at least one second sub-model into two types of sub-models includes: The at least one first sub-model is used as the first type of sub-model among the two types of sub-models, and the at least one second sub-model is used as the second type of sub-model among the two types of sub-models.
5. The processing method according to claim 3, characterized in that: The dividing the at least one first sub-model and the at least one second sub-model into two types of sub-models includes: Calculate the computational amount of each first sub-model in the at least one first sub-model, divide the first target sub-model with the largest computational amount into the first type of sub-model among the two types of sub-models, divide the other first sub-models except the first target sub-model in the at least one first sub-model and the at least one second sub-model into the second type of sub-model among the two types of sub-models.
6. The processing method according to claim 3, characterized in that: The dividing the at least one first sub-model and the at least one second sub-model into two types of sub-models includes: Calculate the computational amount of each first sub-model in the at least one first sub-model, classify the first target sub-model whose computational amount is greater than a preset threshold into the first type of sub-model among the two types of sub-models, and classify the other first sub-models except the first target sub-model in the at least one first sub-model and the at least one second sub-model into the second type of sub-model among the two types of sub-models.
7. The processing method according to claim 3, characterized in that: The dividing the at least one first sub-model and the at least one second sub-model into two types of sub-models includes: Calculate the computational amount of each first sub-model in the at least one first sub-model, and classify the first target sub-model with the largest computational amount and the second target sub-model with a computational amount greater than a preset threshold as first-category sub-models among the two categories of sub-models, and classify the other first sub-models except the first target sub-model and the second target sub-model in the at least one first sub-model and the at least one second sub-model as second-category sub-models among the two categories of sub-models.
8. The processing method according to any one of claims 5 to 7, characterized in that: The calculating the calculation amount of each first sub-model in the at least one first sub-model comprises: The calculation amount of each first sub-model is calculated according to the number of target nodes in each first sub-model, wherein the target nodes include operators of calculation type.
9. The processing method according to any one of claims 4 to 7, characterized in that: The at least two types of sub-models are compiled to obtain at least two types of executable sub-programs, including: Compiling the first type of sub-model to form a first type of executable sub-program, wherein the first type of executable sub-program is used to run on a neural network dedicated processor; The second type of sub-model is compiled to form a second type of executable sub-program, and the second type of executable sub-program is used to run on a non-neural network dedicated processor.
10. The processing method according to any one of claims 1 to 7, characterized in that: The dividing the multiple nodes into at least two categories according to the characteristics of the multiple nodes includes: Setting different tags on the multiple nodes to distinguish different categories; The merging of the multiple nodes according to the division result to form at least two types of sub-models includes: The multiple nodes are merged according to the labels on the multiple nodes to form at least two types of sub-models.
11. The processing method according to claim 10, characterized in that: The step of setting different tags on the multiple nodes to distinguish different categories includes: Setting a first mark on a node among the plurality of nodes that is suitable for running on a neural network dedicated processor, and setting a second mark on a node among the plurality of nodes that is not suitable for running on a neural network dedicated processor; The merging of the multiple nodes according to the labels on the multiple nodes to form at least two types of sub-models includes: Merging adjacent nodes having the first mark among the multiple nodes to form at least one first initial sub-model; Merging adjacent nodes having the second label among the plurality of nodes to form at least one second initial sub-model; Merge and judge the at least one first initial sub-model and the at least one second initial sub-model to form at least one first sub-model and at least one second sub-model; The at least one first sub-model and the at least one second sub-model are divided into two types of sub-models.
12. The processing method according to claim 11, characterized in that: The adjacent nodes with the first mark include a first parent node and a first child node, and merging the adjacent nodes with the first mark among the multiple nodes includes: If the first parent node and the first child node form a ring structure, merge the first parent node and the first child node; If no ring structure is formed between the first parent node and the first child node, merging the first parent node and the first child node; If the first parent node, the first child node and other nodes together form a ring structure, the first parent node, the first child node and other nodes are merged in sequence according to a preset rule.
13. The processing method according to claim 11, characterized in that: No ring structure is formed in each of the at least one first initial sub-model; No ring structure is formed in each of the at least one second initial sub-model.
14. The processing method according to claim 11, characterized in that: The combining and judging the at least one first initial sub-model and the at least one second initial sub-model to form at least one first sub-model and at least one second sub-model includes: If the output node of a first specific initial submodel in the at least one first initial submodel is a specific node, merging the first specific initial submodel with a second specific initial submodel, where the second specific initial submodel is a second initial submodel where a child node of the output node of the first specific initial submodel is located, and using the merged first specific initial submodel and the second specific initial submodel as a second submodel; If the output node of the second specific initial submodel in the at least one second initial submodel is a specific node, the second specific initial submodel is merged with the first specific initial submodel, where the first specific initial submodel is the first initial submodel where the child node of the output node of the second specific initial submodel is located, and the merged first specific initial submodel and the second initial specific submodel are used as a second submodel.
15. The processing method according to claim 14, characterized in that: The combining and judging the at least one first initial sub-model and the at least one second initial sub-model to form at least one first sub-model and at least one second sub-model further includes: If an output node of one of the at least one first initial sub-models is a non-specific node, taking the one first initial sub-model as a first sub-model; If the output node of one of the at least one second initial sub-models is a non-specific node, the one second initial sub-model is used as a second sub-model.
16. The processing method according to any one of claims 1 to 7, characterized in that: The first type of executable subprograms in the executable program runs on a neural network dedicated processor, and the second type of executable subprograms in the executable subprograms runs on a non-neural network dedicated processor.
17. The processing method according to any one of claims 1 to 7, characterized in that: The data is image, voice or text data, and the executing the executable program to infer the data to obtain an inference result includes: The executable program is executed to perform target detection on image, voice or text data to obtain a detection result.
18. A method for reasoning a neural network model, characterized in that: include: Acquire data input by a user through an interface, wherein the interface is an interface between an executable program and a user; Executing at least two types of executable subprograms in the executable program to infer the data and obtain inference results; The at least two types of executable subprograms are used to run on at least two types of processors, and the at least two types of executable subprograms are obtained by classifying and merging multiple nodes in a graph model of a neural network model according to the characteristics of the multiple nodes, and the graph model of the neural network model is used to process the data; The executable program is obtained by connecting each executable subprogram of the at least two types of executable subprograms in series according to the topological structure of the graph model.
19. The inference method according to claim 18, characterized in that: The first type of executable subprograms among the at least two types of executable subprograms runs on a neural network dedicated processor, and the second type of executable subprograms among the at least two types of executable subprograms runs on a non-neural network dedicated processor.
20. The inference method according to claim 19, characterized in that: The first type of executable subprogram is compiled according to the first type of submodel in the graph model of the neural network model, and the second type of executable subprogram is compiled according to the second type of submodel in the graph model of the neural network model; Among them, the first type of sub-model and the second type of sub-model are obtained by merging two types of nodes in the multiple nodes respectively, the first type of nodes in the two types of nodes are nodes suitable for running on a neural network dedicated processor, and the second type of nodes in the two types of nodes are nodes not suitable for running on a neural network dedicated processor.
21. The inference method according to claim 20, characterized in that: The first type of sub-model is obtained by merging the first type of nodes; The second type of sub-model is obtained by merging the second type of nodes.
22. The inference method according to claim 20, characterized in that: The first type of sub-model is a first target sub-model with the largest computational complexity among at least one first sub-model formed after the first type of nodes are merged; The second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
23. The inference method according to claim 20, characterized in that: The first type of sub-model includes a first target sub-model whose calculation amount is greater than a preset threshold in at least one first sub-model formed after the first type of nodes are merged; The second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
24. The inference method according to claim 20, characterized in that: The first type of sub-model includes a first target sub-model with the largest amount of calculation in at least one first sub-model formed after the first type of nodes are merged, and a second target sub-model with a calculation amount greater than a preset threshold; The second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model and the second target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
25. The inference method according to any one of claims 22 to 24, characterized in that: The amount of calculation of each first sub-model in the at least one first sub-model is proportional to the number of target nodes in each first sub-model, wherein the target nodes include operators of calculation type.
26. The inference method according to any one of claims 18 to 24, characterized in that: The data is image, voice or text data, and the reasoning on the data to obtain the reasoning result includes: Perform target detection on image, voice or text data to obtain detection results.
27. A processing device for a neural network model, characterized in that: include: A first acquisition unit is used to acquire a graph model of a neural network model, wherein the graph model includes a plurality of nodes, and each of the plurality of nodes includes an operator; A first processing unit, configured to classify the plurality of nodes into at least two categories according to characteristics of the plurality of nodes; Merging the multiple nodes according to the division result to form at least two types of sub-models; Compiling the at least two types of sub-models to obtain at least two types of executable sub-programs, wherein the at least two types of executable sub-programs are used to run on at least two types of processors; The processing device further includes: a second acquisition unit and a second processing unit; The first processing unit is further used to: connect the at least two types of executable subprograms in series to form an executable program of the neural network model; and establish an interface between the executable program and a user; The second acquisition unit is used to: acquire data through the interface; The first processing unit is further used to execute one type of executable subprogram among the at least two types of executable subprograms, and the second processing unit is used to execute another type of executable subprogram among the at least two types of executable subprograms, so as to infer the data and obtain inference results.
28. The processing device according to claim 27, characterized in that The first processing unit is used to: divide the multiple nodes into two categories according to the characteristics of the multiple nodes, one category is nodes suitable for running on a neural network dedicated processor, and the other category is nodes not suitable for running on a neural network dedicated processor.
29. The processing device according to claim 28, characterized in that The first processing unit is used to: merge nodes suitable for running on a neural network dedicated processor among the plurality of nodes to form at least one first sub-model; Merging nodes among the plurality of nodes that are not suitable for running on a neural network dedicated processor to form at least one second sub-model; The at least one first sub-model and the at least one second sub-model are divided into two types of sub-models.
30. The processing device according to claim 29, characterized in that The first processing unit is used to: use the at least one first sub-model as the first type of sub-model among the two types of sub-models, and use the at least one second sub-model as the second type of sub-model among the two types of sub-models.
31. The processing device according to claim 29, characterized in that The first processing unit is used to: calculate the computational amount of each first sub-model in the at least one first sub-model, divide the first target sub-model with the largest computational amount into the first type of sub-model among the two types of sub-models, and divide the other first sub-models except the first target sub-model in the at least one first sub-model and the at least one second sub-model into the second type of sub-model among the two types of sub-models.
32. The processing device according to claim 29, characterized in that The first processing unit is used to: calculate the computational amount of each first sub-model in the at least one first sub-model, classify the first target sub-model whose computational amount is greater than a preset threshold into the first type of sub-model among the two types of sub-models, and classify the other first sub-models except the first target sub-model in the at least one first sub-model and the at least one second sub-model into the second type of sub-model among the two types of sub-models.
33. The processing device according to claim 29, characterized in that The first processing unit is used to: calculate the computational amount of each first sub-model in the at least one first sub-model, and classify the first target sub-model with the largest computational amount and the second target sub-model with a computational amount greater than a preset threshold as the first type of sub-models among the two types of sub-models, and classify the other first sub-models except the first target sub-model and the second target sub-model in the at least one first sub-model and the at least one second sub-model as the second type of sub-models among the two types of sub-models.
34. A processing device according to any one of claims 31 to 33, characterized in that The first processing unit is used to calculate the calculation amount of each first sub-model according to the number of target nodes in each first sub-model, wherein the target nodes include operators of calculation type.
35. A processing device according to any one of claims 30 to 33, characterized in that The first processing unit is used to compile the first type of sub-model to form a first type of executable sub-program, and the first type of executable sub-program is used to run on a neural network dedicated processor; The second type of sub-model is compiled to form a second type of executable sub-program, and the second type of executable sub-program is used to run on a non-neural network dedicated processor.
36. A processing device according to any one of claims 27 to 33, characterized in that The first processing unit is used to: set different marks on the multiple nodes to distinguish different categories; The multiple nodes are merged according to the labels on the multiple nodes to form at least two types of sub-models.
37. The processing device according to claim 36, characterized in that The first processing unit is used to: set a first mark on a node suitable for running on a neural network dedicated processor among the multiple nodes, and set a second mark on a node not suitable for running on a neural network dedicated processor among the multiple nodes; Merging adjacent nodes having the first mark among the multiple nodes to form at least one first initial sub-model; Merging adjacent nodes having the second label among the plurality of nodes to form at least one second initial sub-model; Merge and judge the at least one first initial sub-model and the at least one second initial sub-model to form at least one first sub-model and at least one second sub-model; The at least one first sub-model and the at least one second sub-model are divided into two types of sub-models.
38. The processing device according to claim 37, characterized in that The adjacent nodes having the first mark include a first parent node and a first child node, and the first processing unit is used for: If the first parent node and the first child node form a ring structure, merge the first parent node and the first child node; If no ring structure is formed between the first parent node and the first child node, merging the first parent node and the first child node; If the first parent node, the first child node and other nodes together form a ring structure, the first parent node, the first child node and other nodes are merged in sequence according to a preset rule.
39. The processing device according to claim 37, characterized in that No ring structure is formed in each of the at least one first initial sub-model; No ring structure is formed in each of the at least one second initial sub-model.
40. The processing device according to claim 37, characterized in that The first processing unit is used for: if the output node of a first specific initial submodel in the at least one first initial submodel is a specific node, merging the first specific initial submodel with a second specific initial submodel, where the second specific initial submodel is a second initial submodel where a child node of the output node of the first specific initial submodel is located, and treating the merged first specific initial submodel and the second initial specific submodel as a second submodel; If the output node of the second specific initial submodel in the at least one second initial submodel is a specific node, the second specific initial submodel is merged with the first specific initial submodel, where the first specific initial submodel is the first initial submodel where the child node of the output node of the second specific initial submodel is located, and the merged first specific initial submodel and the second specific initial submodel are used as a second submodel.
41. The processing device according to claim 40, characterized in that The first processing unit is further configured to: If an output node of one of the at least one first initial sub-models is a non-specific node, taking the one first initial sub-model as a first sub-model; If the output node of one of the at least one second initial sub-models is a non-specific node, the one second initial sub-model is used as a second sub-model.
42. A processing device according to any one of claims 27 to 33, characterized in that The second processing unit is a processing unit on a neural network dedicated processor, and the first processing unit is a processing unit on a non-neural network dedicated processor.
43. A processing device according to any one of claims 27 to 33, characterized in that The data is image, voice or text data, and the first processing unit and the second processing unit are used to perform target detection on the image, voice or text data to obtain a detection result.
44. An inference device for a neural network model, characterized in that: include: An acquisition unit, used to acquire data input by a user, wherein the acquisition unit includes an interface, which is an interface between the executable program and the user; At least two processing units, used to respectively execute at least two types of executable subroutines in the executable program to infer the data and obtain inference results; The at least two processing units are processing units in at least two types of processors, and the at least two types of executable subroutines are obtained by classifying and merging multiple nodes in a graph model of a neural network model according to characteristics of the multiple nodes, and the graph model of the neural network model is used to process the data; The executable program is obtained by connecting each executable subprogram of the at least two types of executable subprograms in series according to the topological structure of the graph model.
45. The inference device according to claim 44, characterized in that The first processing unit of the at least two processing units is a processing unit in a neural network dedicated processor, and the first processing unit is used to execute a first type of executable subprogram of the at least two types of executable subprograms; The second processing unit of the at least two processing units is a processing unit in a non-neural network dedicated processor, and the second processing unit is used to execute the second type of executable subroutine of the at least two types of executable subroutines.
46. The inference device according to claim 45, characterized in that The first type of executable subprogram is compiled according to the first type of submodel in the graph model, and the second type of executable subprogram is compiled according to the second type of submodel in the graph model; The first type of sub-model and the second type of sub-model are obtained by merging two types of nodes in the multiple nodes respectively, the first type of nodes in the two types of nodes are nodes suitable for running on a neural network dedicated processor, and the second type of nodes in the two types of nodes are nodes not suitable for running on a neural network dedicated processor.
47. The inference device according to claim 46, characterized in that The first type of sub-model is obtained by merging the first type of nodes; The second type of sub-model is obtained by merging the second type of nodes.
48. The inference device according to claim 46, characterized in that The first type of sub-model is a first sub-model with the largest computational complexity among at least one first sub-model formed after the first type of nodes are merged; The second type of sub-models includes other first sub-models except the first sub-model with the largest computational complexity among the at least one first sub-model, and at least one second sub-model formed after the second type of nodes are merged.
49. The inference device according to claim 46, characterized in that The first type of sub-model includes a first sub-model whose calculation amount is greater than a preset threshold in at least one first sub-model formed after the first type of nodes are merged; The second type of sub-model includes other first sub-models in the at least one first sub-model except the first sub-model whose calculation amount is greater than a preset threshold, and at least one second sub-model formed after the second type of nodes are merged.
50. The inference device according to claim 46, characterized in that The first type of sub-model includes a first target sub-model with the largest amount of calculation in at least one first sub-model formed after the first type of nodes are merged, and a second target sub-model with a calculation amount greater than a preset threshold; The second type of sub-model includes other first sub-models in the at least one first sub-model except the first target sub-model and the second target sub-model, and at least one second sub-model formed after the second type of nodes are merged.
51. The inference device according to any one of claims 48 to 50, characterized in that The amount of calculation of each first sub-model in the at least one first sub-model is proportional to the number of target nodes in each first sub-model, wherein the target nodes include operators of calculation type.
52. The inference device according to any one of claims 44 to 50, characterized in that The data is image, voice or text data, and at least two processing units are used for: Perform target detection on image, voice or text data to obtain detection results.
53. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory is used to store program code, and the processor is used to call the program code to execute the processing method of the neural network model described in any one of claims 1 to 17.
54. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory is used to store program code, and the processor is used to call the program code to execute the reasoning method of the neural network model described in any one of claims 18 to 26.
55. A computer-readable storage medium, characterized in that: Used to store program code, wherein the program code is used to execute the processing method of the neural network model according to any one of claims 1 to 17.
56. A computer-readable storage medium, characterized in that Used to store program code, wherein the program code is used to execute the inference method of the neural network model according to any one of claims 18 to 26.
Citation Information
Patent Citations
Processing computational graphs
CN108292241A
Operation method and device and related product
CN109684087A