Model execution method and device, electronic equipment and storage medium
By decomposing the large neural network model into multiple small models and specifying the execution order, the NPU executes each small model in sequence, solving the problem that the NPU cannot handle the large model and achieving effective execution of the entire model.
Patent Information
- Application Number
- CN202510330009.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
When the weighted data volume of the neural network model exceeds the amount of data that the neural network processor NPU can read at one time, the NPU cannot execute the neural network model normally.
By decomposing the large model into multiple small models, the amount of model parameter data of each small model is smaller than the amount of data that the NPU can read at one time, and specify the execution order, so that the NPU executes each small model in sequence to complete the execution of the entire model.
It realizes that when the NPU executes the model, it only reads part of the model parameters for processing each time, and finally completes the execution of the entire model, solving the model execution problem caused by excessive data volume.
Smart Images

Figure CN120216189A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies. Specifically, this application relates to a model execution method, apparatus, electronic device, and storage medium. Background Art
[0002] Neural network models have recently become a research hotspot in the field of artificial intelligence. They abstract the human brain neuron network from the perspective of information processing, establish a certain simple model, and form different networks according to different connection methods. The NPU (Neural Network Processing Unit) plays an important role in data processing in neural network models. Generally, a trained neural network model contains multiple weight data. When it is necessary to perform classification calculations on any input data through the neural network model, the data can be input into the neural network model. The NPU will read all the weight data of the neural network model and process the input data according to the trained processing logic through each weight data, and finally output the desired result.
[0003] With the continuous development of neural network model technologies, the weight data contained in new neural network models is also increasing. Since the NPU needs to read all the weight data of the neural network model at one time when normally executing the neural network model, when the weight data is too much and exceeds the one-time readable capacity of the NPU itself, problems will occur in the execution of the neural network model. Summary of the Invention
[0004] The purpose of this application aims to solve at least one of the above technical defects. The technical solutions provided by the embodiments of this application are as follows: In a first aspect, an embodiment of this application provides a model execution method, including: Storing the input information of the first model, each second model, and the execution order of each second model into a memory space respectively, where the data volume of the model parameters of the first model is greater than the data volume that can be read by the neural network processor NPU of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU at one time; each second model is obtained by decomposing the first model; Sending the first memory space address of the input information, the second memory space addresses of each second model, and the third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtains corresponding outputs according to the input of each second model, and uses the output of the last second model as the processing result; Among them, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0005] In a second aspect, an embodiment of the present application provides a model execution method, including: Receiving a first memory space address of the input information of the first model, a second memory space address of each second model, and a third memory space address of the execution order of each second model sent by the CPU; wherein, the input information, each second model, and the execution order are pre-stored in the memory space by the CPU; the data volume of the model parameters of the first model is greater than the data volume that can be read at one time by the NPU of the first model, and the data volume of the model parameters of each second model is less than the data volume that can be read at one time by the NPU of the first model; the execution order is used to indicate the execution order of each second model; each second model is obtained by decomposing the first model; Reading the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, reading each second model from the memory space in sequence according to the execution order and the second memory space address of each second model, obtaining corresponding outputs according to the input of each second model, and using the output of the last second model as the processing result; Among them, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0006] In a third aspect, an embodiment of the present application provides a model execution device, including: A model execution file storage module, configured to store the input information of the first model, each second model, and the execution order of each second model in the memory space respectively, wherein the data volume of the model parameters of the first model is greater than the data volume that can be read at one time by the neural network processor NPU of the first model, and the data volume of the model parameters of each second model is less than the data volume that can be read at one time by the NPU; each second model is obtained by decomposing the first model; An address sending module, configured to send a first memory space address of the input information, a second memory space address of each second model, and a third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space address of each second model, obtains corresponding outputs according to the input of each second model, and uses the output of the last second model as the processing result; Among them, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0007] Fourthly, an embodiment of the present application provides a model execution device, including: An address receiving module, configured to receive a first memory space address of input information of a first model sent by a CPU, second memory space addresses of each second model, and a third memory space address of an execution order of each second model; wherein, the input information, each second model, and the execution order are pre-stored in the memory space by the CPU; the data volume of the model parameters of the first model is greater than the data volume that can be read by the NPU of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU of the first model at one time; the execution order is used to indicate the execution order of each second model; each second model is obtained by decomposing the first model; A model execution file reading module, configured to read the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, read each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtain corresponding outputs according to the inputs of each second model, and use the output of the last second model as the processing result; wherein, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0008] Fifthly, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory; The processor executes the computer program to implement the method provided in the embodiment of the first aspect or any optional embodiment of the first aspect.
[0009] Sixthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method provided in the embodiment of the first aspect or any optional embodiment of the first aspect is implemented.
[0010] The beneficial effects brought by the technical solution provided by the embodiment of the present application are: In the solution provided by the embodiment of the present application, when it is necessary to execute a first model whose data volume of model parameters exceeds the upper limit of the data volume that can be read by the NPU at one time through the NPU, each second model and the execution order corresponding to the first model are obtained, so that the NPU can execute each second model in sequence based on the execution order, and the output of the last second model is used as the final output. Through the above process, the NPU only reads the model parameters of a part of the model for processing each time, and finally completes the execution of the entire model, solving the problem that the neural network model cannot be executed when the data volume that can be read by the NPU at one time is less than the data volume of the model parameters of the entire neural network model. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application.
[0012] Figure 1 It is a schematic flowchart of a model execution method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the content included in the execution file of the first model and the storage method of model parameters in an example of an embodiment of the present application; Figure 3 It is a schematic flowchart of a model execution method provided by an embodiment of the present application; Figure 4 It is a structural block diagram of a model execution device provided by an embodiment of the present application; Figure 5 It is a structural block diagram of a model execution device provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] The following describes the embodiments of the present application in conjunction with the accompanying drawings in the present application. It should be understood that the implementation manners described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0014] Those skilled in the art of this technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used here may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by this technical field. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by this term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0015] To make the purpose, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the drawings.
[0016] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0017] Figure 1 FIG. is a schematic flowchart of a model execution method provided by an embodiment of the present application. The execution subject of this method can be a CPU (Central Processing Unit), such as Figure 1 shown, this method may include: Step S101, storing the input information of the first model, each second model, and the execution order of each second model into the memory space respectively, where the data volume of the model parameters of the first model is greater than the data volume that can be read by the neural network processor NPU of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU at one time; each second model is obtained by decomposing the first model.
[0018] In the embodiments of the present application, the use of the model can be executed in a server. The server includes a CPU and an NPU. The CPU is mainly responsible for scheduling the files required during the execution of the model, while the NPU is responsible for the main execution calculation part in the model. The first model can be any neural network model, and the second model is obtained by decomposing the first model and can be regarded as a sub-model of the first model. The input information can be the input about the first model input by the user. The model parameters can be internal variables automatically learned during the training of the model, including weights, biases, etc. The model parameters are used to calculate the input information and finally obtain the model output corresponding to the input information. The data volume that can be read by the NPU at one time is determined by the register capacity and data bus width of the NPU itself, and the register capacity and data bus width are usually the same and are an integer power of 2. For example, when the register capacity and data bus width are 32 bits, the NPU is 32bit (bit), then the maximum addressing ability of the NPU is 2 32 (that is, at most 2 32 addressing spaces are searched at one time), converted to a size of 4G, then the data volume that can be read by the NPU at one time is 4G.
[0019] Specifically, since one of the conditions for a model to be executed is that all the model parameters included in the model need to be read by the NPU at one time, when the total data volume of the model parameters included in the first model to be used is greater than the data volume that the current NPU can read at one time, it is necessary to first obtain multiple decomposed second models corresponding to the first model. Since the total data volume of the model parameters included in each second model is less than the data volume that the NPU can read at one time, the NPU can read the model parameters of each second model respectively to achieve the purpose of using the first model. Before this, the CPU is mainly responsible for storing the input information of the first model to be used this time, each second model corresponding to the first model, and the execution order of the second models in the memory space, so that the NPU can read from the memory space subsequently.
[0020] Step S102: Send the first memory space address of the input information, the second memory space addresses of each second model, and the third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtains the corresponding output according to the input of each second model, and takes the output of the last second model as the processing result. Among them, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0021] In the embodiments of the present application, the memory space address can represent the storage location of data in the memory space. For example, the first memory space address represents the storage location of the input information in the memory space, the second memory space address represents the storage location of the corresponding second model in the memory space, and the third memory space address represents the storage location of the execution order of each second model in the memory space.
[0022] Specifically, after the CPU obtains the input information, each second model, and the execution order of each second model, it will first store this content in the memory space of the server. After the storage is completed, the CPU will send the positions where the above content is stored in the memory space (i.e., each memory space address) to the NPU, and the NPU can read this data from the memory space in sequence according to the positions sent by the CPU. Specifically, the NPU will first read out the execution order of each second model and the input information according to the first memory space address and the third memory space address, and then read out the first second model according to the execution order, and input the input information into this second model to obtain a temporary output. After that, it will read out the second second model each time according to the execution order, and input the temporary output of the first model into the second second model to obtain the temporary output of the second second model, and so on, until the temporary output of the last second model is obtained, and this temporary output is used as the processing result of the first model.
[0023] It should be noted that each model in the embodiments of the present application may exist in the form of code, and the model parameters of the second model are stored in the memory space. The model parameters of the same second model will be stored in the same memory space.
[0024] In the solution provided by the embodiments of the present application, when it is necessary to execute a first model whose model parameter data volume exceeds the upper limit of the data volume that the NPU can read at one time, obtain each second model and the execution order corresponding to the first model, so that the NPU can sequentially execute each second model based on the execution order, and use the output of the last second model as the final output. Through the above process, the NPU only reads the model parameters of a part of the models for processing each time, and finally completes the execution of the entire model, solving the problem that the neural network model cannot be executed when the data volume that the NPU can read at one time is less than the data volume of the model parameters of the entire neural network model.
[0025] Based on the above various embodiments, as an optional embodiment, before separately storing the input information of the first model, each second model, and the execution order of each second model into the memory space, it further includes the step of decomposing the first model: Obtain the execution logic of the first model, decompose the first model into each second model based on the execution logic, and generate an execution file for the first model; wherein, the execution logic between any two adjacent second models is coherent; Store each second model into a preset model storage space.
[0026] In an embodiment of the present application, the execution logic may be the data flow logic in the model, which describes the full - process processing mechanism of data from input to output, covering links such as pre - processing, internal model calculations, resource scheduling, and result generation. The execution file may be a file in the elf (Executable and Linkable Format) format. The amount of data that can be stored in this execution file is usually the same as the amount of data that the NPU can read at one time. It is usually used to store information after the decomposition of the first model. Since the amount of data of other information except model parameters is small, the execution file itself can store the model parameters of a second model. Exemplarily, it may include the file names of each second model corresponding to the first model, the execution order of each second model, and the model parameters of the first second model in the execution order, etc. The preset model storage space may be the hard disk space of the server, which exists independently of the memory space.
[0027] Specifically, when the CPU of the server detects that the amount of data of the model parameters of the first model is greater than the amount of data that the NPU can read at one time, the decomposition operation of the first model can be started. Specifically, the CPU can input the first model into a preset compiler for model decomposition. The compiler can decompose the first model into multiple second models according to the execution logic of the first model, and at the same time ensure that the amount of data of the model parameters included in each second model is less than the amount of data that the NPC can read at one time. After obtaining multiple second models, these second models are stored in the preset model storage space together for subsequent direct access without further decomposition. When the decomposition of the first model is completed, an execution file of the first model is also generated. Specifically, as Figure 2 shown Figure 2 in, the first model is decomposed into four second models. The execution file of the first model contains weight0 (the model parameters of the first second model), extra_weight0.bin (the file name of the second second model), extra_weight1.bin (the file name of the third second model), extra_weight2.bin (the file name of the fourth second model), and the model parameters of the second to fourth second models (i.e., extra_weight0, extra_weight1, extra_weight2 in the figure) are stored independently in the preset model storage space.
[0028] It should be noted that there is also a compilation process when the compiler in the embodiment of the present application decomposes the model, that is, converting the decomposed second model into a form that the NPU can recognize.
[0029] Based on the above various embodiments, as an alternative embodiment, the first model includes a preset input layer, and the preset input layer is used to receive the input information for the first model; Before storing the input information of the first model, each second model, and the execution order of each second model into the memory space respectively, it further includes steps of obtaining the input information, each second model, and the execution order: Read the execution order and the file names of each second model from the execution file, and obtain the input information from the preset input layer; Based on the file names of each second model, obtain each stored second model from the preset model storage space.
[0030] In the embodiments of the present application, the preset input layer can be the input layer for the first model to receive user input data. The core function of this preset input layer is to convert the original data input by the user into a format that can be processed by the model.
[0031] Specifically, when the user selects the first model to be used and inputs input information to the preset input layer of the first model, the CPU needs to first obtain the relevant data of the first model that needs to be stored in the memory space. Among them, the execution order and the file names of each second model can be directly read from the execution file corresponding to the first model, and the input information is read from the preset input layer after format conversion through the preset input layer; after reading the file names of each second model, the relevant codes and model parameters of each second model can be obtained from the preset model storage space according to the file names, and after obtaining, these relevant data are stored in the memory space for the NPU to read.
[0032] Based on the above various embodiments, as an alternative embodiment, before generating the execution file for the first model, it further includes steps of performing a signature operation on the model parameters of each second model: For each second model, perform a signature operation on the model parameters of the second model to obtain the first signature value of the second model; among them, the first signature value is used to confirm the integrity of the model parameters of the second model; For each second model, associate the first signature value of the second model with the file name of the second model and write it into the execution file.
[0033] In the embodiments of the present application, the first signature value can be obtained based on a preset signature algorithm, which can be RSA (Rivest–Shamir–Adleman, RSA encryption algorithm), DSA (Digital Signature Algorithm, digital signature algorithm), elliptic curve algorithm, etc. The embodiments of the present application do not make limitations here.
[0034] Specifically, after the first model is decomposed to obtain multiple second models, in order to ensure security during subsequent execution, digital signature operations can be performed on the model parameters of each of the decomposed second models to obtain the first signature value of each second model. The first signature value is obtained by performing a one-time signature operation on the model parameters. When any one of the model parameters changes, the corresponding first signature value will change. Therefore, whether the first signature value changes can be used to verify whether the model parameters of the second model have been tampered with (i.e., whether they are complete); the first signature value and the file name of the corresponding second model can be written into the execution file for storage, and the CPU can obtain them together when reading the execution file.
[0035] Based on the above various embodiments, as an optional embodiment, before obtaining the corresponding output according to the input of each second model, it further includes the step of determining that the model parameters of each second model pass the integrity verification: For each second model read from the memory space, perform a signature operation on the model parameters of the second model to obtain the second signature value of the second model; For each second model read from the memory space, determine the corresponding first signature value from the execution file based on the file name of the second model. If the first signature value is the same as the second signature value, it is determined to pass.
[0036] In the embodiments of the present application, the second signature value can be obtained based on the same preset signature algorithm as the one used to obtain the first signature value.
[0037] Specifically, since the model parameters in the second model stored in the preset model storage space may be tampered with or lost due to system failures or other reasons, and incorrect model parameters will cause the final processing result of the model to deviate from the processing result required by the user. Therefore, after the CPU obtains the model parameters of the second model each time, it is necessary to first verify the integrity of the model parameters of the second model. Specifically, a digital signature operation can be performed on each second model obtained by the CPU this time using the same preset signature algorithm as the one used to obtain the first signature value, to obtain the second signature value of each current second model. Then, the first signature value corresponding to the second model file name is read from the executable file, and the first signature value and the second signature value of the second model are compared. If it is found that the first signature value and the second signature value are the same, it indicates that the model parameters of the second model have not been lost or tampered with, the model parameters of the second model are complete, and the integrity verification result of the second model this time is passed. The CPU can then store the second model obtained this time in the memory space. If it is found that the first signature value and the second signature value are different, it indicates that the model parameters of the second model have been tampered with or lost (i.e., incomplete). At this time, the CPU can send an instruction to the NPU to stop reading the second model, and display a prompt message about "abnormal model parameters of the second model" to the user through the display module of the server (such as a preset display screen).
[0038] Figure 3 The following is a schematic flowchart of a model execution method provided by an embodiment of the present application. The execution subject of this method may be an NPU, as Figure 3 shown, this method may include: Step S301: Receive the first memory space address of the input information of the first model, the second memory space addresses of each second model, and the third memory space address of the execution order of each second model sent by the CPU. Among them, the input information, each second model, and the execution order are pre-stored in the memory space by the CPU. The data volume of the model parameters of the first model is greater than the data volume that the NPU of the first model can read at one time, and the data volume of the model parameters of each second model is less than the data volume that the NPU of the first model can read at one time. The execution order is used to indicate the execution order of each second model. Each second model is obtained by decomposing the first model.
[0039] In an embodiment of the present application, the memory space address may represent the storage location of data in the memory space. For example, the first memory space address represents the storage location of the input information in the memory space, the second memory space address represents the storage location of the corresponding second model in the memory space, and the third memory space address represents the storage location of the execution order of each second model in the memory space. The use of the model may be executed in a server, which includes a CPU and an NPU. The CPU is mainly responsible for scheduling the files required during the execution of the model, while the NPU is responsible for the main execution and calculation part in the model. The first model may be any neural network model, and the second model is obtained by decomposing the first model and can be regarded as a sub-model of the first model. The input information may be the input about the first model input by the user. The model parameters may be internal variables automatically learned by the model during training, including weights, biases, etc. The model parameters are used to calculate the input information and finally obtain the model output corresponding to the input information. The amount of data that the NPU can read at one time is determined by the register capacity and data bus width of the NPU itself, and the register capacity and data bus width are usually the same and are an integer power of 2. For example, when the register capacity and data bus width are 32 bits, the NPU is 32bit (bit), then the maximum addressing ability of the NPU is 2 32 (that is, at most 2 32 addressing spaces are searched at one time), and when converted to size, it is 4G. Then the amount of data that the NPU can read at one time is 4G.
[0040] Specifically, since one of the conditions for a model to be executed is that all the model parameters included in the model need to be read by the NPU at one time, when the total data volume of the model parameters included in the first model to be used is greater than the amount of data that the current NPU can read at one time, it is necessary to first obtain multiple decomposed second models corresponding to the first model. Since the total data volume of the model parameters included in each second model is less than the amount of data that the NPU can read at one time, the NPU can read the model parameters of each second model separately to achieve the purpose of using the first model. Before that, the CPU is mainly responsible for storing the input information of the first model to be used this time, each second model corresponding to the first model, and the execution order of the second model in the memory space, so that the NPU can read from the memory space later.
[0041] After the CPU obtains the input information, each second model, and the execution order of each second model, it will first store these contents in the memory space of the server. After the storage is completed, the CPU will send the storage locations of the above contents in the memory space (that is, each memory space address) to the NPU, and the NPU can read these data from the memory space in sequence according to the locations sent by the CPU.
[0042] Step S302: Read the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively. Read each second model from the memory space in sequence according to the execution order and the second memory space address of each second model. Obtain the corresponding output according to the input of each second model, and use the output of the last second model as the processing result. Among them, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0043] Specifically, the NPU first reads the execution order of each second model and the input information according to the first memory space address and the third memory space address, then reads the first second model according to the execution order, inputs the input information into this second model to obtain a temporary output, and then reads the second second model each time according to the execution order, and inputs the temporary output of the first model into the second second model to obtain the temporary output of the second second model, and so on, until the temporary output of the last second model is obtained, and this temporary output is used as the processing result of the first model.
[0044] It should be noted that each model in the embodiments of the present application may exist in the form of code, and the model parameters of the second model are stored in the memory space. The model parameters of the same second model will be stored in the same memory space.
[0045] In the solution provided by the embodiments of the present application, when it is necessary to execute a first model whose model parameter data volume exceeds the upper limit of the data volume that the NPU can read at one time, obtain each second model corresponding to the first model and the execution order, so that the NPU can execute each second model in sequence based on the execution order, and use the output of the last second model as the final output. Through the above process, the NPU only reads the model parameters of a part of the models for processing each time, and finally completes the execution of the entire model, solving the problem that the neural network model cannot be executed when the data volume that the NPU can read at one time is less than the data volume of the model parameters of the entire neural network model.
[0046] Based on the above various embodiments, as an optional embodiment, an execution file of the first model is also generated when the first model is decomposed, and the execution logic of the first model is also included in the execution file. Before obtaining the corresponding output according to the input of each second model, it also includes the step of allocating memory space for the temporary data of each second model: For each second model, determine the temporary data of the second model based on the execution logic. If it is determined based on the execution logic that the temporary data needs to be used by other second models, then determine whether the temporary data has been added with a preset flag; wherein, the temporary data is non-output data generated by the second model during execution. If the temporary data has not been added with a preset flag, add a preset flag to the temporary data; wherein, the preset flag indicates that memory space has been allocated for the temporary data. If the temporary data carries a preset flag, no memory space will be allocated for the temporary data again.
[0047] In the embodiments of the present application, the temporary data can be data used for transition generated by the model during the calculation process. For example, the data X calculated by the model through calculation step S1 still needs to be used in subsequent calculation steps S2 or S3, but is not used as the final model output. When the temporary data is no longer needed to be used by subsequent calculation steps, the NPU will directly clear the temporary data to save the memory space occupied during the calculation.
[0048] Specifically, under the default setting, each second model will allocate memory space for calculation for the data required during its calculation. These data can be temporary data, or the input or output of the model. For some temporary data, it is generated by the previous second model but needs to be used by the subsequent second model. Although the previous second model has allocated memory space for these temporary data when they are generated, according to the default setting, the subsequent second model will still allocate memory space for them again before using these temporary data. At this time, a scenario of "allocating memory space for the same data multiple times" is caused, which will cause a large waste of memory resources and reduce the processing efficiency of the NPU. Therefore, in the embodiments of the present application, the NPU will analyze this part of the temporary data that needs to cross one or more second models based on the execution logic of the first model, and then add a preset flag (such as cross_model tensor, indicating temporary data that crosses models) to these temporary data. Specifically, for each second model, when determining the temporary data in the second model, it will first determine whether the temporary data has been added with a preset flag. If it has been added, directly read the temporary data from the corresponding memory space during the calculation process of the second model and no new memory space will be allocated for the temporary data. If it has not been added, after allocating memory space for the temporary data, add a preset flag to it. Through the above process, the NPU can avoid repeatedly allocating new memory space for the same temporary data, effectively improving the utilization rate of memory resources and the calculation efficiency of the model.
[0049] Figure 4 The structural block diagram of a model execution device provided by the embodiments of the present application is as Figure 4As shown in the figure, the model execution device 400 may include: a model execution file storage module 401 and an address sending module 402, where, The model execution file storage module 401 is configured to store the input information of the first model, each second model, and the execution order of each second model into the memory space respectively, where the data volume of the model parameters of the first model is greater than the data volume that can be read by the neural network processor NPU of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU at one time; each second model is obtained by decomposing the first model; The address sending module 402 is configured to send the first memory space address of the input information, the second memory space address of each second model, and the third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space address of each second model, obtains the corresponding output according to the input of each second model, and takes the output of the last second model as the processing result; Wherein, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0050] In the solution provided by this application, when it is necessary to execute a first model whose data volume of model parameters exceeds the upper limit of the data volume that can be read by the NPU at one time through the NPU, obtain each second model and the execution order corresponding to the first model, so that the NPU can sequentially execute each second model based on the execution order, and take the output of the last second model as the final output. Through the above process, the NPU only reads the model parameters of a part of the model for processing each time, and finally completes the execution of the entire model, solving the problem that the neural network model cannot be executed when the data volume that can be read by the NPU at one time is less than the data volume of the model parameters of the entire neural network model.
[0051] Based on the above embodiments, as an alternative embodiment, the device further includes a model decomposition module, which is specifically configured to: Obtain the execution logic of the first model, decompose the first model into each second model based on the execution logic, and generate an execution file for the first model; where the execution logic between any two adjacent second models is coherent; Store each second model into a preset model storage space.
[0052] Based on the above embodiments, as an alternative embodiment, the first model includes a preset input layer, and the preset input layer is configured to receive the input information of the first model; The device further includes a model information acquisition module, which is specifically configured to: Read the execution order and the file names of each second model from the execution file, and obtain the input information from the preset input layer; Based on the file names of each second model, obtain each stored second model from the preset model storage space.
[0053] Based on the above embodiments, as an alternative embodiment, the device further includes a model parameter signature module, which is specifically configured to: For each second model, perform a signature operation on the model parameters of the second model to obtain a first signature value of the second model; wherein, the first signature value is used to confirm the integrity of the model parameters of the second model; For each second model, associate the first signature value of the second model with the file name of the second model and write it into the execution file.
[0054] Based on the above embodiments, as an alternative embodiment, the device further includes a model parameter verification module, which is specifically configured to: For each second model, perform a signature operation on the model parameters of the second model to obtain a first signature value of the second model; wherein, the first signature value is used to confirm the integrity of the model parameters of the second model; For each second model, associate the first signature value of the second model with the file name of the second model and write it into the execution file.
[0055] Figure 5 The structural block diagram of a model execution device provided by an embodiment of the present application is shown in Figure 5 As shown, the model execution device 500 may include: an address receiving module 501 and a model execution file reading module 502, wherein, The address receiving module 501 is configured to receive the first memory space address of the input information of the first model, the second memory space addresses of each second model, and the third memory space address of the execution order of each second model sent by the CPU; wherein, the input information, each second model, and the execution order are pre-stored in the memory space by the CPU; the data volume of the model parameters of the first model is greater than the data volume that can be read by the NPU of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU of the first model at one time; the execution order is used to indicate the execution order of each second model; each second model is obtained by decomposing the first model; The model execution file reading module 502 is used to read the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, read each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtain the corresponding output according to the input of each second model, and use the output of the last second model as the processing result; Among them, the input of the first second model is the input information, and the input of a non-first second model is the output of the previous second model.
[0056] In the solution provided by this application, when it is necessary to execute a first model whose model parameter data volume exceeds the data volume upper limit that can be read by the NPU at one time through the NPU, obtain each second model and the execution order corresponding to the first model, so that the NPU can execute each second model in sequence based on the execution order, and use the output of the last second model as the final output. Through the above process, the NPU only reads the model parameters of a part of the models for processing each time, and finally completes the execution of the entire model, solving the problem that the neural network model cannot be executed when the data volume that can be read by the NPU at one time is less than the data volume of the model parameters of the entire neural network model.
[0057] Based on the above various embodiments, as an optional embodiment, the device further includes a memory space allocation module, which is specifically used for: For each second model read from the memory space, perform a signature operation on each model parameter of the second model to obtain the second signature value of the second model; For each second model read from the memory space, determine the corresponding first signature value from the execution file based on the file name of the second model. If the first signature value is the same as the second signature value, it is determined to pass.
[0058] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server that executes the Figure 1 shown method) 600 suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0059] The electronic device includes: a memory and a processor. The memory is used to store programs for executing the methods described in the above various method embodiments; the processor is configured to execute the programs stored in the memory. Herein, the processor can be referred to as the processing device 601 described below, and the memory can include at least one of the read-only memory (ROM) 602, random access memory (RAM) 603, and storage device 608 described below, as specifically shown below: As Figure 6 shown, the electronic device 600 can include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, ROM 602, and RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0060] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0061] Specifically, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiments of the present application are executed.
[0062] It should be noted that the above-mentioned computer-readable storage medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0063] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of the communication network include a local area network (“LAN”), a wide area network (“WAN”), the Internet (for example, the Internet), and a peer-to-peer network (for example, an ad hoc peer-to-peer network), as well as any currently known or future-developed network.
[0064] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.
[0065] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to: Store the input information of the first model, each second model, and the execution order of each second model into the memory space respectively, where the data volume of the model parameters of the first model is greater than the data volume that can be read by the neural network processor (NPU) of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU at one time; each second model is obtained by decomposing the first model; send the first memory space address of the input information of the first model, the second memory space addresses of each second model, and the third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtains the corresponding output according to the input of each second model, and uses the output of the last second model as the processing result.
[0066] Or, Receive the first memory space address of the input information of the first model, the second memory space addresses of each second model, and the third memory space address of the execution order of each second model sent by the CPU; where the input information, each second model, and the execution order are pre-stored in the memory space by the CPU; the data volume of the model parameters of the first model is greater than the data volume that can be read by the NPU of the first model at one time, and the data volume of the model parameters of each second model is less than the data volume that can be read by the NPU of the first model at one time; the execution order is used to indicate the execution order of each second model; each second model is obtained by decomposing the first model; Read the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, read each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtain the corresponding output according to the input of each second model, and use the output of the last second model as the processing result.
[0067] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0069] The modules or units described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module or unit does not, in some cases, constitute a limitation on the unit itself. For example, the first constraint acquisition module can also be described as "the module for acquiring the first constraint".
[0070] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0071] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0072] It should be understood that although the steps in the flowcharts of the figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and they can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts of the figures can include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a portion of other steps or sub-steps or stages of other steps.
[0073] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A model execution method, characterized in that: Applied to central processing unit CPU, including: The input information of the first model, each second model and the execution order of each second model are stored in the memory space respectively, wherein the data amount of the model parameters of the first model is greater than the data amount that can be read at one time by the neural network processor NPU of the first model, and the data amount of the model parameters of each second model is less than the data amount that can be read at one time by the NPU; each second model is obtained by decomposing the first model; Sending the first memory space address of the input information, the second memory space addresses of each second model, and the third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address, respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtains a corresponding output according to the input of each second model, and takes the output of the last second model as the processing result; The input of the first second model is the input information, and the input of the non-first second model is the output of the previous second model.
2. The method according to claim 1, characterized in that The step of storing the input information of the first model, each second model and the execution order of each second model in the memory space respectively includes the step of decomposing the first model: Acquire the execution logic of the first model, decompose the first model into the second models based on the execution logic, and generate an execution file for the first model; wherein the execution logic between any two adjacent second models is coherent; The second models are stored in a preset model storage space.
3. The method according to claim 2, characterized in that The first model includes a preset input layer, and the preset input layer is used to receive input information of the first model; Before storing the input information of the first model, each second model and the execution order of each second model in the memory space respectively, the method further includes the step of obtaining the input information, each second model and the execution order: Reading the execution order and the file names of the second models from the execution file, and acquiring the input information from the preset input layer; Based on the file names of the second models, the stored second models are acquired from the preset model storage space.
4. The method according to claim 2, characterized in that Before generating the execution file of the first model, the method further includes the step of performing a signing operation on the model parameters of each second model: For each second model, perform a signing operation on each model parameter of the second model to obtain a first signature value of the second model; wherein the first signature value is used to confirm the integrity of the model parameters of the second model; For each second model, the first signature value of the second model is associated with the file name of the second model and then written into the execution file.
5. The method according to claim 4, characterized in that Before obtaining the corresponding output according to the input of each second model, the method further includes the step of determining whether the model parameters of each second model pass the integrity verification: For each second model read from the memory space, performing a signing operation on each model parameter of the second model to obtain a second signature value of the second model; For each second model read from the memory space, a corresponding first signature value is determined from the execution file based on the file name of the second model. If the first signature value is the same as the second signature value, it is determined to be passed.
6. A model execution method, characterized in that: Applied to NPU, including: Receive the first memory space address of the input information of the first model, the second memory space address of each second model and the third memory space address of the execution order of each second model sent by the CPU; wherein the input information, the second models and the execution order are pre-stored in the memory space by the CPU; the data volume of the model parameters of the first model is larger than the data volume that can be read by the NPU of the first model at one time, and the data volume of the model parameters of each second model is smaller than the data volume that can be read by the NPU of the first model at one time; the execution order is used to indicate the execution order of each second model; each second model is decomposed from the first model; Read the input information and the execution order from the memory space according to the first memory space address and the third memory space address respectively, read each second model from the memory space in sequence according to the execution order and the second memory space address of each second model, obtain a corresponding output according to the input of each second model, and use the output of the last second model as the processing result; The input of the first second model is the input information, and the input of the non-first second model is the output of the previous second model.
7. The method according to claim 6, characterized in that When the first model is decomposed, an execution file of the first model is also generated, and the execution file also includes the execution logic of the first model; Before obtaining the corresponding output according to the input of each second model, the method further includes the step of allocating memory space for temporary data of each second model: For each second model, determining temporary data of the second model based on the execution logic, and if it is determined based on the execution logic that the temporary data needs to be used by other second models, determining whether a preset mark has been added to the temporary data; wherein the temporary data is non-output data generated by the second model during execution; If the temporary data is not marked with a preset tag, the preset tag is added to the temporary data; wherein the preset tag indicates that memory space has been allocated for the temporary data; if the temporary data carries the preset tag, memory space is no longer allocated for the temporary data.
8. A model execution device, characterized in that: include: A model execution file storage module, used to store the input information of the first model, each second model and the execution order of each second model in the memory space respectively, wherein the data amount of the model parameter of the first model is greater than the first data amount, and the data amount of the model parameter of each second model is less than the first data amount, and the first data amount is the data amount that can be read at one time by the neural network processor NPU of the first model; each second model is obtained by decomposing the first model; an address sending module, configured to send the first memory space address of the input information, the second memory space addresses of each second model, and the third memory space address of the execution order to the NPU, so that the NPU reads the input information and the execution order from the memory space according to the first memory space address and the third memory space address, respectively, reads each second model from the memory space in sequence according to the execution order and the second memory space addresses of each second model, obtains a corresponding output according to the input of each second model, and takes the output of the last second model as a processing result; The input of the first second model is the input information, and the input of the non-first second model is the output of the previous second model.
9. A model execution device, characterized in that: include: An address receiving module is used to receive a first memory space address of input information of a first model, a second memory space address of each second model, and a third memory space address of an execution order of each second model sent by the CPU; wherein the input information, the second models, and the execution order are pre-stored in the memory space by the CPU; the data amount of the model parameters of the first model is greater than the first data amount, and the data amount of the model parameters of each second model is less than the first data amount, and the first data amount is the data amount that can be read by the NPU of the first model at one time; the execution order is used to indicate the execution order of each second model; each second model is decomposed from the first model; a model execution file reading module, configured to read the input information and the execution order from the memory space according to the first memory space address and the third memory space address, read each second model from the memory space in sequence according to the execution order and the second memory space address of each second model, obtain a corresponding output according to the input of each second model, and use the output of the last second model as a processing result; The input of the first second model is the input information, and the input of the non-first second model is the output of the previous second model.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.