Method, system and application for expanding tensor processor user-defined operator to realize new model deployment
By abstracting the hardware capabilities of tensor processors as basic instruction configuration methods, the problem that tensor processor firmware programs can only support a limited number of basic operators is solved, and efficient deployment of new models and maximum utilization of hardware capabilities is achieved.
Patent Information
- Application Number
- CN202410298317.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-05-13
AI Technical Summary
Existing tensor processor firmware programs can only support a limited number of basic operators, resulting in the inadmissible operators in the new model being unable to deploy and the hardware capabilities cannot be maximized.
By abstracting the tensor processor hardware capabilities into the basic instruction configuration methods in a firmware program, all operators can be implemented in combination with these methods, so that custom operators can be supported without modifying the firmware program.
The ability to deploy new models without converting basic operator combinations or modifying firmware programs is realized, ensuring the execution efficiency and effectiveness of the model, and reducing the workload of maintaining the consistency of the application and firmware version.
Smart Images

Figure HDA0004743017200000011 
Figure HDA0004743017200000012 
Figure HDA0004743017200000021
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of tensor processors, and relates to a method, system and application for extending a tensor processor to customize an operator and realize new model deployment. Background Art
[0002] The tensor processing unit (TPU) is a special chip customized by Google for machine learning. It is designed specifically for Google's deep learning framework TensorFlow and has a wide range of applications in the field of artificial intelligence. Some domestic companies, such as Suanneng Technology, have also launched their own tensor processing chips for the field of deep learning.
[0003] Operators are the basic computing units in the deep learning programming framework. They generally process tensor-type data. The models and training algorithms in the framework are implemented by the underlying operators. The richness of the operators determines the richness and completeness of the deep learning framework functions. The models in the framework ultimately need to be deployed on hardware to complete inference operations. Therefore, the more basic operators provided by tensor processor chip products, the better the support for deep learning frameworks, and the more types of models that can be deployed.
[0004] In the current common tensor processor architecture, the chip is usually controlled by a microcontroller, and the firmware program in the microcontroller completes specific functions such as sending commands to the chip's internal engine and driving communication, thereby realizing the specific logic of the operator. In other words, when the hardware is determined, the number of operators that can be executed is determined by the firmware program in the microcontroller.
[0005] Limited by the size of the program and the types of operators commonly used when writing programs, only a limited number of basic operators can be implemented in the firmware program. With the rapid development of artificial intelligence, the speed at which new models appear is accelerating, and the number of operators in the models is also increasing. This may result in the hardware supporting the operators but not being able to run them, making it impossible to maximize the hardware capabilities.
[0006] There are currently two main solutions to this problem.
[0007] One is to convert operators that are not supported by the current tensor processor in the model into a combination of one or more basic operators that can output the same results or results within the error range. However, the original operator in the model may be the most efficient and effective implementation method. Converting the operator may reduce the execution efficiency of the operator, affect the output effect of the model, and there is a possibility that the conversion cannot be performed.
[0008] Another method is to add support for new operators by modifying the firmware program, but many chip manufacturers do not provide firmware program source code. Every time an operator is added, it is necessary to find the chip manufacturer for customized modification, which is obviously unrealistic. For this reason, some manufacturers provide a way to add custom operators. For example, Suanneng Technology provides a method to link custom operators with the base library it provides to form an original dynamic library file that can be loaded by its internal controller and called in the host application (TPU working mode—TPU-KERNEL Development Reference Manual Document (sophgo.com)). However, this method requires modifying the programs on both the host and microcontroller sides at the same time, and it is necessary to maintain the version matching problem between the application and the dynamic library. At the same time, due to the limitation of the program size, it may be impossible to add new operators.
[0009] Therefore, there are several main problems in the existing technology: 1. The TPU firmware program can only support a limited number of basic operators. When a new model appears, the model may not be deployed due to the lack of running operators even though the hardware capabilities support it;
[0010] 2. Deploying operators that cannot be supported in the model by converting them into basic operator combinations may reduce the efficiency of operator execution, affect the model effect, and may even make the conversion impossible.
[0011] 3. By modifying the firmware program, the firmware program needs to be recompiled, and additional effort needs to be spent to ensure that the application program and the firmware program version match. Summary of the invention
[0012] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a method for extending the tensor processor to customize operators and implement new model deployment. The method of the present invention abstracts the hardware capabilities of the tensor processor into basic instruction configuration methods in the firmware program, and all operators are implemented by combining these methods, thereby solving the three problems mentioned in the prior art.
[0013] In the prior art, TPU firmware programs can only support a limited number of basic operators. Since all operators in the method of the present invention can be obtained by combining the basic instruction configuration method, there is no problem of limited number. As long as the hardware capabilities support it, it can be deployed. However, converting operators that cannot be supported in the model into basic operator combinations will reduce execution efficiency, affect model effects, and require recompiling the firmware program by modifying the firmware program. These are all imperfections in some existing methods for processing the limited number of basic operators supported by TPU firmware programs in the prior art. This method directly solves the root cause of the limited number of basic operators supported by TPU firmware programs in the prior art.
[0014] The present invention provides a method for implementing new model deployment by extending a tensor processor custom operator. The method does not require modification of the firmware program. By abstracting the hardware capabilities of the tensor processor into a basic instruction configuration method in the firmware program, all operators are allowed to implement the calculation logic of the operator by calling one or more of the basic instructions. An instruction configuration method sequence of the custom operator is generated through a library function interface provided by the host side, and the instruction configuration method sequence is sent to the firmware program through a software interface. The firmware program parses the instruction configuration method sequence and drives the TPU hardware to perform computing tasks, thereby implementing the deployment of the new model.
[0015] In the present invention, the hardware capabilities of the TPU are abstracted as a basic instruction configuration method in a firmware program, which means that its hardware functions are described as a series of basic instructions or operations, which can be used alone or in combination to implement various tensor computing tasks using the underlying hardware of the TPU.
[0016] The parameters in the basic instruction configuration method are defined and described by a structure, which can describe the format of the custom operator instruction sequence. The basic structure of the structure includes the custom operator name, the number of instructions, the instruction length, the cyclic redundancy check code CRC, one or more instruction codes and one or more instruction parameters corresponding to the instruction code;
[0017] in,
[0018] Custom operator name: used to identify the name of the operator, which is convenient for users and firmware programs to identify and call;
[0019] Number of instructions: indicates the number of basic instructions contained in this custom operator, which is used to inform the firmware program of the number of basic instructions that need to be parsed and executed;
[0020] Instruction length: represents the number of bytes occupied by the instruction code and instruction parameters that follow it, and is used by the firmware program to correctly read and parse the instruction sequence;
[0021] CRC: Cyclic Redundancy Check Code, used to detect whether the instruction sequence is damaged during transmission and ensure the integrity of the data;
[0022] Instruction code and instruction parameters: Each instruction code corresponds to a specific operation, and the instruction parameters are specific parameters required to perform the operation. The instruction code combined with the instruction parameters defines the specific behavior of the custom operator;
[0023] In addition, a custom operator instruction sequence can include multiple sets of instruction codes and instruction parameters to support complex custom operators.
[0024] In the present invention, it is necessary to reasonably abstract the hardware capabilities of the TPU, which is a method of simplifying complex systems, processes or problems into core concepts and operations that are easier to understand and operate. By hiding unnecessary details and emphasizing important properties and behaviors, it allows developers to design, understand and maintain systems more efficiently and flexibly without sacrificing functionality and performance. Reasonable abstraction should balance simplification and practicality to ensure that it can provide sufficient information and functionality without being overly complicated.
[0025] Specifically, the abstract method described in the present invention is an experience in the actual development process. There are no specific rules and constraints. The custom operator is decomposed according to the fine-grainedness actually required, and then the decomposed instructions are combined and configured to realize the execution of tasks and the deployment of new models.
[0026] The following is an explanation of the operator for converting an RGB image to a YUV image of type UINT8.
[0027] This is just a simple example and cannot fully reflect the actual workflow of the hardware.
[0028] Y=0.299R+0.587G+0.114B
[0029] U=-0.147R'-0.289G'+0.436B
[0030] V=0.615R'-0.289G'+0.436B
[0031] To drive the TPU to complete the operator, the following operations are required:
[0032] 1. Use DMA to copy the data of the three channels R, G, and B from the system memory to the TPU local memory;
[0033] 2. Multiply the tensor element of the R channel by the conversion factor;
[0034] 3. The tensor element of the G channel is multiplied by the conversion factor and added to the result of the previous step;
[0035] 4. Multiply the tensor element of channel B by the conversion factor and add it to the result of the previous step to get Y;
[0036] 5. Repeat steps 2 to 4 to obtain U and V;
[0037] 6. Use DMA to copy the results from TPU local memory to system memory.
[0038] Each step here can correspond to a kind of TPU hardware capability described by instructions mentioned above. If not in this granularity, it can also be divided at other granularities, such as finer granularity division, using DMA to copy the data of the three channels of R, G, and B from the system memory to the TPU local memory can be subdivided into DMA setup and startup, and the multiplication of tensor elements and conversion factors can be subdivided into data type conversion (UINT8 to FLOAT32) and tensor multiplication, or coarse granularity division, dividing 2 to 4 steps into one step.
[0039] In the method of the present invention, in addition to generating a custom operator by combining one or more basic instruction configuration methods, basic operators in common models are also provided. When there is no need for a custom operator, the basic operator can be directly called through the library function interface for calculation.
[0040] In the present invention, the creation and generation of the custom operator includes the following steps:
[0041] Step 1: Call the custom operator initialization interface to create and initialize the custom operator descriptor;
[0042] Step 2: Configure the instruction structure of the custom operator;
[0043] Step 3: Call the instruction adding interface to add the instruction to the operator descriptor;
[0044] Step 4: Determine whether there are any more instructions to add. If it is the last instruction, call the custom operator sequence generation interface. The interface function will generate the operator instruction sequence according to the description of the custom operator descriptor and return the address and length of the sequence. If there are still instructions to be added, return to step 2.
[0045] The present invention also provides a system for implementing the above method, the system is divided into a host side and a microcontroller side, and comprises:
[0046] Host-side application module: This module is responsible for creating custom operator instruction sequences. It uses the library function interface to generate corresponding operators according to user requirements and sends these operator instruction sequences to the firmware program on the microcontroller side through inter-core communication.
[0047] Library function interface module: includes one or more APIs for host-side applications to call. It abstracts the TPU hardware capabilities, allowing users to describe custom operators through structures, and then generate instruction configuration method sequences.
[0048] Custom operator generation module: On the host side, this module specifically implements the generation of custom operators. It decomposes and combines instruction configuration methods to generate operator instruction sequences based on user requirements and the instruction set supported by the TPU hardware.
[0049] Microcontroller firmware program module: The firmware program receives the instruction sequence from the host, performs verification and analysis, and drives the TPU hardware to execute operators. It includes the implementation of basic operators and provides them to users through library function interfaces. When the user's operator is not within the scope of the basic operator, this module is responsible for processing custom operators.
[0050] Inter-core communication module: This module exchanges data between the host side and the microcontroller side, handling everything from the creation and sending of custom operators to the reception of execution results.
[0051] The present invention also provides a hardware system for implementing the above method, the hardware system comprising: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above method is implemented.
[0052] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.
[0053] The present invention also provides applications of the above-mentioned method, the above-mentioned system, the above-mentioned hardware system, or the above-mentioned computer-readable storage medium in deep learning model deployment, customized operator development, cross-domain model adaptation, etc.
[0054] The beneficial effects of the present invention include: when deploying a model with operators that are not supported by the current TPU firmware program, there is no need to convert the operators, ensuring maximum similarity with the original model; there is no need to recompile the firmware, and there is no need to spend extra effort to maintain the consistency of the application and firmware versions.
[0055] Specifically, the present invention aims to solve the problem in the prior art that TPU firmware programs can only support a limited number of basic operators. In traditional technology, if the operators in the new model are not supported by the TPU firmware, these operators usually need to be converted into a combination of basic operators for deployment. Such conversion may reduce execution efficiency and affect the output effect of the model. In addition, for some operators that cannot be converted, they may not be deployed, resulting in failure to fully utilize the hardware capabilities.
[0056] The present invention abstracts the TPU hardware capabilities into basic instruction configuration methods in the firmware program, so that all operators are implemented by calling one or more such instruction configuration methods. In this way, even operators that are not supported by the current TPU firmware program can be implemented through a user-defined instruction configuration method sequence without converting them into basic operators or modifying the firmware program. This not only ensures the maximum similarity between the newly deployed model and the original model, but also avoids the need to recompile the firmware, while reducing the workload required to maintain the consistency of the application and firmware versions.
[0057] Users can generate a custom operator instruction sequence based on their needs through the provided library function interface, and then send this sequence to the firmware program through the software interface. The firmware program is responsible for parsing these sequences and driving the TPU hardware to perform the corresponding computing tasks. Therefore, this method greatly simplifies the deployment process of new models, improves efficiency and flexibility, and allows users to maximize the computing power of the hardware. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 A flow diagram showing the application's control of the TPU hardware.
[0060] Figure 2 Indicates the process of creating and calling a custom operator.
[0061] Figure 3 Indicates the detailed process of creating a custom operator.
[0062] Figure 4 Indicates the format of a custom operator instruction sequence.
[0063] Figure 5 Represents the test case output of a custom operator.
[0064] Figure 6 Indicates the steps of the model in data processing on TPU. DETAILED DESCRIPTION
[0065] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.
[0066] The present invention provides a method for extending a tensor processor custom operator to implement new model deployment. The system running the method includes an application running on the host side, a library function interface that can be called by the user, and a driver and a firmware program running on a microcontroller. The method abstracts the tensor processor capability into a basic instruction configuration method in the firmware program, and all operators implement their calculation logic by calling one or more instruction configuration methods. When the user needs to add a custom operator, the instruction configuration method sequence of the custom operator can be generated through the library function interface, and then the sequence can be sent to the firmware program through the software interface, and the firmware program drives the tensor processor hardware to implement it. The method can expand the custom operators supported by the tensor processor hardware without modifying the firmware program, making it convenient for the tensor processor to deploy new models.
[0067] As used in the present invention, the terms "comprise" and "include" are open expressions, that is, they include the contents specified in the present invention but do not exclude other contents.
[0068] Specifically, the present invention proposes a method for extending TPU custom operators and implementing new model deployment without modifying the firmware program. The specific scheme is as follows:
[0069] 1. Abstract the hardware capabilities of the tensor processor chip into a basic instruction configuration method in the firmware program. The basic instruction configuration method is used as the basic unit of tensor processor configuration. All operators are implemented by calling one or more of the basic instruction configuration methods.
[0070] The parameters required by each basic instruction configuration method are described by a structure. The user decomposes the custom operator to be configured into one or more basic instruction configuration methods, configures the instruction structure, calls one or more basic operators through the library function interface, and generates an instruction sequence of a preset or custom operator;
[0071] The structure is configured according to the different conditions of each operator. Its basic structure is as shown in the attached figure. Figure 4 As shown, Figure 4 The basic structure of the structure includes a user-defined operator name, the number of instructions, an instruction length, a cyclic redundancy check code CRC, one or more instruction codes, and one or more instruction parameters corresponding to the instruction code;
[0072] 2. Send the generated command sequence of the preset or custom operator to the firmware program through the library function interface;
[0073] 3. After receiving the instruction sequence of the preset or custom operator, the firmware program verifies and parses the sequence, calls the instruction configuration method in the sequence, implements the specific calculation function, and completes the deployment of the new model.
[0074] In addition to the custom operators mentioned above, the method in the present invention also provides basic operators in common models. When the model that the user needs to deploy is composed of basic operators and does not need to use custom operators, it can be directly called through the library function interface provided by this method to meet your needs.
[0075] In general, the technical solution in the present invention is to further decompose the operator into instructions of the TPU hardware capabilities, so as to realize the addition of user-defined operators; specifically, the operator is decomposed according to the fine granularity actually required, and then the decomposed instructions are combined and configured to realize the execution of tasks and the deployment of new models. The method of the present invention requires a reasonable abstraction of the TPU hardware capabilities, how to abstract into sufficiently fine particles, so as to realize the addition of any custom operator, while not making it too trivial and affecting the user experience.
[0076] The method of the present invention is not targeted at specific application scenarios, but rather tends to expose hardware capabilities to users to the maximum extent possible. All applications can add custom operators based on their actual conditions.
[0077] Figure 1 A simple functional block diagram showing an application controlling TPU hardware, where the application, library, and driver run on the host side. The driver refers to the kernel driver, which directly communicates with the firmware program of the microcontroller MCU to control the TPU hardware; the library is a low-level software library encapsulated on the driver, providing an external interface for user development; the application is an application developed by the user.
[0078] The firmware runs on the microcontroller MCU and directly controls the TPU hardware to realize its computing functions. The host and microcontroller programs exchange data through inter-core communication.
[0079] The firmware program also provides the implementation of common basic operators and provides them to users through library function interfaces. When the operators used by the user's application are within the scope of the provided basic operators, the corresponding library function interface can be directly called to control the TPU to perform operations. When the operator required by the user is not within the scope of the basic operator, it is necessary to create a custom operator to realize its own computing needs.
[0080] Figure 2 Indicates the process of creating and calling a custom operator when used by the user. The detailed process is as follows:
[0081] 1. The application creates a custom operator through the library function interface and obtains the address and size of the custom operator instruction sequence;
[0082] 2. The application calls the instruction sending interface and sends the custom operator instruction sequence generated in the previous step to the firmware program through inter-core communication;
[0083] 3. After receiving the command sent by the application, the firmware program verifies and analyzes the custom operator instruction sequence and executes the instruction after confirming that the instruction is correct;
[0084] 4. After the firmware program is executed, it will notify the host end of the completion of the operation through inter-core communication;
[0085] 5. After receiving the notification, the application calls the library function interface to obtain the calculation result.
[0086] Through the above steps, the present invention achieves the purpose of directly creating and calling a custom operator in an application program to complete the required operation without modifying the firmware program.
[0087] Figure 3 This shows the detailed creation process of a custom operator. When creating a custom operator, the user needs to combine the instructions for all instructions and parameters supported by the TPU hardware, decompose the operator into an instruction configuration method that is executed in sequence, and then generate it according to the following steps:
[0088] 1. Call the custom operator initialization interface to initialize the custom operator descriptor. The custom operator descriptor structure is as follows:
[0089] struct tpu_op_desc{ / / Define a structure named tpu_op_desc to describe a custom operator struct tpu_ins_node*head; / / In the custom operator descriptor, define a pointer head pointing to the tpu_ins_node type, which points to the beginning of the instruction sequence
[0090] struct tpu_ins_node*tail; / / Define a pointer tail pointing to the tpu_ins_node type, which points to the end of the instruction sequence
[0091] uint16_t ins_cnt; / / define an unsigned short integer variable ins_cnt to store the number of instructions in the operator
[0092] uint16_t op_size; / / Define an unsigned short integer variable op_size to indicate the number of bytes occupied by the custom operator
[0093] char op_name[MAX_OP_NAME_LEN]; / / Define a character array op_name to store the name of the custom operator. MAX_OP_NAME_LEN defines the maximum length of the name
[0094] }tpu_op_desc_t; / / End the definition of the structure and name the structure type tpu_op_desc_t
[0095] The descriptor describes the name of the custom operator, the size of the custom operator and the number of instructions, and maintains the instructions in it through a linked list;
[0096] 2. Configure the instruction structure and call the instruction adding interface to add the instruction to the operator descriptor. The instruction structure is as follows:
[0097] struct tpu_ins_node{ / / Define a structure called tpu_ins_node to describe an instruction node in the operator
[0098] struct tpu_ins_node*next; / / Define a pointer next pointing to the next tpu_ins_node type to form a linked list structure
[0099] uint16_t ins_size; / / define an unsigned short integer variable ins_size to store the size of the instruction
[0100] void*ins_param; / / Define a pointer ins_param pointing to any type to point to the parameter of the instruction
[0101] }tpu_ins_node_t; / / End the definition of the structure and name the structure type tpu_ins_node_t
[0102] Each time this function is called, the instruction is added to the linked list of custom operator descriptors, the number of instructions increases by 1, and the operator size increases by the size of the instruction;
[0103] 3. Call the custom operator sequence generation interface. The interface function will generate the operator instruction sequence according to the description of the custom operator descriptor and return the address and length of the sequence;
[0104] Apply for space of the operator size in the operator descriptor, traverse the instructions in the operator descriptor in sequence, and copy the instruction structure data to the applied space until the traversal is completed.
[0105] When creating a custom operator, the tpu_op_desc structure is used to store the overall description of the custom operator, such as its name, size, and number of instructions. The tpu_ins_node structure is used to describe each instruction and its parameters in detail, and connect all instructions in series in the form of a linked list. When the custom operator is initialized, these structures will be filled with the corresponding data, and finally generate an instruction sequence that can be parsed by the firmware program and executed on the TPU hardware.
[0106] Figure 4 Indicates the operator instruction sequence format. For the convenience of description, the format is simplified, which is the operator name, the number of instructions in the operator N, the instruction length, CRC, N instruction codes and their corresponding instruction parameters. The instruction length represents the number of bytes occupied by the N instruction codes and their corresponding instruction parameters, and CRC represents the cyclic redundancy check calculated by the N instruction codes and their corresponding instruction parameters to ensure the correctness of the received command.
[0107] Each instruction can have multiple different parameters, which need to be used according to the actual situation.
[0108] Figure 5 It represents the output of the simulation test case that creates and calls a custom operator. The two programs running on the PC simulate the running programs on the host side and the microcontroller side respectively. The message queue and shared memory are used between the two programs to simulate the inter-core communication in actual use.
[0109] For the convenience of demonstration, in this embodiment, when the host creates a custom operator, two identical test instructions with different test data are added. From the results in the figure, it can be seen that the key program on the microcontroller side (MCU side) can parse the commands passed from the host side (Host side) and output the correct results, verifying the feasibility of the solution.
[0110] Figure 5 This is the output after the program is executed. This is just a simple example.
[0111] Since this is a simple example of a simulation program, only a structure is defined here
[0112] struct tpu_ins_test{ / / Define a structure called tpu_ins_test, which is used to represent a test instruction
[0113] uint16_t ins_code; / / ins_code in the structure is an unsigned short integer used to store the instruction code to identify what specific operation or calculation is performed
[0114] uint32_t test_data; / / test_data in the structure is an unsigned integer used to store test data or parameters related to the instruction, representing the actual data to be calculated, or other information that needs to be passed to the instruction
[0115] }
[0116] The host generates a custom operator according to the above steps and sends it to the microcontroller. After parsing, the microcontroller prints out the data of the operator generated by the host.
[0117] Its function is only to print out the data on the host side, but it is enough to prove the feasibility of this method. That is, after the custom operator generated by the host side is passed to the microcontroller side, the microcontroller side can parse and execute it. Complex operators only require more steps.
[0118] Figure 5 Further explanation is as follows:
[0119] Host side:
[0120] this is host: Indicates that these outputs come from the host.
[0121] init custom op named custom test op: Initialize a custom operation named "custom test op".
[0122] add test ins 0: Add a test instruction with instruction code 0.
[0123] test data 5: Set the data value to 5 for the test instruction with instruction code 0.
[0124] add test ins 1: Add another test instruction with instruction code 1.
[0125] test data 8: Set the data value to 8 for the test instruction of instruction code 1.
[0126] The host creates two test instructions, each with a unique instruction code and related test data.
[0127] Microcontroller side:
[0128] This is firmware: Indicates that these outputs come from the firmware.
[0129] custom op name is custom test op: Confirms that the firmware received the name of the custom operation.
[0130] instruction 0: The firmware acknowledges receipt of the first instruction.
[0131] tpu ins test data is 5: shows that the test data associated with instruction 0 is correctly received as 5.
[0132] instruction 1: The firmware acknowledges receipt of the second instruction.
[0133] tpu ins test data is 8: shows that the test data associated with instruction 1 is correctly received as 8.
[0134] The firmware output confirms that it received the commands sent by the host and correctly parsed those commands and the associated test data.
[0135] Figure 5 The execution results show that the communication protocol between the host and the firmware is valid. The host generates custom operations (here, test instructions with specific data), and the firmware receives and confirms these operations, demonstrating a successful communication handshake between the two ends. The struct tpu_ins_test structure is used to encapsulate the instruction code and related data, so that the firmware can parse and execute the custom operation, thereby verifying the feasibility of the method described in the present invention for creating and transmitting custom instructions to the TPU.
[0136] Example
[0137] This embodiment provides a method for deploying a model containing a custom operator on a TPU, such as Figure 6 As shown, the method comprises the following steps:
[0138] 1. Get the model that needs to be deployed to TPU;
[0139] 2. Analyze the operators of the model to be deployed and obtain custom operators that are not supported by inherent operators;
[0140] 3. Analyze the calculation method of the custom operator and decompose it into the instruction set supported by TPU;
[0141] 4. The host application calls the library function interface to generate a custom operator. The specific calling method refers to the description of the method of the present invention;
[0142] 5. The host application sends the generated custom operator to the microcontroller;
[0143] 6. The microcontroller drives the TPU hardware to perform calculations;
[0144] 7. After the TPU calculation is completed, it notifies the microcontroller that the calculation is completed;
[0145] 8. The microcontroller returns the calculation results to the host application.
[0146] Comparative Example
[0147] Provide a comparative example, refer to "TPU working mode - TPU-KERNEL development reference manual document (sophgo.com)"
[0148] Compared with the embodiment of the present invention, this method needs to decompose steps 4 and 5 of the embodiment into the following steps:
[0149] 1. Develop the device-side code of the custom operator in device_*.c;
[0150] 2. The device-side code compiler links the device-side code with the original base library libbm1684x.a to form a complete A53Lite loadable original dynamic library file;
[0151] 3. Load dynamic library;
[0152] 4. Send the call instruction of the custom operator on the host side;
[0153] It can be seen that compared with the comparative embodiment, the method mentioned in the present invention is simpler and only requires modifying the host-side code, without maintaining the host-side and device-side codes at the same time, thus reducing the amount of tools for code maintenance.
[0154] In addition to implementing the client and server in a purely computer-readable program code, the client and server can also implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a client and server can be considered as a hardware component, and the means for implementing various functions included therein can also be considered as a structure within the hardware component. Or even, the means for implementing various functions can be considered as both a software module for implementing the method and a structure within the hardware component.
[0155] It can be seen from the above description of the implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation method of the present application or some parts of the implementation method.
[0156] Each implementation in this specification is described in a progressive manner, and the same or similar parts between the various implementations can be referred to each other, and each implementation focuses on the differences from other implementations. In particular, for the implementation of the client and the server, both can refer to the introduction of the implementation of the aforementioned method for comparative explanation.
[0157] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0158] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.
Claims
1. A method for extending a tensor processor to customize an operator and implement new model deployment, characterized in that: The method abstracts the hardware capabilities of the tensor processor into a basic instruction configuration method in a firmware program, allowing all operators to implement the calculation logic of the operator by calling one or more of the basic instructions; generates an instruction configuration method sequence of the custom operator through a library function interface provided by the host side, and sends the instruction configuration method sequence to the firmware program through the library function interface. The firmware program parses the instruction configuration method sequence and drives the tensor processor hardware to perform calculation tasks, thereby realizing the deployment of a new model.
2. The method according to claim 1, characterized in that Abstracting the hardware capabilities of the tensor processor into a basic instruction configuration method in a firmware program refers to describing the hardware functions of the tensor processor as one or more basic instructions or operations, which are used alone or in combination to implement various tensor computing tasks using the underlying hardware of the tensor processor; Abstracting the hardware capabilities of the tensor processor means decomposing the custom operators according to the fine-grained level actually required, and then combining and configuring the decomposed instructions to implement task execution and deployment of new models.
3. The method according to claim 1, characterized in that The parameters in the basic instruction configuration method are defined and described by a structure, and the basic structure of the structure includes a custom operator name, the number of instructions, the instruction length, a cyclic redundancy check code CRC, one or more instruction codes, and one or more instruction parameters corresponding to the instruction code; The custom operator name is used to identify the name of the operator, which is convenient for users and firmware programs to identify and call; The number of instructions: indicates the number of basic instructions contained in this custom operator, which is used to inform the firmware program of the number of basic instructions that need to be parsed and executed; The instruction length: represents the number of bytes occupied by the instruction code and instruction parameters, which is used for the firmware program to correctly read and parse the instruction sequence; The cyclic redundancy check code is used to detect whether the instruction sequence is damaged during the transmission process to ensure the integrity of the data; The instruction code and instruction parameters: each instruction code corresponds to a specific operation, and the instruction parameters are specific parameters required to perform the operation. The instruction code combined with the instruction parameters defines the specific behavior of the custom operator.
4. The method according to claim 1, characterized in that The method can directly call the basic operator for the calculation of non-customized operators through the library function interface.
5. The method according to claim 1, characterized in that The creation and generation of the custom operator includes the following steps: Step 1: Call the custom operator initialization interface to create and initialize the custom operator descriptor; Step 2: Configure the instruction structure of the custom operator; Step 3: Call the instruction adding interface to add the instruction to the operator descriptor; Step 4: Determine whether there are any more instructions to add. If it is the last instruction, call the custom operator sequence generation interface. The interface function will generate the operator instruction sequence according to the description of the custom operator descriptor and return the address and length of the sequence. If there are still instructions to be added, return to step 2.
6. The method according to claim 1, characterized in that When a user creates and calls a custom operator, the process is as follows: I. The application creates a custom operator through the library function interface and obtains the address and size of the custom operator instruction sequence; II. The application calls the instruction sending interface and sends the custom operator instruction sequence generated in the previous step to the firmware program through inter-core communication; III. After receiving the command sent by the application, the firmware program verifies and analyzes the custom operator instruction sequence and executes the instruction after confirming that the instruction is correct; IV. After the firmware program is executed, it will notify the host end of the completion of the operation through inter-core communication; V. After receiving the notification, the application calls the library function interface to obtain the calculation result.
7. A system for implementing the method according to any one of claims 1 to 6, characterized in that: The system is divided into a host side and a microcontroller side, including a host side application program module, a library function interface module, a custom operator generation module, a microcontroller side firmware program module, and an inter-core communication module; The host application module is responsible for creating a custom operator instruction sequence and using a library function interface to generate corresponding operators according to user requirements; The library function interface module includes one or more APIs, abstracts the hardware capabilities of the tensor processor, and allows users to describe custom operators through structures, thereby generating a sequence of instruction configuration methods; The custom operator generation module decomposes and combines instruction configuration methods according to user requirements and instruction sets supported by TPU hardware to generate operator instruction sequences; The microcontroller-side firmware program module receives the instruction sequence from the host side, performs verification and analysis, and drives the tensor processor hardware to execute the operator; The inter-core communication module performs data exchange between the host side and the microcontroller side, and handles everything from the creation and sending of custom operators to the reception of execution results.
8. A hardware system for implementing the method according to any one of claims 1 to 6, characterized in that: The hardware system comprises: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. Application of the method according to any one of claims 1 to 6, the system according to claim 7, the hardware system according to claim 8, or the computer-readable storage medium according to claim 9 in deep learning model deployment, customized operator development, and cross-domain model adaptation.