Model operation method and device, electronic equipment and storage medium
By converting high-precision computing instructions into low-precision computing instructions, the problem of low TensorCore utilization of artificial intelligence models on low-end GPUs is solved, and the efficiency of model operation is improved.
Patent Information
- Application Number
- CN202510575429.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-19
AI Technical Summary
When artificial intelligence models are migrated from high-end GPUs to other GPUs, TensorCore utilization is too low, resulting in reduced operating efficiency.
Convert high-precision computing instructions to low-precision computing instructions to improve the utilization of TensorCore, specifically converting the floating-point precision from FP32 to FP16 or FP8, and optimizing the model's computing process through low-precision computing instructions.
The model's TensorCore utilization in low-end GPUs has been improved, which has enhanced overall operating efficiency, especially the significant performance improvement on GPUs in embedded systems such as RTX 3080 and Jetson Orin.
Smart Images

Figure CN120670027A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a model operation method, device, electronic device and storage medium. Background Art
[0002] In the field of artificial intelligence technology, some artificial intelligence models perform well on high-end GPUs. However, when they are migrated to other GPUs for operation, there is a technical problem of low utilization of the GPU's TensorCore (a hardware unit used to accelerate matrix operations), resulting in reduced operating efficiency. Summary of the Invention
[0003] In view of this, an object of the present disclosure is to provide a model operation method, device, electronic device and storage medium to improve the operation efficiency of the model.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for operating a model, the method comprising: triggering the operation process of the target model in response to an operation trigger instruction of the target model for the data to be operated; during the operation process, when the first precision operation instruction of the target model is triggered, converting the first precision operation instruction into a second precision operation instruction to obtain the instruction operation result of the second precision operation instruction; wherein, the second precision is lower than the first precision; based on the instruction operation result, obtaining the target operation result of the target model.
[0005] In the second aspect, an embodiment of the present disclosure provides a model computing device, which includes: a trigger module for triggering the computing process of the target model in response to the computing trigger instruction of the target model for the data to be computed; a conversion module for converting the first precision computing instruction into the second precision computing instruction when the first precision computing instruction of the target model is triggered during the computing process, and obtaining the instruction computing result of the second precision computing instruction; wherein the second precision is lower than the first precision; and an computing module for obtaining the target computing result of the target model based on the instruction computing result.
[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the operation method of the above-mentioned model.
[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the operation method of the above-mentioned model.
[0008] The embodiments of the present disclosure bring the following beneficial effects:
[0009] The calculation method, device, electronic device and storage medium of the above-mentioned model convert the high-precision calculation instructions of the model into low-precision calculation instructions, thereby improving the utilization of TensorCore of the high-precision model in a relatively low-end GPU, thereby improving the overall operation efficiency of the model.
[0010] Other features and advantages of the present disclosure will be described in the following description, and in part will become apparent from the description, or understood by practicing the present disclosure. The objectives and other advantages of the present disclosure are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0011] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 A flowchart of an embodiment of a method for calculating a model in an embodiment of the present disclosure;
[0014] Figure 2 A schematic diagram of a computing device for a model provided in an embodiment of the present disclosure;
[0015] Figure 3 A schematic diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0017] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of the present disclosure and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0018] For ease of understanding, the specific process of the embodiment of the present disclosure is described below. Figure 1 , an embodiment of the model operation method in the embodiment of the present disclosure includes:
[0019] Step S10, in response to the target model's operation triggering instruction for the data to be operated, triggering the target model's operation process;
[0020] The target model can be an artificial intelligence model of any structure. As an example but not limitation, the target model can be a model with an attention mechanism, such as a channel attention mechanism, a spatial attention mechanism, a self-attention mechanism, a category attention mechanism, a temporal attention mechanism, etc. The target model can also be a model with other mechanisms, such as a sequence modeling mechanism, a feature extraction mechanism, a generation and adversarial mechanism, a memory and external optimization mechanism, etc., which are not specifically limited here.
[0021] In one embodiment, the target model is an attention mechanism model accelerated by a Flash Attention mechanism (referred to as a Flash Attention model for short), and the target model runs on an image processor of an embedded system.
[0022] It is understandable that the FlashAttention model that is not executed according to the disclosed embodiments has low TensorCore utilization on some GPUs, especially GPUs running on embedded systems, such as RTX 3080 and Jetson Orin, whose TensorCore utilization is less than 50%. On Jetson Orin, when the head size is 32, the TensorCore utilization is less than 45% on average, and when the head size is 64, the TensorCore utilization is less than 60%. Low TensorCore utilization will lead to slower model inference speed.
[0023] The FlashAttention model executed by the embodiment of the present disclosure improves the performance by 1.79 times on the RTX 3080, while significantly improving the utilization of TensorCore. The average TensorCore utilization in the attention core is as high as 79%. On Jetson Orin, the performance when the attention head size is 32 is improved by 1.38 times, and the performance when the attention head size is 64 is improved by 1.11 times.
[0024] The data to be calculated refers to the data that can be used as input for the target model. Depending on the different functions of the target model, the data to be calculated will also be different. For example, assuming that the target model is a vehicle recognition model, then the data to be calculated can be an image of the vehicle to be recognized. If the target model is a semantic recognition model, then the data to be calculated can be speech or text of the semantics to be recognized. There is no specific limitation here.
[0025] When the target model is triggered to operate on the operation data to obtain the target operation result of the target model, the operation process of the target model is triggered, so that the target model performs the operation on the operation data according to its operation process to obtain the target result. For example, the target result of the vehicle recognition model is to identify the vehicle, and the target result of the semantic recognition model is to identify the semantics. The specific details are not limited here.
[0026] Step S20: During the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction to obtain an instruction operation result of the second precision operation instruction; wherein the second precision is lower than the first precision;
[0027] During the operation of the target model, when the first-precision operation instruction is triggered, the first-precision operation instruction is first converted into the second-precision operation instruction, and then the instruction operation result is obtained through the second-precision operation instruction, so that the operation result of the high-precision model is obtained through the lower-precision operation instruction, which improves the utilization rate of TensorCore by the high-precision model and thus improves the operation efficiency of the model.
[0028] Arithmetic instructions of different precision refer to calculation instructions with different number of bits of the value. Similarly, models of different precision also refer to models with different number of bits of the operation value. For example, precision can include floating-point precision and quantization precision. Floating-point precision can include double precision (FP64), single precision (FP32, TF32), half precision (FP16, BF16), 8-bit precision (FP8), and 4-bit precision (FP4, NF4). Quantization precision can include INT8, INT4, INT3, INT5, INT6, etc., which are not limited here.
[0029] In this embodiment, the first precision and the second precision can be any of the above-mentioned precisions, wherein the first precision is smaller than the second precision. For example, the first precision can be FP32, and the second precision can be FP16, FP8, etc., which are not specifically limited here. In one embodiment, the first precision and the second precision are both floating-point precisions, and the second precision is one level smaller than the first precision. For example, assuming that the first precision is FP32, then the second precision is the next level of precision below FP32, that is, FP16, which is not specifically limited here.
[0030] Operation instructions of different precisions can include any operation instructions in the operation process of the target model. Any operation instruction for the first precision value can be converted into an operation instruction for the second precision value, so that the operation of high-precision values can be converted into the operation of low-precision values, thereby improving the model's utilization of TensorCore and improving the model's operation efficiency.
[0031] As an example and not a limitation, the operation instructions may include matrix operation instructions, non-tensor computing core (TensorCore) operation instructions, single instruction multiple data (half2) operations, etc., which are not specifically limited here.
[0032] Among them, matrix operation instructions may include: matrix multiplication operation instructions, matrix addition operation instructions, matrix subtraction operation instructions, etc., and non-TensorCore operation instructions may include: activation function operation instructions, exponential function operation instructions, element-by-element floating multiplication and addition operation instructions, reduction operation instructions, cropping instructions, and scaling instructions, etc., which are not limited here.
[0033] When obtaining the instruction operation result of the second precision operation instruction, the data to be operated can be first converted into second precision data, and then the second precision data can be operated through the second precision operation instruction to obtain the instruction operation result of the second precision operation instruction, thereby improving the operation efficiency of the model.
[0034] Step S30: Obtain a target operation result of the target model based on the instruction operation result.
[0035] Based on the above-mentioned instruction operation results, when the operation process of the target model is completed, the target operation result of the target model can be obtained. Among them, the target operation results obtained by target models with different functions are also different. For example, the target operation result of the vehicle recognition model is vehicle information, and the target operation result of the semantic recognition model is semantic information. The specific details are not limited here.
[0036] The model calculation method provided by the above embodiment converts the high-precision calculation instructions of the model into low-precision calculation instructions, thereby improving the utilization of TensorCore of the high-precision model in a relatively low-end GPU, thereby improving the overall operation efficiency of the model.
[0037] Next, the specific operation method of the model is explained.
[0038] In one embodiment, the first precision operation instruction is a matrix multiplication operation instruction for the first precision; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the step of obtaining the instruction operation result of the second precision operation instruction includes: during the operation process, when the matrix multiplication operation instruction for the first precision is triggered, the target matrix to be multiplied is converted into a matrix of the second precision; the matrix of the second precision is multiplied and accumulated through the matrix multiplication and accumulation instruction of the second precision to obtain the instruction operation result of the second precision operation instruction.
[0039] The General Matrix Multiply (GEMM) instruction, also known as the general matrix multiplication instruction, is a core operation in linear algebra. Its basic form is to calculate the product of two matrices and add them to a third matrix.
[0040] During the operation of the target model, when the GEMM instruction of the first precision is triggered, the target matrix to be multiplied can be converted into a matrix of the second precision first, and then the converted second precision matrix can be subjected to matrix multiplication operation through the GEMM instruction of the second precision, so as to obtain the instruction operation result of the second precision operation instruction, thereby improving the performance of the target model.
[0041] For example, assuming that the precision of the target matrix to be multiplied is FP32, it can be first converted into a matrix with precision FP16, and then the FP16 GEMM operation is performed to obtain the FP16 instruction operation result. The specific details are not limited here.
[0042] In one embodiment, when multiplying and accumulating a matrix of second precision through a matrix multiplication and accumulation instruction of second precision to obtain the instruction operation result of the second precision operation instruction, it includes: multiplying the matrix of second precision through the matrix multiplication instruction of second precision, accumulating the product into the matrix of second precision, and obtaining the instruction operation result of the second precision operation instruction.
[0043] Taking the first precision of FP32 as an example, assume that the target matrices A and B with precision FP32 are first converted to FP16 precision, and then the calculation instruction of the FP16 accumulator is used as the matrix multiplication instruction of the second precision, A and B are multiplied and accumulated into the FP16 accumulator, thereby obtaining the FP16 instruction operation result.
[0044] In one embodiment, the first precision operation instruction is a non-tensor computing core operation instruction; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained, including: during the operation process, when the non-tensor computing core operation instruction is triggered, the data to be operated is converted into second precision data; through the second precision non-tensor computing core operation instruction, the second precision data is operated to obtain the instruction operation result of the second precision operation instruction.
[0045] During the calculation process, when the non-TensorCore operation instruction of the target model is triggered, the data to be calculated can be converted into second-precision data first, and then the second-precision non-TensorCore operation instruction is used to perform corresponding operations on the second-precision data, thereby obtaining the instruction operation result of the second-precision operation instruction, that is, the instruction operation result of the non-TensorCore operation instruction, thereby improving the performance of the model.
[0046] Among them, non-TensorCore operations include: activation function operations, exponential function operations, element-by-element floating multiplication and addition operations, reduction operations, clipping, and scaling. Specifically, non-TensorCore operations include all operations in the activation function softmax, involving exponential functions, and element-by-element multiplication and accumulation instructions (Fused Multiply–accumulate operation, FMA) and reduction operations (maximum and sum), which can be executed within a single thread and across multiple threads in a warp.
[0047] Non-TensorCore operations also include other non-softmax operations, such as cropping and scaling, so that the entire operation process of the FlashAttention model is optimized for second-precision execution, thereby improving the overall performance of the model.
[0048] In one embodiment, when obtaining the instruction operation result of the second precision operation instruction, it includes: converting the data to be operated corresponding to the second precision operation instruction into the data type corresponding to the single instruction multiple data instruction; performing single instruction multiple data operation on the data to be operated through the single instruction multiple data instruction to obtain the instruction operation result of the second precision operation instruction.
[0049] Single Instruction Multiple Data (HALF2) is an efficient data type and operation mode in GPU programming. It is mainly used to optimize the calculation and memory access of half-precision floating-point numbers (FP16). Its core idea is to pack two second-precision values into a first-precision unit and process two data simultaneously through a single instruction, thereby improving parallelism and memory bandwidth utilization.
[0050] When obtaining the instruction operation result of the second precision operation instruction, the instruction operation result is actually obtained through the second precision operation instruction, the half2 instruction, which can improve the performance of the program by reducing the use of memory bandwidth and increasing the computing density, especially when processing a large amount of half-precision data.
[0051] Corresponding to the above method embodiment, see Figure 2 A schematic diagram of a computing device of a model shown, the device includes: a trigger module 22, used to trigger the computing process of the target model in response to the computing trigger instruction of the target model for the data to be computed; a conversion module 24, used to convert the first precision computing instruction into the second precision computing instruction when the first precision computing instruction of the target model is triggered during the computing process, and obtain the instruction computing result of the second precision computing instruction; wherein the second precision is lower than the first precision; an computing module 26, used to obtain the target computing result of the target model based on the instruction computing result.
[0052] The computing device of the above model converts the high-precision computing instructions of the model into low-precision computing instructions, thereby improving the utilization of TensorCore of the high-precision model in a relatively low-end GPU, thereby improving the overall operating efficiency of the model.
[0053] Optionally, the first-precision operation instruction is a matrix multiplication operation instruction for the first precision; the above-mentioned conversion module 24 includes: a conversion unit, which is used to convert the target matrix to be multiplied into a matrix of the second precision when the matrix multiplication operation instruction for the first precision is triggered during the operation process; a multiplication and addition unit, which is used to multiply and accumulate the matrix of the second precision through the matrix multiplication and accumulation instruction of the second precision to obtain the instruction operation result of the second-precision operation instruction.
[0054] Optionally, the multiplication and addition unit is specifically used to: multiply the second precision matrix through the second precision matrix multiplication instruction, accumulate the product into the second precision matrix, and obtain the instruction operation result of the second precision operation instruction.
[0055] Optionally, the first precision operation instruction is a non-tensor computing core operation instruction; the above-mentioned conversion module 24 is used to: during the operation process, when the non-tensor computing core operation instruction is triggered, convert the data to be operated into second precision data; through the second precision non-tensor computing core operation instruction, operate on the second precision data to obtain the instruction operation result of the second precision operation instruction.
[0056] Optionally, the non-tensor computing core operation instructions include: activation function operation instructions, exponential function operation instructions, element-by-element floating multiplication and addition operation instructions, reduction operation instructions, cropping instructions, and scaling instructions.
[0057] Optionally, the step of obtaining the instruction operation result of the second precision operation instruction includes: converting the data to be operated corresponding to the second precision operation instruction into the data type corresponding to the single instruction multiple data instruction; performing a single instruction multiple data operation on the data to be operated through the single instruction multiple data instruction to obtain the instruction operation result of the second precision operation instruction.
[0058] Optionally, the target model is an attention mechanism model accelerated by a flash memory attention mechanism, and the target model runs on an image processor of an embedded system.
[0059] This embodiment further provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned model operation method. The electronic device can be a server or a terminal device.
[0060] See also Figure 3 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine executable instructions that can be executed by the processor 100. The processor 100 executes the machine executable instructions to implement the operation method of the above model.
[0061] Furthermore, Figure 3 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .
[0062] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0063] The processor 100 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 100 or by instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register, which is well-known in the art. The storage medium is located in the memory 101. The processor 100 reads the information in the memory 101 and, in conjunction with its hardware, performs the steps of the method of the aforementioned embodiment, for example:
[0064] In response to the operation trigger instruction of the target model for the data to be operated, the operation process of the target model is triggered; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained; wherein, the second precision is lower than the first precision; based on the instruction operation result, the target operation result of the target model is obtained.
[0065] In this method, by converting the model's high-precision computing instructions into low-precision computing instructions, the high-precision model's utilization of TensorCore in a relatively low-end GPU is improved, thereby improving the overall operating efficiency of the model.
[0066] Optionally, the first precision operation instruction is a matrix multiplication operation instruction for the first precision; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained. The step includes: during the operation process, when the matrix multiplication operation instruction for the first precision is triggered, the target matrix to be multiplied is converted into a matrix of the second precision; the second precision matrix is multiplied and accumulated through the second precision matrix multiplication and accumulation instruction to obtain the instruction operation result of the second precision operation instruction.
[0067] Optionally, the step of multiplying and accumulating the matrix of the second precision through the matrix multiplication and accumulation instruction of the second precision to obtain the instruction operation result of the second precision operation instruction includes: multiplying the matrix of the second precision through the matrix multiplication instruction of the second precision, accumulating the product into the matrix of the second precision, and obtaining the instruction operation result of the second precision operation instruction.
[0068] Optionally, the first precision operation instruction is a non-tensor computing core operation instruction; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained. The step includes: during the operation process, when the non-tensor computing core operation instruction is triggered, the data to be operated is converted into second precision data; the second precision data is operated by the second precision non-tensor computing core operation instruction to obtain the instruction operation result of the second precision operation instruction.
[0069] Optionally, the non-tensor computing core operation instructions include: activation function operation instructions, exponential function operation instructions, element-by-element floating multiplication and addition operation instructions, reduction operation instructions, cropping instructions, and scaling instructions.
[0070] Optionally, the step of obtaining the instruction operation result of the second precision operation instruction includes: converting the data to be operated corresponding to the second precision operation instruction into the data type corresponding to the single instruction multiple data instruction; performing a single instruction multiple data operation on the data to be operated through the single instruction multiple data instruction to obtain the instruction operation result of the second precision operation instruction.
[0071] Optionally, the target model is an attention mechanism model accelerated by a flash memory attention mechanism, and the target model runs on an image processor of an embedded system.
[0072] This embodiment further provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the operation method of the above-mentioned model, for example:
[0073] In response to the operation trigger instruction of the target model for the data to be operated, the operation process of the target model is triggered; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained; wherein, the second precision is lower than the first precision; based on the instruction operation result, the target operation result of the target model is obtained.
[0074] In this method, by converting the model's high-precision computing instructions into low-precision computing instructions, the high-precision model's utilization of TensorCore in a relatively low-end GPU is improved, thereby improving the overall operating efficiency of the model.
[0075] Optionally, the first precision operation instruction is a matrix multiplication operation instruction for the first precision; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained. The step includes: during the operation process, when the matrix multiplication operation instruction for the first precision is triggered, the target matrix to be multiplied is converted into a matrix of the second precision; the second precision matrix is multiplied and accumulated through the second precision matrix multiplication and accumulation instruction to obtain the instruction operation result of the second precision operation instruction.
[0076] Optionally, the step of multiplying and accumulating the matrix of the second precision through the matrix multiplication and accumulation instruction of the second precision to obtain the instruction operation result of the second precision operation instruction includes: multiplying the matrix of the second precision through the matrix multiplication instruction of the second precision, accumulating the product into the matrix of the second precision, and obtaining the instruction operation result of the second precision operation instruction.
[0077] Optionally, the first precision operation instruction is a non-tensor computing core operation instruction; during the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction, and the instruction operation result of the second precision operation instruction is obtained. The step includes: during the operation process, when the non-tensor computing core operation instruction is triggered, the data to be operated is converted into second precision data; the second precision data is operated by the second precision non-tensor computing core operation instruction to obtain the instruction operation result of the second precision operation instruction.
[0078] Optionally, the non-tensor computing core operation instructions include: activation function operation instructions, exponential function operation instructions, element-by-element floating multiplication and addition operation instructions, reduction operation instructions, cropping instructions, and scaling instructions.
[0079] Optionally, the step of obtaining the instruction operation result of the second precision operation instruction includes: converting the data to be operated corresponding to the second precision operation instruction into the data type corresponding to the single instruction multiple data instruction; performing a single instruction multiple data operation on the data to be operated through the single instruction multiple data instruction to obtain the instruction operation result of the second precision operation instruction.
[0080] Optionally, the target model is an attention mechanism model accelerated by a flash memory attention mechanism, and the target model runs on an image processor of an embedded system.
[0081] The computer program product of the model operation method, device, electronic device and storage medium provided in the embodiments of the present disclosure includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be found in the method embodiments and will not be repeated here.
[0082] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0083] In addition, in the description of the embodiments of the present disclosure, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in the present disclosure based on the specific circumstances.
[0084] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0085] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate the description of this disclosure and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0086] Finally, it should be noted that the above embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A model calculation method, characterized in that: The method comprises: In response to a calculation trigger instruction of the target model for the data to be calculated, triggering a calculation process of the target model; During the operation process, when the first precision operation instruction of the target model is triggered, the first precision operation instruction is converted into a second precision operation instruction to obtain an instruction operation result of the second precision operation instruction; wherein the second precision is lower than the first precision; Based on the instruction operation result, a target operation result of the target model is obtained.
2. The method according to claim 1, characterized in that The first precision operation instruction is a matrix multiplication operation instruction for the first precision; During the operation process, when the first precision operation instruction of the target model is triggered, the step of converting the first precision operation instruction into a second precision operation instruction and obtaining the instruction operation result of the second precision operation instruction includes: During the operation, when a matrix multiplication operation instruction for the first precision is triggered, the target matrix to be multiplied is converted into a matrix of the second precision; The matrix of the second precision is multiplied and accumulated by a matrix multiplication and accumulation instruction of the second precision to obtain an instruction operation result of the second precision operation instruction.
3. The method according to claim 2, characterized in that The step of performing multiplication and accumulation on the matrix of the second precision by using the matrix multiplication and accumulation instruction of the second precision to obtain the instruction operation result of the second precision operation instruction includes: The matrix of the second precision is multiplied by a matrix multiplication instruction of the second precision, and the product is added to the matrix of the second precision to obtain an instruction operation result of the second precision operation instruction.
4. The method according to claim 1, wherein The first precision operation instruction is a non-tensor calculation core operation instruction; During the operation process, when the first precision operation instruction of the target model is triggered, the step of converting the first precision operation instruction into a second precision operation instruction and obtaining the instruction operation result of the second precision operation instruction includes: During the operation process, when the non-tensor calculation core operation instruction is triggered, the data to be operated is converted into the second precision data; The second-precision data is operated by a second-precision non-tensor computing core operation instruction to obtain an instruction operation result of the second-precision operation instruction.
5. The method according to claim 4, characterized in that The non-tensor computing core operation instructions include: activation function operation instructions, exponential function operation instructions, element-by-element floating multiplication and addition operation instructions, reduction operation instructions, cropping instructions, and scaling instructions.
6. The method according to claim 1, characterized in that The step of obtaining the instruction operation result of the second-precision operation instruction includes: Converting the to-be-operated data corresponding to the second-precision operation instruction into a data type corresponding to a single-instruction-multiple-data instruction; A single instruction multiple data operation is performed on the data to be operated through the single instruction multiple data instruction to obtain an instruction operation result of the second precision operation instruction.
7. The method according to claim 1, characterized in that The target model is an attention mechanism model accelerated by a flash memory attention mechanism, and the target model runs on an image processor of an embedded system.
8. A model computing device, characterized in that: The device comprises: a trigger module, configured to trigger a computation process of the target model in response to a computation trigger instruction of the target model for the data to be computed; a conversion module, configured to, during a calculation process, when a first-precision calculation instruction of the target model is triggered, convert the first-precision calculation instruction into a second-precision calculation instruction, and obtain an instruction calculation result of the second-precision calculation instruction; wherein the second precision is lower than the first precision; The operation module is used to obtain the target operation result of the target model based on the instruction operation result.
9. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the operation method of the model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the operation method of the model described in any one of claims 1 to 7.
Citation Information
Patent Citations
Computer processor for higher precision computations using a mixed-precision decomposition of operations
CN110955404A
RISC-V universal processor supporting high-throughput multi-precision multiplication operation
CN112506468A
Propagating reduced-precision on computation graphs
US20200249924A1