An operation unit lock control method, device and electronic device

By executing the to-authorized instructions and the to-work end instructions in the neural network processor, locking and unlocking of the computing unit is realized, which solves the problem of waste of computing power and low flexibility in the time division multiplexing method, and improves the computing efficiency and resource utilization efficiency of the NPU.

CN119883649BActive Publication Date: 2025-06-13AXERA SEMICON (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510360496.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-13
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When multiplexing the computing unit through time division multiplexing method in existing neural network processors (NPUs), the computing power is wasteful and flexibility is low, thereby reducing the chip's computing efficiency and resource utilization efficiency.

Method used

A method for locking the operation unit is provided. By executing the authorized instructions and the end-of-work instructions, it determines whether the target vNPU can obtain the control rights of the operation unit. After the target vNPU releases the control rights, other vNPUs are allowed to obtain the control rights, thereby realizing locking and unlocking the operation unit.

Benefits of technology

The idle time of the computing unit is avoided, and the utilization rate of the computing unit is improved, thereby improving the working efficiency of the NPU.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883649B_ABST
    Figure CN119883649B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of neural network processors, and provides a method, device, and electronic device for controlling an operation unit lock. The method includes: in response to an authorization instruction to be processed, obtaining the computing power of a target virtual neural network processor that the allocator is running; if the computing power of the target virtual neural network processor meets a preset condition, enabling the target virtual neural network processor to obtain the control right for at least one operation unit; in response to an instruction indicating the end of work, obtaining the working state of at least one operation unit; if the working state of all at least one operation unit is the end of work, releasing the control right of the target virtual neural network processor for at least one operation unit. This method locks the operation unit by executing the authorization instruction to be processed; unlocks the operation unit by executing the instruction indicating the end of work, improves the utilization rate of the operation unit, and further improves the working efficiency of the neural network processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of neural network processors, and in particular, to a method and apparatus for controlling an operation unit lock and an electronic device. Background Art

[0002] In the application scenario of current Neural - Processing Unit (NPU) chips, there is a need for multiple users to run simulations simultaneously. The hard - slicing method can be used to achieve this, but the hard - slicing implementation requires multiple copies of operation units and memories, resulting in an increase in chip area and manufacturing cost. Therefore, the time - division multiplexing method can be used to achieve the multiplexing of operation units.

[0003] The current time - division multiplexing method is that multiple virtual Neural - Processing Units (vNPUs) take turns using the same operation unit. Taking the case of having two vNPUs (i.e., vNPU0 and vNPU1) as an example, the priority right of vNPU0 is directly fixed, and a timer is used to control that the operation unit is used by vNPU0 within a certain specific time period and by vNPU1 within another specific time period, so as to achieve the time - sharing sharing of the same operation unit by multiple vNPUs.

[0004] However, the above - mentioned time - division multiplexing method will cause waste of computing power: First, since the time allocated to a certain vNPU is longer than the actual simulation running time of the vNPU, the operation unit remains idle after the vNPU actually completes the simulation task, and this idle time cannot be utilized by other vNPUs, reducing the overall computing efficiency of the chip. Second, this design has low flexibility. The directly fixed hardware priority and the rotation method cannot be dynamically adjusted according to the actual running situation, restricting the performance and resource utilization efficiency of the chip in different situations. Summary of the Invention

[0005] To solve the above problems, the present application provides a method and apparatus for controlling an operation unit lock and an electronic device, which can solve the technical problem that the utilization rate of the operation unit in the NPU is low, resulting in low computing efficiency of the NPU.

[0006] To achieve the above object, in a first aspect, the present application provides an arithmetic unit lock control method, which is applied to an NPU. The NPU includes at least one arithmetic unit and a distributor. The distributor runs at least one vNPU. The arithmetic unit lock control method includes: in response to a pending authorization instruction, obtaining the computing power of the target vNPU that the distributor is running, where the pending authorization instruction is used to instruct the target vNPU to obtain authorization; if the computing power of the target vNPU meets a preset condition, the pending authorization instruction is executed successfully, and the target vNPU is enabled to obtain the control right for at least one arithmetic unit; in response to a pending work end instruction, obtaining the working states of at least one arithmetic unit, where the pending work end instruction is used to instruct at least one arithmetic unit to feedback the working states, and the working states include work end; if the working states of all at least one arithmetic unit are work end, the pending work end instruction is executed successfully, and the control right of the target vNPU for at least one arithmetic unit is released.

[0007] The arithmetic unit lock control method provided by the present application can determine whether the target vNPU can obtain the control right of the arithmetic unit by executing the pending authorization instruction. When the target vNPU obtains the control right of the arithmetic unit, other vNPUs cannot control the arithmetic unit anymore, realizing the locking of the arithmetic unit; then, by executing the pending work end instruction, it is judged whether the target vNPU releases the control right of the arithmetic unit. After the target vNPU releases the control right, other vNPUs can obtain the control right of the arithmetic unit according to the pending authorization instruction, realizing the unlocking of the arithmetic unit. This avoids the idle time of the arithmetic unit, improves the utilization rate of the arithmetic unit, and further improves the working efficiency of the NPU.

[0008] In an implementable manner of the first aspect, the distributor includes a current limiting module; in response to the pending authorization instruction, obtaining the computing power of the target vNPU that the distributor is running includes: determining the target vNPU according to the pending authorization instruction; reading the computing power information related to the target vNPU through the current limiting module, where the computing power information includes at least one of the computing power being used by the target vNPU, the remaining computing power, and the computing power usage historical data.

[0009] In the above method, by obtaining the computing power information of the target vNPU through the current limiting module, it can be judged whether the target NPU can obtain the control right of the arithmetic unit according to the computing power information, realizing the locking of the arithmetic unit.

[0010] In an implementable manner of the first aspect, the preset condition includes: within a preset time, the remaining computing power of the target vNPU is greater than or equal to the computing power required to execute the current work; where the computing power required to execute the current work is estimated according to the complexity of the current work and the required computing resources.

[0011] In the above method, the preset time can be the time before the arithmetic unit starts to execute the work of the target vNPU, or the time after the arithmetic unit finishes executing the work of the target vNPU. Judging whether the target vNPU can obtain the control right of the arithmetic unit according to the preset conditions can improve the flexibility of locking the arithmetic unit.

[0012] In an implementable manner of the first aspect, the successful execution of the instruction to be authorized includes: the instruction type of the instruction to be authorized is correct; the instruction to be authorized passes the verification; the target vNPU successfully obtains the usage information of all at least one arithmetic unit.

[0013] In the above method, when the instruction to be authorized is correctly executed, it can obtain the usage information of all arithmetic units, lock the arithmetic units, and improve the utilization rate of the arithmetic units.

[0014] In an implementable manner of the first aspect, if the computing power of the target vNPU meets the preset conditions, after the step of successfully executing the instruction to be authorized and enabling the target vNPU to obtain the control right of at least one arithmetic unit, the arithmetic unit lock control method further includes: obtaining a calculation instruction, where the calculation instruction is used to instruct at least one arithmetic unit to start the arithmetic function; controlling the arithmetic unit to execute the work of the target vNPU according to the calculation instruction.

[0015] In the above method, after the target vNPU obtains the control right of the arithmetic unit, it starts to execute the calculation instruction, causing the arithmetic unit to start executing the work in the target vNPU, realizing the control of the target vNPU over the arithmetic unit.

[0016] In an implementable manner of the first aspect, in response to the instruction for ending work, obtaining the working state of at least one arithmetic unit includes: sending a working state query request to at least one arithmetic unit, where the working state query request is used to request at least one arithmetic unit to continuously send the working state; receiving the working state sent by at least one arithmetic unit.

[0017] In an implementable manner of the first aspect, the successful execution of the instruction for ending work includes: the instruction type of the instruction for ending work is correct; the instruction for ending work passes the verification; the working state of at least one arithmetic unit is successfully obtained.

[0018] In the above method, by executing the instruction for ending work, it can be judged whether the arithmetic unit has completed the task of the target vNPU. If the task is completed, the target vNPU can release the control right of the arithmetic unit, unlock it, avoid idle time of the arithmetic unit, and improve the utilization rate of the arithmetic unit.

[0019] In an implementable manner of the first aspect, the instruction to be authorized is a 64-bit binary instruction, including: bits 0 to 7 are the instruction type bits, used to indicate that the type of the binary instruction is the instruction to be authorized; bits 8 to 12 are the check bits, used to check whether the reading of the instruction to be authorized is correct; bits 13 to 17 are the reserved bits, used for function expansion; bit 18 is the new job judgment bit, used to judge whether the current job is a new job; bit 19 is the marking judgment bit, used to judge whether it is necessary to print the time point when the instruction to be authorized ends; bits 20 to 35 are the arithmetic unit enable flag bits, used to mark whether all at least one arithmetic unit is enabled; bits 36 to 63 are the reserved bits, used for function expansion.

[0020] In the above method, the instruction to be authorized can obtain the enable status of the arithmetic unit, and can also implement extended functions through the reserved bits, improving the working efficiency and scalability of the NPU.

[0021] In an implementable manner of the first aspect, the instruction for job end is a 64-bit binary instruction, including: bits 0 to 7 are the instruction type bits, used to indicate that the type of the binary instruction is the instruction for job end; bits 8 to 12 are the check bits, used to check whether the reading of the instruction for job end is correct; bits 13 to 17 are the reserved bits, used for function expansion; bit 18 is the new job judgment bit, used to judge whether the current job is a new job; bit 19 is the marking judgment bit, used to judge whether it is necessary to print the time point when the instruction for job end ends; bits 20 to 27 are the arithmetic unit working status check bits, used to mark whether all at least one arithmetic unit needs to check the working status; bits 28 to 63 are the reserved bits, used for function expansion.

[0022] In the above method, the instruction for job end can obtain the working status of the arithmetic unit, and can also implement extended functions through the reserved bits, improving the working efficiency and scalability of the NPU.

[0023] Second aspect, the present application further provides an arithmetic unit lock control device, which is applied to an NPU. The NPU includes at least one arithmetic unit and a distributor, and the distributor runs at least one vNPU. The arithmetic unit lock control device includes: a computing power acquisition module, configured to: in response to an authorization instruction to be processed, acquire the computing power of the target vNPU that the distributor is running, and the authorization instruction to be processed is used to instruct the target vNPU to obtain authorization; a locking module, configured to: if the computing power of the target vNPU meets a preset condition and the authorization instruction to be processed is successfully executed, enable the target vNPU to obtain the control right over at least one arithmetic unit; a working state acquisition module, configured to: in response to an instruction for indicating the end of work, acquire the working state of at least one arithmetic unit, and the instruction for indicating the end of work is used to instruct at least one arithmetic unit to feedback the working state, and the working state includes the end of work; an unlocking module, configured to: if the working states of all at least one arithmetic unit are the end of work and the instruction for indicating the end of work is successfully executed, release the control right of the target vNPU over at least one arithmetic unit.

[0024] Third aspect, the present application further provides an electronic device, which includes: one or more processors; a memory, configured to store one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the arithmetic unit lock control method in the first aspect and any one of its optional implementation manners.

[0025] It can be understood that for the beneficial effects that can be achieved by the technical solutions provided in the above second aspect and third aspect, reference can be made to the beneficial effects in the first aspect and any one of its optional implementation manners, and details are not described herein again.

[0026] As can be seen from the above technical solutions, the present application provides an arithmetic unit lock control method, device and electronic device, which are applied to an NPU. The NPU includes at least one arithmetic unit and a distributor, and the distributor runs at least one vNPU. The arithmetic unit lock control method includes: in response to an authorization instruction to be processed, acquire the computing power of the target vNPU that the distributor is running, and the authorization instruction to be processed is used to instruct the target vNPU to obtain authorization; if the computing power of the target vNPU meets a preset condition and the authorization instruction to be processed is successfully executed, enable the target vNPU to obtain the control right over at least one arithmetic unit; in response to an instruction for indicating the end of work, acquire the working state of at least one arithmetic unit, and the instruction for indicating the end of work is used to instruct at least one arithmetic unit to feedback the working state, and the working state includes the end of work; if the working states of all at least one arithmetic unit are the end of work and the instruction for indicating the end of work is successfully executed, release the control right of the target vNPU over at least one arithmetic unit.

[0027] The arithmetic unit lock control method provided by this application can determine whether the target vNPU can obtain the control right of the arithmetic unit by executing the instruction to be authorized. When the target vNPU obtains the control right of the arithmetic unit, other vNPUs cannot control the arithmetic unit anymore, realizing the locking of the arithmetic unit. Then, it is judged whether the target vNPU releases the control right of the arithmetic unit by executing the instruction for the end of work. After the target vNPU releases the control right, other vNPUs can obtain the control right of the arithmetic unit according to the instruction to be authorized, realizing the unlocking of the arithmetic unit. This can avoid the idle time of the arithmetic unit, improve the utilization rate of the arithmetic unit, and further improve the working efficiency of the NPU. Brief Description of the Drawings

[0028] In order to more clearly illustrate the technical solutions of this application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0029] Figure 1 Schematic diagram of an NPU structure provided by an embodiment of this application;

[0030] Figure 2 Schematic diagram of an arithmetic unit lock control method provided by an embodiment of this application;

[0031] Figure 3 Schematic diagram of a method for determining whether a target vNPU obtains the control right of an arithmetic unit provided by an embodiment of this application;

[0032] Figure 4 Schematic diagram of an arithmetic unit lock control device provided by an embodiment of this application;

[0033] Figure 5 Schematic diagram of an electronic device provided by an embodiment of this application. Detailed Embodiments

[0034] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with this application.

[0035] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the embodiments described next, rather than intending to limit the embodiments of this application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.

[0036] In this application, terms such as "first", "second", "third", etc. in the specification and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms used can be interchanged under appropriate circumstances.

[0037] For the convenience of understanding the solution, the following explains relevant terms:

[0038] Neural network processor: A neural network processor is a processor specifically designed to process neural network algorithms. It can efficiently execute mathematical operations involved in deep learning and machine learning tasks, such as matrix multiplication, convolution operations, etc.

[0039] Allocator: The allocator is a key hardware in the neural network processor. Its main function is to allocate and schedule data and computing resources to ensure the flow of data and computing tasks between various modules of the neural network processor.

[0040] Arithmetic unit: The arithmetic unit is the core hardware responsible for executing various mathematical operations in the neural network algorithm. It is the basic execution unit for the NPU to complete computing tasks and directly participates in the computing process of the neural network, such as performing operations like matrix multiplication, addition, convolution operation, pooling operation, activation function calculation, etc.

[0041] Instruction: Instructions are used to indicate to the neural network processor how to perform various basic operations, such as data reading, storage, arithmetic operations (addition, subtraction, multiplication, division), logical operations (AND, OR, NOT), control flow operations (jump, loop, branch), etc.

[0042] The NPU can process neural network algorithms through the arithmetic unit to implement the functions of neural network algorithms, such as image processing, speech recognition, etc. When using the NPU, there are requirements for simultaneously processing multiple different functions and data, that is, the same NPU needs to perform image processing and speech recognition at the same time. This requirement can be achieved through the hard segmentation method, but the hard segmentation implementation method requires multiple copies of arithmetic units and memories, resulting in an increase in the chip area of the NPU and an increase in manufacturing costs. Therefore, the time-division multiplexing method can be used to achieve the multiplexing of arithmetic units.

[0043] The basic logic of the time-division multiplexing method is that multiple vNPUs can run in the same NPU, and each vNPU represents a neural network algorithm with different functions. Multiple vNPUs take turns using the same arithmetic unit. Taking the case of having two vNPUs (i.e., vNPU0 and vNPU1) as an example, directly fix the priority usage right of vNPU0, and use a timer to control that the arithmetic unit is used by vNPU0 within a certain specific time period and by vNPU1 within another specific time period, so as to achieve the time-sharing sharing of the same arithmetic unit by multiple vNPUs.

[0044] However, since the time allocated to a certain vNPU is longer than the time when the vNPU actually runs the simulation, the arithmetic unit remains idle after the vNPU actually completes the simulation task, and this idle time cannot be utilized by other vNPUs, reducing the overall arithmetic efficiency and utilization rate of the NPU.

[0045] To solve the above problems, the embodiments of the present application provide an arithmetic unit lock control method, device and electronic device. By adding a to-be-authorized instruction, the target vNPU executing the to-be-authorized instruction obtains the control right of the arithmetic unit; and by the to-work-end instruction, the vNPU currently having the control right of the arithmetic unit releases its control right, so as to realize time-division multiplexing of the arithmetic unit based on lock control, avoid the generation of idle time of the arithmetic unit, and improve the arithmetic efficiency and utilization rate of the NPU.

[0046] In some embodiments, the instructions for instructing the NPU to work can be generated by a programming language (such as C, C++, Java, etc.), packaged and then sent to the NPU.

[0047] Figure 1 This is a schematic diagram of the NPU structure provided by the embodiments of the present application.

[0048] As Figure 1 shown, the NPU includes an on-chip memory (On Chip Memory, OCM), a distributor, and at least one arithmetic unit. For example, arithmetic unit 1, arithmetic unit 2, arithmetic unit 3, and arithmetic unit 4. Among them, the distributor includes a current limiting module, a decoding module, and an instruction loading module, and the arithmetic unit includes a control module and a logic processing module (taking arithmetic unit 1 as an example in the figure).

[0049] In some embodiments, the instructions are first stored in a double data rate synchronous dynamic random access memory (Double Data Rate Synchronous Dynamic Random Access Memory, DDR). The OCM in the NPU needs to obtain the instructions from the DDR to instruct the NPU to run the neural network algorithm. Figure 1 The NPU in is in a two-level control mode. First, the instruction loading module in the distributor reads the first-level instructions, and then the decoding module decodes the first-level instructions. The decoded first-level instructions are used to instruct the arithmetic unit to perform corresponding work. At this time, the control module in the arithmetic unit obtains and decodes the second-level instructions according to the instructions of the first-level instructions, and the decoded second-level instructions are used to instruct the logic processing module to perform specific arithmetic work.

[0050] Next, based on Figure 1The NPU shown below introduces the specific implementation method of the operation unit lock control method. It should be understood that the operation unit lock control method provided by the embodiments of the present application is used to meet the requirement that an NPU can process multiple different functions and data simultaneously. Therefore, multiple vNPUs can run in the allocator of the NPU, and different vNPUs are neural network models for implementing different functions.

[0051] Figure 2 This is a schematic diagram of an operation unit lock control method provided by the embodiments of the present application. As Figure 2 shown, the operation unit lock control method includes steps S100 - S400.

[0052] S100: In response to the instruction to be authorized, obtain the computing power of the target vNPU running in the allocator.

[0053] In some embodiments, to achieve time - division multiplexing of multiple vNPUs for the operation units in the NPU and improve the flexibility of multiple vNPUs running in the allocator of the NPU to obtain the control right of the operation units, when generating instructions for instructing the NPU to work by a programming language, an instruction to be authorized can be added to indicate the target vNPU to obtain authorization.

[0054] Specifically, the instruction to be authorized can be a 64 - bit binary instruction, including: bits 0 to 7 are instruction type bits, used to indicate that the type of the binary instruction is an instruction to be authorized; bits 8 to 12 are check bits, used to check whether the reading of the instruction to be authorized is correct; bits 13 to 17 are reserved bits for function expansion; bit 18 is a new work judgment bit, used to judge whether the current work is a new work; bit 19 is a marking judgment bit, used to judge whether it is necessary to print the time point when the instruction to be authorized ends; bits 20 to 35 are operation unit enable flag bits, used to mark whether all at least one operation unit is enabled; bits 36 to 63 are reserved bits for function expansion.

[0055] Exemplarily, Table 1 shows the instruction format of an instruction to be authorized.

[0056] Table 1

[0057] 。

[0058] In the to-be-authorized instruction shown in Table 1, bits 0 to 7 indicate that the instruction type is a to-be-authorized instruction, which is represented by an 8-bit hexadecimal number in the table and converted to binary as 10101010; bits 8 to 12 are the CRC check value of the instruction, which is used to check whether an error occurs during the instruction reading process; bits 13 to 17 indicate reserved empty positions, which can be used to add other functions to the instruction subsequently; bit 18 indicates whether the current work is a new work. For example, 0 indicates it is not a new work, and 1 indicates it is a new work; bit 19 indicates whether to print the time point when the instruction ends. For example, 0 indicates not to print, and 1 indicates to print; bits 20 to 35 indicate the enabling status of each arithmetic unit during the control process of the current vNPU for the arithmetic units. For example, there are 16 arithmetic units in total in the NPU, and each bit in bits 20 to 35 represents the enabling status of an arithmetic unit. 0 indicates not enabled, and 1 indicates enabled. In addition, if the number of arithmetic units in the NPU is less than 16, only the values of the same number of bit positions as the number of arithmetic units are set, and the other bit positions are left empty. For example, if there are 8 arithmetic units in the NPU, only the values of bits 20 to 27 are set, and bits 28 to 35 are left empty. Then, the values of bits 20 to 35 can be 0000000010011101, indicating that during the control process of the current vNPU for the arithmetic units, the 0th, 2nd, 3rd, 4th, and 7th arithmetic units are enabled; bits 36 to 63 are reserved empty positions, which are used to set the enabling status of the arithmetic units exceeding 16 when the number of arithmetic units in the NPU exceeds 16. For example, if there are 20 arithmetic units in the NPU, bits 36 to 39 are used to represent the enabling status of the additional 4 arithmetic units exceeding 16. 0 indicates not enabled, and 1 indicates enabled.

[0059] It should be noted that the to-be-authorized instruction can also be in other instruction formats, as long as it can represent the enabling status of the arithmetic units. In the embodiments of the present application, no specific limitations are made.

[0060] In some embodiments, assume that two vNPUs, vNPU0 and vNPU1, can run in the allocator, and the currently running vNPU in the current allocator is vNPU0. Then, vNPU0 is denoted as the target vNPU. At this time, the target vNPU needs to obtain the control right of the arithmetic units. Therefore, when the allocator runs to the to-be-authorized instruction of the target vNPU, it is necessary to determine whether the to-be-authorized instruction can be executed successfully according to the computing power of the target vNPU, that is, whether the target vNPU can obtain the control right of the arithmetic units. Therefore, it is necessary to obtain the computing power of the target vNPU.

[0061] In some embodiments, to obtain the computing power of the target vNPU, it is first necessary to determine the target vNPU according to the to-be-authorized instruction. For example, it is determined that the target vNPU is vNPU0.

[0062] Secondly, read the computing power information related to the target vNPU through the flow control module. The computing power information includes at least one of the computing power currently used by the target vNPU, the remaining computing power, and the historical data of computing power usage. In some embodiments, the computing power information of the vNPU can be determined according to the token bucket principle. Specifically, a token bucket is allocated for each vNPU, and a preset number of tokens are added to the token bucket at every preset time interval; during the operation of the vNPU, a preset number of tokens are deducted from the token bucket of the vNPU in each operation cycle. Based on this, the flow control module can obtain the number of tokens in the token bucket of the target vNPU at any time, that is, the computing power of the target vNPU.

[0063] It should be noted that, in addition to the token bucket principle, other methods can also be used to obtain the computing power information related to the target vNPU, which is not specifically limited in the embodiments of the present application.

[0064] S200: If the computing power of the target vNPU meets the preset conditions, the to-be-authorized instruction is executed successfully, and the target vNPU obtains the control right for at least one arithmetic unit.

[0065] In some embodiments, to improve the accuracy of locking the arithmetic unit, it is necessary to determine whether the target vNPU can obtain the control right of the arithmetic unit according to the computing power of the target vNPU.

[0066] Figure 3 It is a schematic diagram of a method for determining whether the target vNPU obtains the control right of the arithmetic unit provided by the embodiments of the present application.

[0067] As Figure 3 shown, first, step S301 is executed: obtain the computing power of the target vNPU. For example, there are 10 tokens in the token bucket of the target vNPU.

[0068] Then, step S302 is executed: determine whether the computing power of the target vNPU meets the preset conditions. In some embodiments, the preset conditions may be: within a preset time, the remaining computing power of the target vNPU is greater than or equal to the computing power required to execute the current work; wherein, the computing power required to execute the current work is estimated according to the complexity of the current work and the required computing resources.

[0069] It should be noted that the computing power required for the target vNPU to execute the current work can be set according to actual needs. For example, it is set that there are no less than 1 token in the token bucket of the target vNPU, or it is set that the number of tokens in the token bucket of the target vNPU is greater than or equal to the sum of the number of tokens required for the target vNPU to complete the current work, which is not specifically limited in the embodiments of the present application.

[0070] Exemplarily, if the preset time is the time point when the to-be-authorized instruction is successfully executed, that is, the time point when the target vNPU obtains the control right of the arithmetic unit and controls the arithmetic unit to start executing the current work of the target vNPU, then the number of tokens in the token bucket of the target vNPU is not less than 1, which means that the preset condition is satisfied.

[0071] Exemplarily again, if the preset time is the time point when the arithmetic unit finishes executing the current work of the target vNPU, then the number of tokens in the token bucket of the target vNPU needs to be greater than or equal to the total number of tokens required for the target vNPU to complete the current work to meet the preset condition.

[0072] In some embodiments, if the computing power of the target vNPU meets the preset condition, then step S303 is executed: the to-be-authorized instruction is successfully executed, and the target vNPU obtains the control right for at least one arithmetic unit.

[0073] Specifically, for the to-be-authorized instruction to be successfully executed, the following conditions need to be met: the instruction type of the to-be-authorized instruction is correct; the to-be-authorized instruction passes the verification; the target vNPU successfully obtains the usage information for all at least one arithmetic unit. For example, if the instruction type of the to-be-authorized instruction is 0XAA and it passes the verification, the bit information indicating the enabling status of all arithmetic units in the NPU in the to-be-authorized instruction is correct, and the allocator determines that the computing power of the target vNPU meets the preset condition, then the to-be-authorized instruction is successfully executed. At this time, the target vNPU running this to-be-authorized instruction can obtain the control right of the arithmetic unit, that is, the target vNPU starts to control the arithmetic unit to execute the work, and other vNPUs except the target vNPU cannot use the arithmetic unit, realizing the locking of the arithmetic unit by the target vNPU.

[0074] It should be noted that the conditions for the successful execution of the to-be-authorized instruction are determined by its instruction format, and are not specifically limited in the embodiments of the present application.

[0075] In some embodiments, if the computing power of the target vNPU does not meet the preset condition, then step S301 is executed again: obtain the computing power of the target vNPU. Until the computing power of the target vNPU meets the preset condition, the target vNPU can obtain the control right of the arithmetic unit through steps S302 and S303.

[0076] In some embodiments, if the computing power of the target vNPU meets the preset condition and the to-be-authorized instruction is successfully executed, after the target vNPU obtains the control right for at least one arithmetic unit, the arithmetic unit starts to execute the work of the target vNPU: First, the arithmetic unit reads the calculation instruction, starts the arithmetic function, and then executes the corresponding work according to the calculation instruction.

[0077] In some embodiments, the computing instructions are used to instruct the arithmetic units to perform corresponding tasks. The arithmetic units are mainly divided into two categories, namely the data processing unit and the data transmission unit. Among them, the data processing unit is used to implement data calculations, including convolution operations, addition operations, multiplication operations, division operations, etc. The data transmission unit is used to implement data transmission tasks, such as transmitting the data required to implement the target vNPU from the DDR to the OCM.

[0078] S300: In response to the pending work end instruction, obtain the working status of at least one arithmetic unit.

[0079] In some embodiments, after all the arithmetic units required for the target vNPU to execute the current work have been started by the computing instructions, the target vNPU will execute the pending work end instruction and instruct the started arithmetic units to feedback the working status, where the working status includes work end.

[0080] Specifically, the pending work end instruction can be a 64-bit binary instruction, including: bits 0 to 7 are the instruction type bits, used to indicate that the type of the binary instruction is the pending work end instruction; bits 8 to 12 are the parity bits, used to check whether the pending work end instruction is read correctly; bits 13 to 17 are the reserved bits, used for function expansion; bit 18 is the new work judgment bit, used to judge whether the current work is a new work; bit 19 is the marking judgment bit, used to judge whether it is necessary to print the time point when the pending work end instruction ends; bits 20 to 27 are the arithmetic unit working status check bits, used to mark whether all at least one arithmetic unit needs to check the working status; bits 28 to 63 are the reserved bits, used for function expansion.

[0081] Exemplarily, Table 2 shows the instruction format of a pending work end instruction.

[0082] Table 2

[0083] 。

[0084] In the pending work completion instruction shown in Table 2, bits 0 to 7 indicate that the instruction type is a pending work completion instruction, which is represented by an 8-bit hexadecimal number in the table and converted to binary as 10100011; bits 8 to 12 are the CRC check value of the instruction, used to check whether an error occurs during the instruction reading process; bits 13 to 17 represent reserved empty positions, which can be used to add other functions to the instruction in the future; bit 18 indicates whether the current work is a new work, for example, 0 means it is not a new work, and 1 means it is a new work; bit 19 indicates whether the time point when the instruction ends needs to be printed, for example, 0 means no printing is required, and 1 means printing is required; bits 20 to 27 represent the arithmetic units for which the current pending work completion instruction needs to wait for the end of work. For example, there are 8 arithmetic units in the NPU in total. Each bit in bits 20 to 27 indicates whether to wait for the end of work of an arithmetic unit. 0 means not waiting for the end of work of the arithmetic unit, and 1 means waiting for the end of work of the arithmetic unit.

[0085] It should be noted that in some embodiments, since the arithmetic units are divided into two major categories: data processing units and data transmission units, and during the operation of the NPU, the working states of the arithmetic units in the same category can be synchronized. Therefore, to simplify the instructions and improve the working efficiency of the NPU, each bit in bits 20 to 27 can indicate whether to wait for the end of work of the arithmetic units in the same category. For example, bit 20 indicates whether to wait for the end of work of the data processing units, and bit 21 indicates whether to wait for the end of work of the data transmission units.

[0086] Specifically, assume that there are 8 arithmetic units in the NPU, including 6 data processing units and 2 data transmission units. Then bit 20 indicates whether to wait for the end of work of the 6 data processing units, and bit 21 indicates whether to wait for the end of work of the 2 data transmission units. When there is also an arithmetic unit with a third function in the NPU, bit 22 can be used to indicate whether to wait for its end of work.

[0087] Based on this, if the number of arithmetic unit classifications in the NPU is less than 8, only the values of the same number of bits as the number of arithmetic unit classifications are set, and the other bits are left empty. For example, if there are 2 types of arithmetic unit classifications in the NPU, only the values of bits 20 and 21 are set, and bits 22 to 27 are left empty. Then the values of bits 20 to 27 can be 00000011, indicating that it is necessary to wait for both the data transmission units and the data processing units to end their work.

[0088] It should be noted that when the target vNPU is running, multiple arithmetic units are required to cooperate to achieve the function. For example, data transmission needs to be realized through the data transmission unit, and then the calculation function is realized through the data processing unit. After the data transmission unit completes the data transmission required for the first task in the target vNPU, the data processing unit can process the data required for the first task. At this time, the data transmission unit can continue to transmit the data required for the second task in the target vNPU. Therefore, to improve the working efficiency of the NPU, the instruction to wait for the end of the task does not need to wait until both the data processing unit and the data transmission unit have completed their work before it can be successfully executed. Instead, it can be successfully executed after one of the types of arithmetic units has completed its work, so as to start executing the next task. For example, bits 20 and 27 of the instruction to wait for the end of the task can be 00000010, indicating that only the data transmission unit needs to wait for its work to end, and there is no need to wait for the data processing unit to end its work. If the instruction is successfully executed, the data transmission unit can execute the next task and transmit the data required for the next task; or, bits 20 and 27 of the instruction to wait for the end of the task can be 00000001, indicating that only the data processing unit needs to wait for its work to end, and there is no need to wait for the data transmission unit to end its work. If the instruction is successfully executed, the data processing unit can execute the next task and perform the data calculation for the next task.

[0089] In some embodiments, bits 28 to 63 of the instruction to wait for the end of the task are reserved empty positions, which are used to set whether to wait for the end of the work of the arithmetic unit categories exceeding 8 when the number of arithmetic unit categories in the NPU exceeds 8. For example, if there are 9 types of arithmetic units in the NPU, bit 28 is used to indicate whether to wait for the end of the work of the additional 1 arithmetic unit category exceeding 8 types. 0 indicates that there is no need to wait for the work to end, and 1 indicates that it is necessary to wait for the work to end.

[0090] It should be noted that the instruction to wait for the end of the task can also be in other instruction formats, as long as it can represent the working state of the arithmetic unit. There is no specific limitation in the embodiments of the present application.

[0091] In some embodiments, when the instruction to wait for the end of the task is executed, it is necessary to check the working state continuously sent by the arithmetic unit so that the vNPU can release the control right of the arithmetic unit in a timely manner.

[0092] S400: If the working state of all at least one arithmetic unit is the end of the work, the instruction to wait for the end of the task is successfully executed, and the control right of the target vNPU for at least one arithmetic unit is released.

[0093] In some embodiments, for the successful execution of the end-of-work instruction, the following conditions need to be met: the instruction type of the end-of-work instruction is correct; the end-of-work instruction passes the verification; and the target vNPU successfully obtains the working status of at least one type of arithmetic unit. For example, if the instruction type of the end-of-work instruction is 0XA3 and it passes the verification, and the bit information indicating the working status of at least one type of arithmetic unit in the NPU in the end-of-work instruction is correct, and the work of the arithmetic units to be waited for is all completed, then the end-of-work instruction is successfully executed.

[0094] It should be noted that the conditions for the successful execution of the end-of-work instruction are determined by its instruction format, and no specific limitations are made in the embodiments of the present application.

[0095] In some embodiments, since the end-of-work instruction can be set separately for different types of arithmetic units, in order to ensure that all arithmetic units in the NPU are unlocked simultaneously for subsequent use by the vNPU, it is necessary to wait for the work of all types of arithmetic units to end, that is, after the end-of-work instruction indicating the end of the work of all types of arithmetic units is successfully executed, the target vNPU can release the control right of all arithmetic units to achieve the unlocking of the arithmetic units.

[0096] Specifically, assume that there are two end-of-work instructions for the target vNPU. The 20th to 27th bits of the first instruction are 00000001, indicating the end of the work of the data processing unit; the 20th to 27th bits of the second instruction are 00000010, indicating the end of the work of the data transmission unit. When both of the above two end-of-work instructions are successfully executed, the target vNPU can release the control right of the arithmetic unit, that is, the target vNPU no longer controls the arithmetic unit to execute work. At this time, all vNPUs running in the NPU can obtain the control right of the arithmetic unit according to the authorization instruction to achieve the unlocking of the arithmetic unit.

[0097] Some embodiments of the present application also provide an arithmetic unit lock control device, which is applied to the NPU. The NPU includes at least one arithmetic unit and a distributor, and the distributor runs at least one vNPU. Figure 4 For the schematic diagram of an arithmetic unit lock control device provided by the embodiments of the present application. As Figure 4 shown, the arithmetic unit lock control device includes:

[0098] The computing power acquisition module 401 is configured to: in response to an authorization pending instruction, acquire the computing power of the target vNPU that the allocator is running, where the authorization pending instruction is used to instruct the target vNPU to obtain authorization; the locking module 402 is configured to: if the computing power of the target vNPU meets a preset condition and the authorization pending instruction is successfully executed, enable the target vNPU to obtain the control right over at least one arithmetic unit; the working state acquisition module 403 is configured to: in response to an instruction pending work end, acquire the working state of at least one arithmetic unit, where the instruction pending work end is used to instruct at least one arithmetic unit to feedback the working state, and the working state includes work end; the unlocking module 404 is configured to: if the working state of all at least one arithmetic unit is work end and the instruction pending work end is successfully executed, release the control right of the target vNPU over at least one arithmetic unit.

[0099] Some embodiments of the present application further provide an electronic device, including: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the arithmetic unit lock control method provided by some embodiments of the present application.

[0100] For example, Figure 5 is a schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown, the electronic device includes a processor 501, at least one communication bus 502, a user interface 503, at least one external communication interface 504, and a memory 505. Among them, the communication bus 502 is configured to implement connection communication between these components. Among them, the user interface 503 may include a display screen, and the external communication interface 504 may include a standard wired interface and a wireless interface. Among them, a computer program is stored in the memory 505. The processor 501 is used to execute the computer program stored in the memory 505.

[0101] As can be seen from the above technical solutions, the present application provides an arithmetic unit lock control method, device, and electronic device, which are applied to an NPU. The NPU includes at least one arithmetic unit and an allocator, and the allocator runs at least one vNPU. The arithmetic unit lock control method includes: in response to an authorization pending instruction, acquiring the computing power of the target vNPU that the allocator is running, where the authorization pending instruction is used to instruct the target vNPU to obtain authorization; if the computing power of the target vNPU meets a preset condition and the authorization pending instruction is successfully executed, enabling the target vNPU to obtain the control right over at least one arithmetic unit; in response to an instruction pending work end, acquiring the working state of at least one arithmetic unit, where the instruction pending work end is used to instruct at least one arithmetic unit to feedback the working state, and the working state includes work end; if the working state of all at least one arithmetic unit is work end and the instruction pending work end is successfully executed, releasing the control right of the target vNPU over at least one arithmetic unit.

[0102] The arithmetic unit lock control method provided by this application can determine whether the target vNPU can obtain the control right of the arithmetic unit by executing the to-be-authorized instruction. When the target vNPU obtains the control right of the arithmetic unit, other vNPUs cannot control the arithmetic unit anymore, realizing the locking of the arithmetic unit; then, it is judged whether the target vNPU releases the control right of the arithmetic unit by executing the to-be-work-ended instruction. After the target vNPU releases the control right, other vNPUs can obtain the control right of the arithmetic unit according to the to-be-authorized instruction, realizing the unlocking of the arithmetic unit. This can avoid the idle time of the arithmetic unit, improve the utilization rate of the arithmetic unit, and further improve the working efficiency of the NPU.

[0103] For the similar parts between the embodiments provided by this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.

Claims

1. A method for controlling a computing unit lock, characterized in that: Applied to a neural network processor, the neural network processor is a processor for processing a neural network algorithm, capable of performing mathematical operations involved in deep learning and machine learning tasks, the neural network processor includes at least one operation unit and a distributor, the operation unit is the core hardware responsible for performing various mathematical operations in the neural network algorithm, is the execution unit of the neural network processor to complete the computing task, and directly participates in the computing process of the neural network, the distributor is the hardware in the neural network processor, distributes and schedules data and computing resources, ensures that data and computing tasks flow between various modules of the neural network processor, and the distributor runs at least one virtual neural network processor, the method includes: In response to the pending authorization instruction, obtaining the computing power of the target virtual neural network processor being run by the allocator, wherein the pending authorization instruction is used to instruct the target virtual neural network processor to obtain authorization; If the computing power of the target virtual neural network processor meets the preset conditions, the to-be-authorized instruction is executed successfully, so that the target virtual neural network processor obtains the control right over the at least one computing unit; the preset conditions include that within a preset time, the remaining computing power of the target virtual neural network processor is greater than or equal to the computing power required to perform the current work; wherein the computing power required to perform the current work is estimated according to the complexity of the current work and the required computing resources; In response to a waiting work completion instruction, acquiring a working state of the at least one operation unit, the waiting work completion instruction is used to instruct the at least one operation unit to feedback a working state, the working state including the completion of work; If the working status of all of the at least one computing unit is completion of work, the waiting work completion instruction is executed successfully, and the control of the target virtual neural network processor over the at least one computing unit is released.

2. The operation unit lock control method according to claim 1, characterized in that: The distributor includes a current limiting module; The step of obtaining the computing power of the target virtual neural network processor being run by the allocator in response to the instruction to be authorized includes: Determine the target virtual neural network processor according to the instruction to be authorized; The computing power information related to the target virtual neural network processor is read through the current limiting module, and the computing power information includes at least one of the computing power currently being used by the target virtual neural network processor, the remaining computing power, and the computing power usage history data.

3. The operation unit lock control method according to claim 1, characterized in that: The authorization-to-be-authorized instruction is successfully executed, including: The instruction type of the instruction to be authorized is correct; The instruction to be authorized is verified and passed; The target virtual neural network processor successfully obtains usage information of all of the at least one computing units.

4. The operation unit lock control method according to claim 1, characterized in that: After the step of if the computing power of the target virtual neural network processor meets the preset conditions and the to-be-authorized instruction is executed successfully, so that the target virtual neural network processor acquires the control right over the at least one computing unit, the method further includes: Acquire a calculation instruction, where the calculation instruction is used to instruct the at least one calculation unit to start a calculation function; The at least one computing unit is controlled according to the computing instruction to execute the work of the target virtual neural network processor.

5. The operation unit lock control method according to claim 1, characterized in that: The step of obtaining the working state of the at least one computing unit in response to the instruction to complete the work includes: Sending a working status query request to the at least one computing unit, wherein the working status query request is used to request the at least one computing unit to continuously send the working status; Receive the working status sent by the at least one computing unit.

6. The operation unit lock control method according to claim 1, characterized in that: The execution of the waiting work completion instruction is successful, including: The instruction type of the pending work completion instruction is correct; The waiting work completion instruction is verified to be passed; The working status of the at least one computing unit is obtained successfully.

7. The operation unit lock control method according to claim 1, characterized in that: The instruction to be authorized is a 64-bit binary instruction, including: Bits 0 to 7 are instruction type bits, used to indicate that the type of the binary instruction is the instruction to be authorized; Bits 8 to 12 are check bits, used to check whether the reading of the instruction to be authorized is correct; Bits 13 to 17 are reserved for function expansion; Bit 18 is the new job judgment bit, which is used to judge whether the current job is a new job; Bit 19 is a dot determination bit, which is used to determine whether it is necessary to print the time point at which the instruction to be authorized ends; Bits 20 to 35 are operation unit enable flag bits, used to mark whether all of the at least one operation unit is enabled; Bits 36 to 63 are reserved for function expansion.

8. The operation unit lock control method according to claim 1, characterized in that: The waiting work completion instruction is a 64-bit binary instruction, including: Bits 0 to 7 are instruction type bits, used to indicate that the type of the binary instruction is the to-be-finished instruction; Bits 8 to 12 are check bits, used to check whether the reading of the pending work completion instruction is correct; Bits 13 to 17 are reserved for function expansion; Bit 18 is the new job judgment bit, which is used to judge whether the current job is a new job; Bit 19 is a dot determination bit, which is used to determine whether it is necessary to print the time point at which the instruction to wait for the end of the work is finished; Bits 20 to 27 are bits for checking the working status of the computing unit, used to mark whether the working status of all of the at least one computing unit needs to be checked; Bits 28 to 63 are reserved for function expansion.

9. A computing unit lock control device, characterized in that: Applied to a neural network processor, the neural network processor is a processor for processing neural network algorithms, capable of performing mathematical operations involved in deep learning and machine learning tasks, the neural network processor includes at least one operation unit and a distributor, the operation unit is the core hardware responsible for performing various mathematical operations in the neural network algorithm, is the execution unit of the neural network processor to complete the computing tasks, and directly participates in the computing process of the neural network, the distributor is the hardware in the neural network processor, allocates and schedules data and computing resources, ensures that data and computing tasks flow between various modules of the neural network processor, the distributor runs at least one virtual neural network processor, the device includes: A computing power acquisition module is configured to: acquire the computing power of the target virtual neural network processor being run by the allocator in response to an instruction to be authorized, wherein the instruction to be authorized is used to instruct the target virtual neural network processor to obtain authorization; The locking module is configured to: if the computing power of the target virtual neural network processor meets a preset condition, the to-be-authorized instruction is executed successfully, so that the target virtual neural network processor obtains control over the at least one computing unit; the preset condition includes that within a preset time, the remaining computing power of the target virtual neural network processor is greater than or equal to the computing power required to perform the current work; wherein the computing power required to perform the current work is estimated according to the complexity of the current work and the required computing resources; A working status acquisition module is configured to: acquire the working status of the at least one operation unit in response to a waiting work completion instruction, wherein the waiting work completion instruction is used to instruct the at least one operation unit to feedback the working status, wherein the working status includes the completion of the work; The unlocking module is configured to: if the working status of all of the at least one computing unit is completed, the waiting work completion instruction is executed successfully, and the control of the target virtual neural network processor over the at least one computing unit is released.

10. An electronic device, characterized in that: include: one or more processors; a memory configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the operation unit lock control method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Neural network model processing method and device

    CN116187391A

  • Strategy neural network training and role control method and device, and electronic equipment

    CN117414580A